Model deployment method and communication device
By receiving and utilizing the identification information of the target model group, deploying and obtaining description files to achieve joint inference, multiple model problems in the prior art that joint inference cannot be deployed, and effective joint inference and intelligent deployment of the model group are realized.
Patent Information
- Application Number
- CN202311777250.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art cannot effectively deploy multiple models that can jointly implement functions, resulting in the inability to implement joint inference functions.
By receiving the identification information of the target model group, deploying a model group that has been jointly trained and/or joint tested based on the identification information, and obtaining a description file about the model group to achieve joint inference between the model groups.
The deployment of the target model group that can jointly implement inference functions is realized, ensuring that the model group can effectively perform joint inference and improving the intelligence level of network equipment.
Smart Images

Figure CN120200902A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communications, and in particular, to a model deployment method and a communication device. Background Art
[0002] In order to improve the intelligence and automation levels of the network, artificial intelligence (AI) and machine learning (ML) technologies are also gradually being applied. For example, there may be a type of service module (such as an inference function) in an access network device that implements functions through an ML model. The access network device can implement corresponding functions based on the ML model to improve the intelligence level of the access network device.
[0003] Generally speaking, each model can independently support a certain specific type of function. Some model deployment methods can deploy one or more independent models for a device that needs to deploy a model (such as the above-mentioned access network device). However, in practice, the implementation of a function may require multiple models to cooperate. For example, using the output of one model as the input of another model to form a sequence of interconnected models, or providing outputs in parallel through multiple models, or merging the parallel outputs of multiple models as the input of another model, and realizing functions through such a combination of multiple models. These multiple models can be jointly trained and / or jointly tested models.
[0004] However, currently, there is no deployment method for such models that can jointly implement functions. Summary of the Invention
[0005] This application provides a model deployment method and a communication device, which are beneficial to deploying a model group capable of implementing a joint inference function.
[0006] In a first aspect, this application provides a model deployment method. This method can be executed by a second device or a component (such as a chip, a chip system, etc.) in the second device. The second device has the ability to provide a model loading service. The second device can be a network management system, a network element management system, a network device, a radio access network device, or any other device or system that can provide a model loading service. This application does not make specific limitations in this regard. The method includes: receiving identification information of a target model group, where the target model group includes multiple models, and the multiple models are jointly trained and / or jointly tested; and deploying the target model group based on the identification information of the target model group.
[0007] It should be understood that the purpose of deploying a model is to make the deployed model available on the entity that implements the inference function. In a possible implementation, when deploying a model, the entity that implements the inference function reads the description file of the model, which can be a description file for the capabilities of the model and needs to be obtained based on the identification information of the model. Although the existing model deployment methods can deploy multiple models, on the one hand, the multiple models may be independently trained or independently tested, and the purpose of joint inference cannot be achieved. On the other hand, even if these multiple models are jointly trained and jointly tested, since the multiple models are deployed in their own independent ways based on their respective model identifications, only the description files of each model itself can be read, and the description file of the model group cannot be obtained, that is, the joint method between the multiple models cannot be obtained. That is to say, the multiple models can each implement the inference functions they support, but the joint inference function cannot be achieved.
[0008] In the embodiments of the present application, during the process of deploying the target model group, the identification information of the target model group indicates that a target model group that has been jointly trained and / or jointly tested and can jointly implement the inference function is deployed. Thus, the description file of the target model group can be obtained based on the identification information of the target model group during the deployment process. The description file may include the cooperation method of multiple models in the target model group when jointly implementing the inference function. For example, the cooperation method may be using the output of one model in the multiple models as the input of another model to form an interconnected model sequence, or the multiple models provide outputs in parallel, or the parallel outputs of some of the multiple models are merged and used as the input of another part of the models, etc. In this way, the multiple models of the deployed target model group can implement joint inference based on the description file obtained during the deployment process. It can be seen that the model deployment method provided by the present application can achieve the deployment of the target model group that can jointly implement the inference function, which is beneficial for the target model group to implement joint inference.
[0009] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: sending the first identification information and the identification information of the first process, where the first identification information is associated with the message carrying the identification information of the target model group, and the first identification information has a mapping relationship with the first process, and the first process is used to manage the deployment process of the target model group.
[0010] In a possible implementation manner, the identification information of the target model group is carried by a request message, and the first identification information may be the identification information of the request instance created based on the request message.
[0011] In another possible implementation manner, the identification information of the target model group is carried by a configuration message, and the first identification information may be the identification information of the instance created based on the configuration message.
[0012] In combination with the first aspect, in some implementations of the first aspect, receive first indication information, where the first indication information is used to indicate the deployment process management method of the target model group.
[0013] In a possible implementation, the first indication information indicates that the deployment process management method of the target model group is batch management.
[0014] In another possible implementation, the default deployment process management method of the target model group is batch management. That is, no explicit indication is required from either party, for example, it can be agreed upon by the protocol.
[0015] It should be understood that the target model group includes multiple models, and "batch management" can be understood as deploying these multiple models as a whole during the model deployment process.
[0016] In combination with the first aspect, in some implementations of the first aspect, when the deployment process management method of the target model group is batch management, the first process corresponds to the overall deployment process of the target model group.
[0017] Optionally, the number of process instances corresponding to the first process can be one, but this application does not limit this.
[0018] In the embodiments of this application, the overall deployment process of the target model group can be completed through one process, which is beneficial to saving the resources and energy consumption of the second device.
[0019] In combination with the first aspect, in some implementations of the first aspect, the method further includes: sending first information, where the first information includes the first identification information and the status of the first process, or the identification information of the first process and the status of the first process, and the status includes: running, canceling, canceled, paused, or ended.
[0020] In a possible implementation, the first information is actively reported by the second device.
[0021] In another possible implementation, the first information is reported by the second device in response to a request from the first device. The method further includes: before sending the first information, receiving a request message for requesting the status of the first process, where the request message includes the first identification information and the status of the first process, or the identification information of the first process and the status of the first process.
[0022] In the embodiments of the present application, when the model deployment process management method is batch management, the overall deployment process of the target model group can be uniquely identified by the first identification information or the identification information of the first process. When the second device reports the status of the first process, it can indicate the correspondence between the currently reported status information and the model deployment process based on the first identification information or the identification information of the first process. When a target model group is deployed and the target model group includes multiple models, the second device can also report only the status information of one process when reporting the first information, which helps to reduce the signaling overhead between the second device and the device interacting with the second device.
[0023] In combination with the first aspect, in some implementation manners of the first aspect, the method further includes: receiving a second information, where the second information is used to control the overall deployment process of the target model group, and the second information includes the first identification information and a control instruction for the first process, or the identification information of the first process and the control instruction, and the control instruction includes cancel, pause, or start.
[0024] In the embodiments of the present application, when the model deployment process management method is batch management, the overall deployment process of the target model group can be controlled by one control instruction, which helps to improve the model deployment efficiency.
[0025] In combination with the first aspect, in some implementation manners of the first aspect, the default model deployment process management method for the target model group is independent management.
[0026] In combination with the first aspect, in some implementation manners of the first aspect, the first indication information indicates that the model deployment process management method for the target model group is independent management.
[0027] "Independent management" can be understood as that for each model included in the target model group, an independent process is created for deployment process management.
[0028] In combination with the first aspect, in some implementation manners of the first aspect, the first process includes multiple processes, the multiple processes correspond one-to-one to the identification information of each model in the multiple models, the multiple processes respectively correspond to the deployment processes of the multiple models, and the identification information of the first process includes the multiple identification information corresponding to the multiple processes.
[0029] In combination with the first aspect, in some implementations of the first aspect, the method further includes: sending third information, where the third information includes the first identification information, identification information of each model corresponding to all or part of the multiple processes, and the status of each of the all or part of the processes; or, the identification information of the all or part of the processes, and the status of each of the all or part of the processes, where the status includes: running, canceling, canceled, paused, or ended.
[0030] In a possible implementation, the third information is actively reported by the second device.
[0031] In another possible implementation, the third information is reported by the second device in response to a request from the first device. The method further includes: before sending the third information, receiving a message for requesting the status of all or part of the multiple processes, where the message includes the first identification information, identification information of each model corresponding to all or part of the multiple processes; or, the identification information of the all or part of the processes.
[0032] In combination with the first aspect, in some implementations of the first aspect, the method further includes: receiving fourth information, where the fourth information is used to control the deployment processes of all or part of the multiple models, and the fourth information includes the first identification information, identification information of each of the all or part of the models, and control instructions for the deployment processes of each of the all or part of the models; or, the identification information of the deployment processes of each of the all or part of the models, and the control instructions.
[0033] In the embodiments of the present application, through an independently managed model deployment method, the first process may include multiple processes, and the multiple processes correspond one-to-one to the identification information of each of the multiple models, and the multiple processes respectively correspond to the deployment processes of the multiple models. In this way, the first device can respectively obtain and control the deployment statuses of the respective models in the target model group, and can reflect the different deployment statuses of the respective models. In the case where some models in the target model group fail to be deployed, it can be determined which models have failed to be deployed. Based on the identification information of the partially failed models, these models can be redeployed without redeploying the entire target model group, which is beneficial to improving management accuracy and reducing unnecessary resource waste.
[0034] In combination with the first aspect, in some implementations of the first aspect, the identification information of the target model group is sent through a request message or a configuration message.
[0035] When the identification information of the target model group is sent through a request message, the request message can be used to indicate the deployment of the target model group; when the identification information of the target model group is sent through a configuration message, the configuration message may further include the deployment policy of the target model group.
[0036] Optionally, the deployment policy of the target model group may be a requirement for the inference ability of the target model group, or parameters of each functional entity involved in the model deployment process, etc. The deployment policy may be determined by the first device or reported by the entity responsible for the inference function, and this application does not make any limitations in this regard. Exemplarily, the deployment policy may be that the accuracy of the target model group is greater than 90%, or the time reaches a preset time point or preset period for deploying the target model group, or it is detected that a certain network performance indicator is lower than a preset threshold (such as: energy efficiency is less than 500 bits per joule), etc. This application does not make specific limitations in this regard.
[0037] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: receiving the identification information of each model in the target model group.
[0038] In a possible implementation manner, the identification information of each model in the target model group is sent through a request message or a configuration message.
[0039] In another possible implementation manner, among the various functional entities participating in the deployment of the target model group, a mapping relationship between the identification information of the target model group and the identification information of each model in the target model group is stored. In this way, after any party obtains the identification information of the target model group, it can search locally based on the identification information of the target model group or request to obtain the identification information of each model in the target model group from other entities.
[0040] In a second aspect, this application further provides a model deployment method. This method can be executed by a first device or a component in the first device (such as a chip, a chip system, etc.). The first device has the ability to call the model loading service. The first device may be a network management system, a network element management system, a network device, or a radio access network device, or any other device or system with the ability to call the model loading service. This application does not make specific limitations in this regard. The method includes: determining the identification information of a target model group, where the target model group includes multiple models that have been jointly trained and / or jointly tested; sending the identification information of the target model group.
[0041] Optionally, the manner in which the first device determines the identification information of the target model group may include: the first device determines the identification information of the target model group corresponding to the requirement based on the requirement reported by the target inference functional entity; or, the first device determines the identification information of the target model group corresponding to the target inference function based on the operating conditions of the target inference functional entity monitored, etc. This application does not make any limitations in this regard.
[0042] Optionally, the identification information of the target model group may be stored locally by the first device or obtained by the first device from other devices. This application also does not make any limitations in this regard.
[0043] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: receiving the first identification information and the identification information of the first process, where the first identification information is associated with a message carrying the identification information of the target model group, and there is a mapping relationship between the first identification information and the first process, and the first process is used to manage the deployment process of the target model group.
[0044] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: sending first indication information, where the first indication information is used to indicate the manner of managing the deployment process of the target model group.
[0045] In combination with the second aspect, in some implementation manners of the second aspect, the manner of managing the deployment process is batch management.
[0046] In combination with the second aspect, in some implementation manners of the second aspect, the manner of managing the deployment process of the target model group is default to batch management.
[0047] In combination with the second aspect, in some implementation manners of the second aspect, the first process corresponds to the overall deployment process of the target model group.
[0048] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: receiving first information, where the first information includes the first identification information and the status of the first process, or the identification information of the first process and the status of the first process, and the status includes: running, canceling, canceled, paused, or ended.
[0049] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: sending second information, where the second information is used to control the overall deployment process of the target model group, and the second information includes the first identification information and a control instruction for the first process, or the identification information of the first process and the control instruction, and the control instruction includes cancel, pause, or start.
[0050] In combination with the second aspect, in some implementation manners of the second aspect, the deployment process management manner of the target model group is default to independent management.
[0051] In combination with the second aspect, in some implementation manners of the second aspect, the deployment process management manner is independent management.
[0052] In combination with the second aspect, in some implementation manners of the second aspect, the first process includes multiple processes, the multiple processes correspond one-to-one to the identification information of each model in the multiple models, the multiple processes respectively correspond to the deployment processes of the multiple models, and the identification information of the first process includes the multiple identification information corresponding to the multiple processes.
[0053] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: receiving third information, where the third information includes the first identification information, the identification information of each model corresponding to all or part of the multiple processes, and the states of the all or part of the processes respectively; or, the identification information of the all or part of the processes, and the states of the all or part of the processes respectively, where the states include: running, canceling, canceled, paused or ended.
[0054] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: sending fourth information, where the fourth information is used to control the deployment processes of all or part of the multiple models, and the fourth information includes the first identification information, the identification information of each model in the all or part of the models, and control instructions for the deployment processes of each model in the all or part of the models; or, the identification information of the deployment processes of each model in the all or part of the models, and the control instructions.
[0055] In combination with the second aspect, in some implementation manners of the second aspect, the identification information of the target model group is sent through a request message or a configuration message.
[0056] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: sending the identification information of each model in the target model group.
[0057] In a third aspect, the present application provides a communication device, including a module for executing the method according to any one of the first aspect or the second aspect.
[0058] In a fourth aspect, the present application further provides a communication device, including a processor, which is coupled to a memory and can be used to execute instructions in the memory to implement the method in any possible implementation manner of the above first aspect or the second aspect. Optionally, the device further includes a memory. Optionally, the device further includes a communication interface, and the processor is coupled to the communication interface.
[0059] In a fifth aspect, a processor is provided, including: an input circuit, an output circuit, and a processing circuit. The processing circuit is configured to receive a signal through the input circuit and transmit the signal through the output circuit, so that the processor executes the method in any one of the possible implementation manners in the first aspect or the second aspect described above.
[0060] In a specific implementation process, the above-mentioned processor may be a chip, the input circuit may be an input pin, the output circuit may be an output pin, and the processing circuit may be transistors, gate circuits, flip-flops, and various logic circuits, etc. The input signal received by the input circuit may be received and input by, for example but not limited to, a receiver. The signal output by the output circuit may be output to, for example but not limited to, a transmitter and transmitted by the transmitter. Moreover, the input circuit and the output circuit may be the same circuit, which serves as the input circuit and the output circuit at different times respectively. The embodiments of the present application do not limit the specific implementation manners of the processor and various circuits.
[0061] In a sixth aspect, a processing device is provided, including a processor and a memory. The processor is configured to read instructions stored in the memory, and may receive a signal through a receiver and transmit a signal through a transmitter to execute the method in any one of the possible implementation manners in the first aspect or the second aspect described above.
[0062] Optionally, there is one or more processors, and one or more memories.
[0063] Optionally, the memory may be integrated with the processor, or the memory and the processor are separately arranged.
[0064] In a specific implementation process, the memory may be a non-transitory memory, such as a read only memory (ROM), which may be integrated with the processor on the same chip or may be separately arranged on different chips. The embodiments of the present application do not limit the type of the memory and the arrangement manner of the memory and the processor.
[0065] It should be understood that relevant data interaction processes, such as sending indication information, may be a process of outputting indication information from the processor, and receiving capability information may be a process of the processor receiving input capability information. Specifically, the data output by the processing may be output to the transmitter, and the input data received by the processor may come from the receiver. Among them, the transmitter and the receiver may be collectively referred to as a transceiver.
[0066] The processing device in the sixth aspect described above may be a chip, and the processor may be implemented by hardware or software. When implemented by hardware, the processor may be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor may be a general-purpose processor that is implemented by reading software code stored in a memory. The memory may be integrated in the processor or may be located outside the processor and exist independently.
[0067] In a seventh aspect, a computer program product is provided. The computer program product includes: a computer program (which may also be referred to as code or instructions). When the computer program is run, it causes a computer to execute the method in any one of the possible implementation manners in the first aspect or the second aspect described above.
[0068] In an eighth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program (which may also be referred to as code or instructions). When it runs on a computer, it causes the computer to execute the method in any one of the possible implementation manners in the first aspect or the second aspect described above.
[0069] The beneficial effects and possible implementation manners of the third aspect to the eighth aspect may refer to the descriptions of the first aspect and the second aspect, and will not be elaborated here. Description of the Drawings
[0070] Figure 1 A model workflow provided by an embodiment of the present application;
[0071] Figure 2 A network resource model working diagram provided by an embodiment of the present application;
[0072] Figure 3 A network resource model working diagram provided by an embodiment of the application;
[0073] Figure 4 Another network resource model working diagram provided by an embodiment of the application;
[0074] Figure 5 A schematic diagram of a communication system applicable to an embodiment of the application;
[0075] Figure 6 A schematic flowchart of a model deployment method provided by an embodiment of the present application;
[0076] Figure 7 Another schematic flowchart of a model deployment method provided by an embodiment of the present application;
[0077] Figure 8 Another schematic flowchart of a model deployment method provided by an embodiment of the present application;
[0078] Figure 9A schematic flowchart of another model deployment method provided by an embodiment of the present application;
[0079] Figure 10 A schematic flowchart of a model deployment method provided by an embodiment of the present application;
[0080] Figure 11 A schematic block diagram of a communication device provided by an embodiment of the present application;
[0081] Figure 12 A schematic block diagram of another communication device provided by an embodiment of the present application. Detailed implementation manners
[0082] Next, the technical solutions in the present application will be described with reference to the accompanying drawings.
[0083] Before introducing the present application, the following points are explained first.
[0084] First, in the embodiments shown below, each term and English abbreviation, such as reference data or differential data, etc., are all exemplary examples given for convenience of description, and should not constitute any limitation to the present application. The present application does not exclude the possibility of defining other terms that can achieve the same or similar functions in existing or future protocols.
[0085] Second, in the embodiments shown below, the first, second, and various numerical numbers are only for the convenience of description and are not used to limit the scope of the embodiments of the present application.
[0086] Third, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or its similar expression means any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, and c can represent: a, or b, or c, or a and b, or a and c, or b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0087] To improve the intelligence and automation level of the network, artificial intelligence (AI) and machine learning (ML) technologies are also gradually being applied to promote network intelligence.
[0088] Figure 1An exemplary AI / ML workflow is shown, mainly including a training phase, an emulation phase, a deployment phase, and an inference phase. It should be understood that generally, the AI / ML workflow is implemented sequentially according to the training phase, emulation phase, deployment phase, and inference phase. However, in some possible implementation manners, the AI / ML model can execute the inference phase after the training phase and / or emulation phase. Or, after the training phase, the emulation phase is skipped and the deployment phase and inference phase are executed. Or, after the emulation phase or inference phase, the model enters the training phase again. Among them, the deployment phase is the process of loading the model into the inference function. In one possible implementation manner, after the model goes through the training phase and emulation phase, it is loaded into the corresponding inference function through the deployment phase to enter the inference phase and execute the inference process.
[0089] In some examples, the role of providing the ML entity loading management service (MnS) is called the producer, and the role of invoking the ML entity loading management service is called the consumer. Two sub-use cases are defined for the deployment phase and solutions for these two sub-use cases.
[0090] The two sub-use cases include:
[0091] Sub-use case 1: The consumer requests the producer to perform model loading (consumer requested ML entity loading).
[0092] Sub-use case 2: The consumer configures a policy for the producer, and the producer triggers the model loading by itself (control of producer-initiated ML entity loading).
[0093] Solutions for the two sub-use cases:
[0094] The current solution defines 3 information object classes (IOCs), including: MLEntityLoadingRequest, MLEntityLoadingPolicy, and MLEntityLoadingProcess which is used to represent the ML model loading process. Among them, MLEntityLoadingRequest and MLEntityLoadingProcess are applied to the scenario of the above-mentioned sub-use case 1, and MLEntityLoadingPolicy and MLEntityLoadingProcess are applied to the scenario of the above-mentioned sub-use case 2.
[0095] The attributes of the above 3 information object classes are described in detail below.
[0096] 1. MLEntityLoadingRequest< <ioc>>
[0097] MLEntityLoadingRequest< <ioc>>Represents an ML model loading request created by a consumer. The consumer uses this IOC request to ask the producer to load the ML model into the target inference function. The attributes of this IOC are shown in Table 1.
[0098] Table 1
[0099]
[0100] Among them, the requestStatus attribute represents the status of the request. The value of the requestStatus attribute can be any of the following: Not Started (NOT_STARTED), Loading in Progress (LOADING_IN_PROGRES), Suspended (SUSPENDED), Finished Successfully (FINISHED_SUCCESS), Finished with Failure (FINISHED_FAILED), or Cancelled (CANCELLED). The cancelRequest attribute represents whether the consumer cancels the ML entity loading request, with values of Yes (TRUE) or No (FALSE). The role-related attribute mLEntityToLoadRef represents the identifier of the model to be loaded for the ML model loading request created by the consumer.
[0101] For the values of "Support Qualifier", "Readable", and "Writable" in the above table, they can be understood as follows: The values of "Support Qualifier" include Mandatory (M), Optional (O), and Conditional Mandatory (CM). A value of M indicates that the attribute is a mandatory attribute, a value of O indicates that the attribute is an optional attribute, and a value of CM indicates conditional mandatory, that is, the attribute is mandatory when a certain condition is met; The values of "Readable" include True (TRUE, T) and False (FALSE, F). When the value is T, it means the attribute is readable, and when the value is F, it means the attribute is not readable; The values of "Writable" also include TRUE and FALSE. When the value is T, it means the attribute is writable, and when the value is F, it means the attribute is not writable. This explanation also applies to the following tables and will not be repeated later.
[0102] 2. MLEntityLoadingPolicy< <ioc>>
[0103] MLEntityLoadingPolicy< <ioc>>Indicates the ML model loading strategy set by the consumer for the producer. The consumer uses this IOC to set the conditions for the producer to trigger the ML model loading. Only when the ML model meets the conditions can the producer trigger the model loading. The attributes of this IOC are shown in Table 2.
[0104] Table 2
[0105]
[0106] Among them, the inferenceType attribute is used to indicate the inference type of the initial model associated with the ML model loading strategy set by the consumer for the producer. This attribute is a conditionally required attribute, that is, when the model associated with the ML model loading strategy set by the consumer for the producer is an initial training model, this attribute is a required attribute. It should be understood that the initial training model is the model trained for the first time, and the initial training model has not been assigned a model identifier. Therefore, the inference type of the initial training model can be used to identify the model.
[0107] The mLEntityId attribute is used to indicate the identifier of the retrained model corresponding to the ML model loading strategy set by the consumer for the producer. This attribute is a conditionally required attribute, that is, when the model associated with the ML model loading strategy set by the consumer for the producer is a retrained model, this attribute is a required attribute. It should be understood that an mLEntityId can uniquely identify a model that has undergone initial training.
[0108] The policyForLoading attribute is used to indicate the ML model loading strategy set by the consumer for the producer. This strategy can be a series of thresholds. For example, this strategy can be that the accuracy of the ML model is greater than 90%, or it reaches a preset time point or preset period, or a certain network performance index is lower than a preset threshold (such as: energy efficiency is less than 500 bits per joule).
[0109] 3. MLEntityLoadingProcess< <ioc>>
[0110] MLEntityLoadingProcess< <ioc>>Represents the ML model loading process. In the scenario of the above Sub-use Case 1, the producer can use this IOC to instantiate one or more ML model loading processes for each ML model loading request of the consumer; in the scenario of the above Sub-use Case 2, the producer can use this IOC to instantiate one or more ML model loading processes and associate these one or more ML model loading processes with the ML model loading policy set by the consumer for the producer. The attributes of this IOC are shown in Table 3.
[0111] Table 3
[0112]
[0113] Among them, the progressStatus attribute represents the status of the ML model loading process, and the value of this attribute can be any one of the following: Running, Cancelling, Suspended, Finished, or Cancelled. The cancelProcess attribute represents whether the consumer cancels the ML entity loading process, and the value is true or false. The suspendProcess attribute represents whether the consumer suspends the ML entity loading process, and the value is TRUE or FALSE. The resumeProcess attribute represents whether the consumer resumes the ML entity loading process, and the value is TRUE or FALSE.
[0114] The role-related attribute MLEntityLoadingRequestRef is the loading request identifier associated with the ML entity loading process, and it is a mandatory attribute when the ML entity loading process corresponds to the scenario of the above Sub-use Case 1. MLEntityLoadingPolicyRef is the loading policy identifier associated with the ML entity loading process, and it is a mandatory attribute when the ML entity loading process corresponds to the scenario of the above Sub-use Case 2. LoadedMLEntityRef is the model identifier associated with the ML entity loading process.
[0115] Figure 2 Exemplarily shows the network resource model (NRM) relationship diagrams of the respective IOCs corresponding to the above Table 1, Table 2, and Table 3. As Figure 2 As shown, the AiMlInferenceFunction represents a function that can perform the ML inference function and is also an IOC. In the figure, the connections between the MLEntityLoadingRequest object, the MLEntityLoadingProcess object, the MLEntityLoadingPolicy object, and the AiMlInferenceFunction object are solid diamonds on the side of the AiMlInferenceFunction object. This solid diamond is used to represent the inclusion relationship, that is, the AiMlInferenceFunction object includes the MLEntityLoadingRequest object, the MLEntityLoadingProcess object, and the MLEntityLoadingPolicy object. Further, the connections between the MLEntityLoadingRequest object and the MLEntityLoadingProcess object and the MLEntity object are arrows on the side of the MLEntity object. This arrow is used to represent the association relationship, that is, the MLEntityLoadingRequest object and the MLEntityLoadingProcess object are respectively associated with the MLEntity object. Therefore, it can be further understood that the AiMlInferenceFunction object includes the MLEntity object, which is represented in the figure as a solid diamond on the side of the AiMlInferenceFunction object for the connection between the AiMlInferenceFunction object and the MLEntity object.
[0116] In addition, the numbers and symbols on the connections between the various objects in the figure represent the cardinality relationships between the intended object attributes with inclusion and association relationships. For example Figure 2 On the connection line between the MLEntityLoadingRequest object and the AiMlInferenceFunction object shown in [figure], the cardinality on the MLEntityLoadingRequest object side is "*", and the cardinality on the AiMlInferenceFunction object side is "1". On the connection line between the MLEntityLoadingRequest object and the MLEntity object, the cardinality on both sides of the MLEntityLoadingRequest object and the MLEntity object is "1", which means that an AiMlInferenceFunction object can contain multiple MLEntityLoadingRequest objects, and each MLEntityLoadingRequest object can only be associated with one MLEntity object. Figure 2 The connection rules of other connection lines in [figure] and the subsequent NRM diagrams can also be explained according to this rule, so no more details will be given.
[0117] That is to say, when loading a model for an inference function, only one model can be loaded through a single model loading request initiated by the consumer in the above sub-use case 1. Generally speaking, each model can independently support a specific type of function. However, in actual applications, the implementation of an inference function may require the cooperation of multiple models. For example, the output of one model is used as the input of another model to form a sequence of interconnected models, or multiple models provide outputs in parallel, or the parallel outputs of multiple models are merged and used as the input of another model. Through such a combination of multiple models, the function can be realized. These multiple models can be jointly trained and / or jointly tested models.
[0118] According to Figure 2 the design, although multiple models can be loaded through multiple model loading requests, on the one hand, the models loaded by multiple model loading requests may be independently trained or independently tested, and the purpose of joint inference cannot be achieved. On the other hand, even if these multiple models are jointly trained and jointly tested, since the multiple models are deployed in their own independent ways based on their respective model identifiers, only the description files of each model itself can be read, and the joint method between these multiple models cannot be obtained. That is to say, multiple models can each implement the inference functions they support, but the joint inference function cannot be achieved.
[0119] In view of this, the present application provides a model deployment method and a communication device, which can load a set of target model groups for a target inference function. The target model groups include multiple models jointly trained and / or jointly tested. During the deployment process, a description file about the target model groups can be read based on the identification information of the target model groups. The description file can include the cooperation mode of the multiple models in the target model groups when jointly implementing the inference function. In this way, the multiple models of the deployed target model groups can implement joint inference based on the description file obtained during the deployment process, thus achieving the purpose of deploying a target model group that can support joint inference for the target inference function.
[0120] Next, first, a detailed description of the IOC design proposed in the present application will be given. There are the following two design methods.
[0121] Design 1:
[0122] 1. Add a new object that can model the joint inference model group.
[0123] Exemplarily, the object can be MLEntityCoordinationGroup< <ioc>>, but this application does not limit this. The MLEntityCoordinationGroup< <ioc>> can represent a target model group, which may include multiple models that can be jointly trained and / or jointly tested to achieve joint inference.
[0124] Optionally, the model can be an AI model or an ML model (also referred to as an ML entity), and this application does not limit it. In the subsequent description of this application, only the model being an ML model is taken as an example, but this example does not limit this application.
[0125] It should be noted that one MLEntityCoordinationGroup is associated with at least two ML models.
[0126] In one possible implementation, MLEntityCoordinationGroup < <ioc>The attributes of > are shown in Table 4.
[0127] Table 4
[0128]
[0129] Optionally, the role-related attribute memberMLEntityRefList in Table 4 above can represent a list of model identifiers in a target model group, but this application does not make specific limitations in this regard. The model identifiers in the target model group can also be reflected in other forms.
[0130] 2. Add the following two attributes to the model loading request object created for consumers:
[0131] Attribute 1: Used to indicate the management method during the model deployment process (such as batch management or independent management);
[0132] Attribute 2: Used to indicate that the model loading request created by the consumer is used to request the deployment of a target model group, and multiple models in this target model group can achieve joint inference.
[0133] Exemplarily, the model loading request created by the consumer can be through MLEntityLoadingRequest< <ioc>>Indicates that Table 5 exemplarily shows an MLEntityLoadingRequest provided by an embodiment of the present application< <ioc>>Contained attributes.
[0134] Table 5
[0135]
[0136] In a possible implementation, the above attribute 1 can be groupBatchIndicator, which can be an optional attribute and can indicate the management method during the model deployment process. For example, batch management can be performed on the target model group, or independent management can be performed on multiple models in the target model group; the above attribute 2 can be the role-related attribute mLEntityCoordinationGroupToLoadRef, which represents the identifier of the target model group that the ML model loading request created by the consumer needs to load. At the same time, modify the support qualifier of mLEntityToLoadRef to CM. In this way, when the model that the ML model loading request created by the consumer needs to load is not the target model group, mLEntityToLoadRef is a required attribute, and when the model that the ML model loading request created by the consumer needs to load is the target model group, mLEntityCoordinationGroupToLoadRef is a required attribute. However, it should be understood that the specific manifestation forms of attribute 1 and attribute 2 in this application are not specifically limited.
[0137] 3. Add the following two attributes to the model loading policy object set by the consumer for the producer:
[0138] Attribute 3: Used to indicate the management method during the model deployment process (such as batch management or independent management);
[0139] Attribute 4: Used to indicate that the model associated with the model loading policy set by the consumer for the producer is the target model group, and multiple models in this target model group can achieve joint inference.
[0140] Exemplarily, the model loading policy set by the consumer for the producer can be through MLEntityLoadingPolicy< <ioc>>Indicates that Table 6 exemplarily shows an MLEntityLoadingPolicy provided by an embodiment of the present application< <ioc>>Included attributes.
[0141] Table 6
[0142]
[0143] In a possible implementation, the above-mentioned attribute 3 can be groupBatchIndicator, which can be an optional attribute and can indicate the management method during the model deployment process through groupBatchIndicator; the above-mentioned attribute 4 can be the role-related attribute mLEntityCoordinationGroupToLoadRef, which represents the identifier of the target model group associated with the ML model loading policy set by the consumer for the producer. This attribute can be a conditionally required attribute, that is, when the model associated with the ML model loading policy set by the consumer for the producer is the target model group, mLEntityCoordinationGroupToLoadRef is a required attribute. However, it should be understood that the specific forms of attribute 3 and attribute 4 in this application are not specifically limited.
[0144] Optionally, the value limit of the identifier of the retrained model corresponding to the model loading policy set by the consumer for the producer can also be modified. In a possible implementation, the identifier of the retrained model corresponding to the model loading policy set by the consumer for the producer can be represented by mLEntityId, as Figure 2 shown in, the value limit of mLEntityId is 1, and one MLEntityLoadingPolicy object can only be associated with one MLEntity object. In a possible implementation of this design, the value limit of mLEntityId is cancelled, so that one MLEntityLoadingPolicy object can be associated with one or more MLEntity objects.
[0145] 4. Add the following two attributes to the model loading process object:
[0146] Attribute 5: Used to associate with the deployment process of the target model group;
[0147] Attribute 6: Used to associate with the deployment processes of multiple models in the target model group during the deployment process of the target model group.
[0148] Optionally, the model loading process can also be referred to as the model deployment process, and this application does not limit this.
[0149] Exemplarily, the model loading process object can be through MLEntityLoadingProcess< <ioc>>Indicates that Table 7 exemplarily shows an MLEntityLoadingProcess provided by an embodiment of the present application< <ioc>>Contained attributes.
[0150] Table 7
[0151]
[0152]
[0153] In a possible implementation, the above-mentioned attribute 5 may be LoadedMLEntityCoordinationGroupRef, which is a conditionally required attribute. For example, when the required loaded model corresponding to the above-mentioned sub-use case 1 or sub-use case 2 is the target model group, LoadedMLEntityCoordinationGroupRef is a required attribute; the above-mentioned attribute 6 may be LoadedMemberMLEntityRef, which may also be a conditionally required attribute. For example, when the required loaded model corresponding to the above-mentioned sub-use case 1 or sub-use case 2 is the target model group and the management method during the model deployment process is to independently manage multiple models in the target model group, this attribute is a required attribute. However, it should be understood that the present application does not limit the specific forms of attribute 5 and attribute 6.
[0154] Figure 3 Exemplarily shows the network resource model relationship diagram of each IOC corresponding to Design One provided by the embodiments of the present application. Figure 3 Compared with Figure 2 a new MLEntityCoordinationGroup object is added, and the cardinality relationships between the MLEntityLoadingRequest object, the MLEntityLoadingProcess object, and the MLEntityLoadingPolicy object and the MLEntityCoordinationGroup object are defined respectively.
[0155] Among them, on the arrow line where the MLEntityLoadingRequest object points to the MLEntityCoordinationGroup object and the MLEntity object, the cardinality is "1" on the side close to the MLEntityLoadingRequest object, and the cardinality is "0..1" on the side close to the MLEntityCoordinationGroup object and the MLEntity object, indicating that one MLEntityLoadingRequest object can be associated with one MLEntityCoordinationGroup object, or one MLEntityLoadingRequest object can be associated with one MLEntity object.
[0156] On the arrow line where the MLEntityLoadingPolicy object points to the MLEntityCoordinationGroup object and the MLEntity object, the cardinality is "1" on the side close to the MLEntityLoadingPolicy object, and the cardinality is "0..1" on the side close to the MLEntityCoordinationGroup object and the MLEntity object, indicating that one MLEntityLoadingPolicy object can be associated with one MLEntityCoordinationGroup object, or one MLEntityLoadingPolicy object can be associated with one MLEntity object.
[0157] On the arrow line where the MLEntityLoadingProcess object points to the MLEntityCoordinationGroup object and the MLEntity object, the cardinality is "0..1" on the side close to the MLEntityCoordinationGroup object and the MLEntity object, indicating that the MLEntityLoadingProcess object can be associated with one MLEntityCoordinationGroup object, or the MLEntityLoadingProcess object can be associated with one MLEntity object. On the arrow line where the MLEntityLoadingProcess object points to the MLEntityCoordinationGroup object, the cardinality is "1..*" on the side close to the MLEntityLoadingProcess object, indicating that when the MLEntityLoadingProcess object is associated with the MLEntityCoordinationGroup object, the number of MLEntityLoadingProcess objects can be at least one ( "at least one" can also be described as "one or more"); on the arrow line where the MLEntityLoadingProcess object points to the MLEntity object, the cardinality is "1" on the side close to the MLEntityLoadingProcess object, indicating that when the MLEntityLoadingProcess object is associated with the MLEntity object, the number of MLEntityLoadingProcess objects is one.
[0158] In addition, on the arrow line pointing from the MLEntityLoadingProcess object to the MLEntityLoadingPolicy object, and on the arrow line pointing from the MLEntityLoadingProcess object to the MLEntityLoadingRequest object, the cardinality on the side close to the MLEntityLoadingProcess object is changed from "*" ( Figure 2 as shown in) to "1..*", indicating that the number of MLEntityLoadingProcess objects associated with one MLEntityLoadingPolicy object can be at least one, and the number of MLEntityLoadingProcess objects associated with one MLEntityLoadingRequest object can also be at least one.
[0159] In a possible implementation, based on the above Design One, when deploying a target model group for an inference function, it can be achieved by associating an MLEntityCoordinationGroup object with an MLEntityLoadingRequest object, or it can be achieved by associating an MLEntityCoordinationGroup object with an MLEntityLoadingPolicy object.
[0160] Design Two:
[0161] 1. Modification to the model loading request object created by the consumer:
[0162] 1) Modify the value of the identifier of the model to be loaded in the model loading request created by the consumer;
[0163] 2) Add Attribute 7: Used to indicate that the model loading request created by the consumer is for requesting the deployment of a target model group, and multiple models in this target model group can perform joint inference;
[0164] 3) Add Attribute 8: Used to indicate the management method during the model deployment process (such as batch management or independent management).
[0165] Exemplarily, the model loading request created by the consumer can be through MLEntityLoadingRequest< <ioc>>Indicates that Table 8 exemplarily shows another MLEntityLoadingRequest provided by the embodiments of the present application< <ioc>>Contained attributes.
[0166] Table 8
[0167]
[0168] In a possible implementation, the identifier of the model to be loaded by the model loading request created by the consumer can be represented by mLEntityToLoadRef. Figure 2 As shown in, the value range of mLEntityToLoadRef is restricted to 1, and one MLEntityLoadingRequest object can only be associated with one MLEntity object. In a possible implementation of this design, the value range restriction of mLEntityToLoadRef is removed, so that one MLEntityLoadingRequest object can be associated with one or more MLEntity objects. In addition, a new attribute 7 is added to indicate that the mLEntityToLoadRef corresponding to one or more MLEntity objects associated with one MLEntityLoadingRequest object belongs to a target model group that can implement the joint inference function. In this way, when the model to be loaded by the ML model loading request created by the consumer is the target model group, the above attribute 7 is a required attribute. Optionally, the above attribute 7 can be represented by mLEntityGroupInfo, but this application does not make any restrictions on this.
[0169] In a possible implementation, the above attribute 8 can be groupBatchIndicator, and this attribute can be an optional attribute. The management method during the model deployment process can be indicated by groupBatchIndicator. Exemplarily, this management method can be batch management for the target model group, or independent management for multiple models in the target model group, but this application does not make specific restrictions on this.
[0170] 2. Modifications to the model loading policy object set by the consumer for the producer:
[0171] 1) Modify the value range of the identifier of the retraining model corresponding to the model loading policy set by the consumer for the producer;
[0172] 2) Add a new attribute 9: used to indicate that the model loading request created by the consumer is for requesting the deployment of a target model group, and multiple models in this target model group can implement joint inference;
[0173] 3) Add a new attribute 10: used to indicate the management method during the model deployment process (such as batch management or independent management).
[0174] Exemplarily, the model loading policy set by the consumer for the producer can be through MLEntityLoadingPolicy< <ioc>>Indicates that Table 9 exemplarily shows another MLEntityLoadingPolicy provided by the embodiments of the present application< <ioc>>Contained attributes.
[0175] Table 9
[0176]
[0177] In a possible implementation, the identifier of the retrained model corresponding to the model loading policy set by the consumer for the producer can be represented by mLEntityId. As shown in Figure 2 , the value range of mLEntityId is restricted to 1, and an MLEntityLoadingPolicy object can only be associated with one MLEntity object. In a possible implementation of this design, the value range restriction of mLEntityId is removed, that is, an MLEntityLoadingPolicy object can be associated with one or more MLEntity objects. In addition, a new attribute 9 is added to indicate that the mLEntityIds corresponding to one or more MLEntity objects associated with an MLEntityLoadingPolicy object belong to a target model group that can implement the joint inference function. In this way, when the model associated with the ML model loading policy set by the consumer for the producer is the target model group, the above-mentioned attribute 9 is a required attribute. Optionally, the above-mentioned attribute 9 can be represented by mLEntityGroupInfo, but this application does not make any restrictions on this.
[0178] In a possible implementation, the above-mentioned attribute 10 can be groupBatchIndicator, and this attribute can be an optional attribute. The management method during the model deployment process can be indicated by groupBatchIndicator. Exemplarily, this management method can be batch management for the target model group, or independent management for multiple models in the target model group, but this application does not make specific restrictions on this.
[0179] 3. Modification for the model loading process object: Modify the value of the model identifier associated with the loading process.
[0180] Exemplarily, the model loading process object can be passed through MLEntityLoadingProcess< <ioc>>indicates that Table 10 exemplarily shows an MLEntityLoadingProcess provided by an embodiment of the present application< <ioc>>Contained attributes.
[0181] Table 10
[0182]
[0183] In a possible implementation, the model identifier associated with the loading process can be represented by LoadedMLEntityRef. In some examples, the value of LoadedMLEntityRef is restricted to 1, that is, one MLEntityLoadingProcess object corresponds to only one LoadedMLEntityRef. In a possible implementation of this design, the value restriction of LoadedMLEntityRef is removed, so that one MLEntityLoadingProcess object can correspond to one or more LoadedMLEntityRefs. Exemplarily, when the deployment process management method is batch management, one MLEntityLoadingProcess object can correspond to multiple LoadedMLEntityRefs, and at this time, the multiple LoadedMLEntityRefs respectively correspond to the identification information of multiple models in the target model group; when the deployment process management method is independent management, one MLEntityLoadingProcess object can correspond to one LoadedMLEntityRef, and at this time, the LoadedMLEntityRef corresponds to the identification information of one model in the target model group, but this application does not make any limitations in this regard. Optionally. The models in the target model group can also be referred to as sub-models of the target model group, and this application does not make any limitations in this regard.
[0184] Figure 4 Exemplarily shows the network resource model relationship diagram of each IOC corresponding to Design 2 provided by the embodiments of the present application. As Figure 4 As shown, on the arrow line where the MLEntityLoadingRequest object points to the MLEntity object, the cardinality is "1" on the side close to the MLEntityLoadingRequest object and "1..*" on the side close to the MLEntity object, indicating that one MLEntityLoadingRequest object can be associated with at least one MLEntity object; on the arrow line where the MLEntityLoadingPolicy object points to the MLEntity object, the cardinality is "1" on the side close to the MLEntityLoadingPolicy object and "1..*" on the side close to the MLEntity object, indicating that one MLEntityLoadingPolicy object can be associated with at least one MLEntity object; on the arrow line where the MLEntityLoadingProcess object points to the MLEntity object, the cardinality is "1" on the side close to the MLEntityLoadingProcess object and "1..*" on the side close to the MLEntity object, indicating that one MLEntityLoadingProcess object can be associated with at least one MLEntity object.
[0185] In addition, on the arrow line where the MLEntityLoadingProcess object points to the MLEntityLoadingPolicy object and on the arrow line where the MLEntityLoadingProcess object points to the MLEntityLoadingRequest object, the cardinality on the side close to the MLEntityLoadingProcess object is modified from "*" ( Figure 2 shown in) to "1..*", indicating that the number of MLEntityLoadingProcess objects associated with one MLEntityLoadingPolicy object can be at least one, and the number of MLEntityLoadingProcess objects associated with one MLEntityLoadingRequest object can also be at least one.
[0186] In a possible implementation, based on the above Design 2, when deploying a target model group for an inference function, it can be achieved by associating at least one MLEntity object through an MLEntityLoadingRequest object, or it can be achieved by associating at least one MLEntity object through an MLEntityLoadingPolicy object. When deploying the target model group, the number of at least one MLEntity object can be greater than or equal to two.
[0187] Figure 5 Schematic diagram of a communication system 500 applicable to the embodiments of the present application. As Figure 5 shown, the communication system 500 includes a first device 510, a second device 520, and a third device 530. Communication can be established between the first device 510, the second device 520, and the third device 530.
[0188] In a possible implementation, the role of the first device 510 can be a model loading consumer (MLloading consumer) that can call the model loading service. The role of the second device 520 can be a model loading producer (ML loading producer) that can provide the model loading service. The third device 530 can be a role that implements the inference function based on the model (it can also be understood as the bearer of the inference function, the executor of the inference function, the acquirer of the inference result. In some examples, this role can be directly referred to as the inference function), but the present application does not make specific limitations in this regard.
[0189] It should be understood that Figure 5 in the shown communication system, the number of the first device, the second device, or the third device can be one or more, and the present application does not make limitations in this regard.
[0190] In a possible scenario, the first device 510 can call the model loading service provided by the second device 520 to instruct to load the target model into the specified third device 530, and the third device 530 is responsible for inputting data into the model to obtain the corresponding output.
[0191] In a possible implementation, the first device 510 (or the second device 520, or the third device 530) can include one or more of the following: a module that can implement the functions corresponding to the model loading consumer, a module that can implement the functions corresponding to the model loading producer, or a module that can implement the inference function based on the model. The present application does not make specific limitations in this regard.
[0192] Optionally, the first device 510 may be a network management system (NMS), an element management system (EMS), or a radio access network (RAN) device, or any other device or system that can act as a model loading consumer. This application does not make specific limitations thereon.
[0193] Optionally, the second device 520 may be a network management system, an element management system, or a radio access network device, or any other device or system that can act as a model loading producer. This application does not make specific limitations thereon.
[0194] Optionally, the third device 530 may be a network management system, an element management system, or a radio access network device, or any other device or system that can implement an inference function based on a model. This application does not make specific limitations thereon.
[0195] Optionally, the above-mentioned element management system can manage the network, be responsible for the operation, management, and function maintenance of the network, and may also be referred to as a cross-domain management system. This application does not make limitations thereon.
[0196] Optionally, the element management system can be used to manage one or more network elements of a certain category, and may also be referred to as a domain management system or a single-domain management system. This application does not make limitations thereon.
[0197] Optionally, the above access network device may be a transmission reception point (TRP), or may be an evolved NodeB (eNB or eNodeB) in a Long Term Evolution (LTE) system, or may be a home base station (for example, home evolved NodeB, or home Node B, HNB), a base band unit (BBU), or may be a radio controller in a Cloud Radio Access Network (CRAN) scenario, or the access network device may be a relay station, an access point, a vehicle-mounted device, a wearable device, and an access network device in a 5G network or an access network device in a future evolved Public Land Mobile Network (PLMN) network, etc. It may be an access point (AP) in a Wireless Local Area Network (WLAN), may be a gNB in a New Radio (NR) system, may be a satellite base station in a satellite communication system, etc. The embodiments of the present application do not limit this.
[0198] Optionally, the above access network device may include a centralized unit (CU) node, or a distributed unit (DU) node, or an access network device including a CU node and a DU node, or an access network device including a control plane CU node (CU-CP node), a user plane CU node (CU-UP node), and a DU node. Among them, the access network device including a CU node and a DU node can split the protocol layer of the access network device, place the functions of some protocol layers under centralized control of the CU, and distribute the functions of the remaining part or all protocol layers in the DU, with the CU centrally controlling the DU. As an implementation, the protocol stack deployed by the CU includes a radio resource control (RRC) layer, a packet data convergence protocol (PDCP) layer, and a service data adaptation protocol (SDAP) layer. The protocol stack deployed by the DU includes a radio link control (RLC) layer, a media access control (MAC) layer, and a physical layer (PHY) layer. Thus, the CU has the processing capabilities of RRC, PDCP, and SDAP. The DU has the processing capabilities of RLC, MAC, and PHY. The above function split is only an example and does not constitute a limitation on the CU and the DU. That is to say, there can be other ways of splitting functions between the CU and the DU, which are not elaborated in this embodiment of the present application. The functions of the CU can be implemented by one entity or by different entities. For example, the functions of the CU can be further split. For example, the control plane (CP) and the user plane (UP) can be separated, that is, the control plane (CU-CP) and the CU user plane (CU-UP) of the CU. For example, the CU-CP and the CU-UP can be implemented by different functional entities, and the CU-CP and the CU-UP can be coupled with the DU to jointly complete the functions of the access network device. In a possible way, the CU-CP is responsible for the control plane functions, mainly including RRC and PDCP-C, where PDCP-C is mainly responsible for encryption, decryption, integrity protection, data transmission, etc. of control plane data. The CU-UP is responsible for the user plane functions, mainly including SDAP and PDCP-U, where SDAP is mainly responsible for processing the data of the core network device and mapping the data flow to the bearer. PDCP-U is mainly responsible for encryption, decryption, integrity protection, header compression, sequence number maintenance, data transmission, etc. of the data plane. Among them, the CU-CP and the CU-UP are connected through the E1 interface. The CU-CP represents that the access network device is connected to the core network device through the interface between the core network device and the access network device. It is connected to the DU through F1-C (control plane).The CU-UP is connected to the DU via F1-U (user plane). In addition, there is also a possible implementation where PDCP-C is also in the CU-UP, which is not limited in this application.
[0199] To make the objectives and technical solutions of this application clearer and more intuitive, the model deployment method and communication device provided in the embodiments of this application will be described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0200] Next, first in combination with Figures 6 to 10 A model deployment method provided in the embodiments of this application will be described in detail.
[0201] The model deployment method provided in the embodiments of this application can be implemented based on the above two designs for IOC provided in this application. This method can be applied to a communication system 500 as shown in Figure 5 or other communication systems, which is not limited in this application. Exemplarily, in the embodiments of this application, taking the role of the first device as a consumer, the role of the second device as a producer, and the third device as a role that implements an inference function based on a model as an example, the model deployment method provided in this application will be described from the perspective of interaction between devices.
[0202] Figure 6 FIG. is a schematic flowchart of a model deployment method 600 provided in the embodiments of this application. The method 600 includes the following steps:
[0203] S601. The first device determines the identification information of the target model group, and the target model group includes multiple models that have been jointly trained and / or jointly tested.
[0204] It should be noted that the model in this application can be a machine learning model (ML model) or can be embodied in the form of an entity, such as a machine learning entity (ML entity), which is not limited in this application.
[0205] S602. The first device sends the identification information of the target model group. Correspondingly, the second device receives the identification information of the target model group.
[0206] S603. The second device deploys the target model group based on the identification information of the target model group.
[0207] It should be noted that the deployment of the target model group by the second device can be understood as the second device loading the target model group into one or more target inference functions (such as the third device and the gNB mentioned below), where the target model group will be used to perform inference.
[0208] In the embodiments of the present application, during the process of the second device deploying the target model group, the identity information of the target model group is used to indicate that the deployed target model group has been jointly trained and / or jointly tested and can jointly implement the inference function. Thus, a description file about the target model group can be obtained based on the identity information of the target model group during the deployment process. The description file may include the cooperation mode of multiple models in the target model group when jointly implementing the inference function. For example, the cooperation mode may be using the output of one model among the multiple models as the input of another model to form an interconnected model sequence, or the multiple models providing outputs in parallel, or the parallel outputs of some of the multiple models being merged and used as the input of another part of the models, etc. In this way, the multiple models of the deployed target model group can implement joint inference based on the description file obtained during the deployment process. It can be seen that the model deployment method provided in the present application can realize the deployment of the target model group that can jointly implement the inference function, which is beneficial for the target model group to implement joint inference.
[0209] In a possible scenario, the purpose of the first device determining the identity information of the target model group in S601 above is to deploy the target model group for the third device to implement the target inference function on the third device. Exemplarily, the manner in which the first device determines the identity information of the target model group may include: the first device determines the identity information of the target model group corresponding to the target inference function required by the third device based on the requirements reported by the third device; or, the first device determines the identity information of the target model group corresponding to the target inference function required by the third device based on the monitored operating conditions of the third device, etc. The present application does not limit this. Optionally, the identity information of the target model group may be stored locally by the first device or obtained by the first device from other devices. The present application also does not limit this.
[0210] It should be understood that the purpose of deploying the model is to make the deployed model available on the entity with the inference function (such as the third device). A possible implementation of the above S603 may be: the second device deploys the target model group for one or more third devices based on the identification information of the target model group. In some examples, the second device sends a second indication message to the third device, and the indication message is used to instruct the third device to activate the existing target model group, and the indication message may include the identification information of the target model group; in other examples, the second device sends a third indication message to the third device, and the indication message is used to instruct the third device to download and activate the target model group, and the indication message may include the identification information of the target model group and the storage address of the target model group.
[0211] As an optional embodiment, after the above S602, the method 600 further includes: the second device sends the first identification information and the identification information of the first process, the first identification information is associated with the message carrying the identification information of the target model group, and the first identification information has a mapping relationship with the first process, and the first process is used to manage the deployment process of the target model group. Correspondingly, the first device receives the first identification information and the identification information of the first process.
[0212] Optionally, the first identification information and the identification information of the first process may be sent through a response message, and the response message may be a response to the message carrying the identification information of the target model group, but this application does not limit this.
[0213] It should be understood that the embodiments of this application are applicable to the scenarios corresponding to the two sub-use cases defined in the model deployment stage mentioned above. In different scenarios, the identification information of the target model group may be carried by different types of messages, so the meaning of the first identification information is different in different scenarios.
[0214] Scenario 1: The scenario corresponding to the sub-use case 1 mentioned above. In a possible implementation of this scenario, the identification information of the target model group in the above S601 is sent by the first device through a request message, and the request message is used to request the second device to deploy the target model group for the third device.
[0215] It should be understood that in such a scenario, the first device determines the timing of sending the above request message to the second device based on the deployment policy of the target model group. The deployment policy can be a requirement for the inference ability of the target model group, or parameters of each functional entity involved in the model deployment process, etc. The deployment policy can be determined by the first device or reported by the third device, and this application does not make any limitations in this regard. Exemplarily, the deployment policy can be that the accuracy of the target model group is greater than 90%, or it reaches a preset time point or preset period for deploying the target model group in terms of time, or a certain network performance index of the third device is monitored to be lower than a preset threshold (e.g., energy efficiency is less than 500 bits per joule), etc., and this application does not make specific limitations in this regard.
[0216] In a possible implementation manner, when the first device determines that the target model group and / or the third device have met the conditions corresponding to the deployment policy, the first device sends a request message to the second device. The request message is used to request the deployment of the target model group, and the request message includes the identification information of the target model group. Correspondingly, the second device receives the request message and creates an instance of the request message and a first process instance.
[0217] In some possible implementation manners, the request message further includes the identification information of each model in the target model group.
[0218] It should be understood that the first process is used to manage the deployment process of the target model group. The first process can be called the deployment process or the loading process of the target model group, and this application does not make specific limitations on the specific name of the first process. In addition, in the embodiments of this application, "loading" and "deploying", without special explanations, can be understood to have the same meaning and will not be repeatedly explained hereinafter.
[0219] In this scenario, there can be the following two examples of the IOC and attributes corresponding to the request message:
[0220] In the first example, the request message can correspond to the MLEntityLoadingRequest object described in Design One above. By associating the MLEntityCoordinationGroup object through the MLEntityLoadingRequest object, it can be indicated that the request message requests the deployment of a model group, and the identification information of the target model group can be represented by the value of the role-related attribute mLEntityCoordinationGroupToLoadRef of the MLEntityLoadingRequest object.
[0221] Optionally, the identification information of each model in the target model group can be represented by the value of the role-related attribute memberMLEntityRefList of the MLEntityCoordinationGroup object associated with the MLEntityLoadingRequest object.
[0222] Correspondingly, the first process can be a process created based on the MLEntityLoadingProcess object described in Design One. The request message can be identified by the value of the role-related attribute MLEntityLoadingRequestRef of the MLEntityLoadingProcess object, and the identification information of the target model group corresponding to the first process can be represented by the value of the role-related attribute LoadedMLEntityCoordinationGroupRef of the MLEntityLoadingProcess object.
[0223] In the second example, the request message can correspond to the MLEntityLoadingRequest object described in Design Two above. The mLEntityGroupInfo attribute of the MLEntityLoadingRequest object is used to indicate that the target model group is requested to be deployed, and the value of mLEntityGroupInfo can be used to represent the identification information of the target model group.
[0224] Optionally, the identification information of each model in the target model group can be correspondingly represented by multiple values of the role-related attribute mLEntityToLoadRef in the MLEntityLoadingRequest object.
[0225] Correspondingly, the first process can be a process created based on the MLEntityLoadingProcess object described in Design Two. The request message can be identified by the value of the role-related attribute MLEntityLoadingRequestRef of the MLEntityLoadingProcess object, and the identification information of the target model group or the identification information of multiple models in the target model group can be correspondingly indicated by one or more values of the attribute LoadedMLEntityRef of the MLEntityLoadingProcess object.
[0226] In the two examples corresponding to Scenario 1, the instance of the request message can be an MLEntityLoadingRequest instance, the first identification information can be the identification of the MLEntityLoadingRequest instance, the instance of the first process can be an MLEntityLoadingProcess instance, and the identification information of the first process can be the identification of the MLEntityLoadingProcess instance.
[0227] Scenario 2: The scenario corresponding to Sub-use case 2 mentioned above. In a possible implementation of this scenario, the identification information of the target model group in S601 above is sent by the first device through a configuration message, and the deployment policy of the target model group is also included in this configuration message.
[0228] In a possible implementation, the first device sends the identification information of the target model group and the deployment policy of the target model group to the second device through a configuration message. After receiving the identification information of the target model group and the deployment policy of the target model group, the second device creates an instance of this configuration message, and creates an instance of the first process after determining that the target model group and / or the third device have met the conditions corresponding to the deployment policy.
[0229] In a possible implementation, this configuration message further includes the identification information of each model in the target model group.
[0230] Optionally, the first device can send the identification information of the target model group and the deployment policy of the target model group through one configuration message, or can send the identification information of the target model group and the deployment policy of the target model group through two configuration messages respectively. This application does not limit the quantity, content, and sequence of the configuration messages.
[0231] Optionally, the configuration policy of the target model group in this scenario can be that the accuracy of the target model group is greater than 90%, or the time reaches the preset time point or preset period for deploying the target model group, or a certain network performance index of the third device is monitored to be lower than the preset threshold (such as: energy efficiency is less than 500 bits per joule), etc. This application does not make specific limitations on this.
[0232] In this scenario, there can be the following two examples for the IOC and attributes corresponding to the configuration message:
[0233] In the first example, the configuration message can correspond to the MLEntityLoadingPolicy object described in Design 1 above. The deployment policy included in the configuration message corresponding to the target model group can be represented by associating the MLEntityCoordinationGroup object through the MLEntityLoadingPolicy object. The identification information of the target model group can be represented by the value of the role-related attribute mLEntityCoordinationGroupToLoadRef of the MLEntityLoadingPolicy object, and the deployment policy of the target model group can be identified by the attribute policyForLoading.
[0234] In a possible implementation, the identification information of each model in the target model group is represented by the value of the role-related attribute memberMLEntityRefList of the MLEntityCoordinationGroup object associated through the MLEntityLoadingPolicy object.
[0235] In another possible implementation, the identification information of multiple models in the target model group is correspondingly indicated by multiple values of the mLEntityId attribute of the MLEntityLoadingPolicy object.
[0236] Correspondingly, the first process can be a process created based on the MLEntityLoadingProcess object described in Design 1 above. The configuration message can be identified by the value of the role-related attribute MLEntityLoadingPolicyRef of the MLEntityLoadingProcess object, and the identification information of the target model group can be represented by the value of the role-related attribute LoadedMLEntityCoordinationGroupRef of the MLEntityLoadingProcess object.
[0237] In the second example, the configuration message can correspond to the MLEntityLoadingPolicy object described in Design 2 above. The identification information of the target model group can be represented by the value of the attribute mLEntityGroupInfo of the MLEntityLoadingPolicy object, and the deployment policy of the target model group can be identified by the attribute policyForLoading.
[0238] Optionally, the identification information of multiple models in the target model group can be correspondingly indicated by multiple values of the mLEntityId attribute.
[0239] Correspondingly, the first process, which is created based on the MLEntityLoadingProcess object described in Design 2, can identify the configuration message through the value of the role-related attribute MLEntityLoadingPolicyRef of the MLEntityLoadingProcess object, and can correspond to the identification information of multiple models in the target model group through multiple values of the attribute LoadedMLEntityRef of the MLEntityLoadingProcess object.
[0240] In the two examples corresponding to Scenario 2, the instance of the configuration message can be an MLEntityLoadingPolicy instance, the first identification information can be the identification of the MLEntityLoadingPolicy instance, the instance of the first process can be an MLEntityLoadingProcess instance, and the identification information of the first process can be the identification of the MLEntityLoadingProcess instance.
[0241] The first process in the above two scenarios is used to manage the deployment process of the target model group. Optionally, the management method of the deployment process of the target model group can include two methods: "batch management" and "independent management", but this application does not make specific limitations on this. Under different deployment process management methods, the meaning of the first process is different. The following takes "batch management" and "independent management" as examples for detailed description.
[0242] The first method: Batch management
[0243] The target model group includes multiple models. "Batch management" can be understood as regarding these multiple models as a whole during the model deployment process and performing overall deployment, that is, deploying / loading the models in the trained / tested model group into the target inference function at the same time. In this case, the first process corresponds to the overall deployment process of the target model group. Exemplarily, the second device creates an MLEntityLoadingProcess instance to manage the deployment process of the target model group.
[0244] It should be understood that the MLEntityLoadingProcess instance has a status attribute progressStatus, and the value of this attribute can be any of the following: RUNNING, CANCELLING, SUSPENDED, FINISHED, or CANCELLED. Among them, when the value of progressStatus of the MLentityLoadingProcess is RUNNING, it can be understood that the model deployment process is in progress. Therefore, RUNNING can also be understood as "deploying" or "loading", and this application does not make a limitation on this.
[0245] Exemplarily, when the second device creates the MLEntityLoadingProcess instance, or after creating the MLEntityLoadingProcess instance, and the value of progressStatus of this MLEntityLoadingProcess instance is RUNNING, the second device can send the above-mentioned second indication information or third indication information to the third device.
[0246] Optionally, the second device can report the status of the first process to the first device.
[0247] In the case where the deployment process management method is batch management, the number of the first processes can be one, and the identification information of the first process identifies a process instance that can represent the overall deployment process of the target model group.
[0248] In a possible implementation manner, the second device can actively report the status of the first process to the first device. The method further includes: the second device sends a first message, and the first message includes first identification information and the status of the first process, or the identification information of the first process and the status of the first process. Correspondingly, the first device receives this first message.
[0249] It should be understood that the first message is used to report the status of the first process. The first identification information is associated with the message carrying the identification information of the target model group and has a mapping relationship with the first process. Therefore, the second device can use the first identification information or the identification information of the first process to identify the first process. The status of the first process can include: RUNNING, CANCELLING, CANCELLED, SUSPENDED, or FINISHED.
[0250] In another possible implementation, the second device may report the status of the first process to the first device in response to a first request from the first device. Specifically, before the second device sends the first information, the second device receives a first request from the first device, where the first request is used to request the status of the first process, and the first request includes first identification information, or the identification information of the first process.
[0251] Optionally, the first device may control the first process. In one possible implementation, the method further includes: the first device sends second information, where the second information is used to control the overall deployment process of the target model group, and the third information includes first identification information and a control instruction for the first process, or the identification information of the first process and a control instruction, and the control instruction includes cancel, pause, or start. Correspondingly, the second device receives the second information.
[0252] In the embodiments of the present application, through the model deployment method of batch management, the number of the first processes may be one, and the second device may only maintain one MLEntityLoadingProcess instance to implement the deployment of the target model group. During the status request and reporting process of the first process, the target model group may be regarded as a whole, and during the control process of the first process, the target model group may also be controlled as a whole, without the need to carry the information of each model in the target model group, which is beneficial to reducing the signaling overhead between devices and improving the deployment efficiency of the target model group.
[0253] The second method: independent management
[0254] "Independent management" can be understood as creating independent processes for the management of the deployment process for each model included in the target model group. In some examples, these multiple models may be deployed sequentially during the deployment process of the target model group, that is, the multiple models in the trained target model group are deployed to the target inference function sequentially; in other examples, these multiple models may be deployed simultaneously during the deployment process of the target model group, that is, the multiple models in the trained target model group are deployed to the target inference function simultaneously. The present application does not make specific limitations on the deployment order of the multiple models in the target model group.
[0255] In this case, the first process includes multiple processes, and these multiple processes respectively correspond to the deployment processes of multiple models in the target model group. The identification information of the first process includes multiple identification information corresponding to the multiple processes. Exemplarily, the second device creates multiple MLEntityLoadingProcess instances for managing the deployment process of the target model group, and these multiple MLEntityLoadingProcess instances respectively correspond to multiple models in the target model group.
[0256] When the deployment process management method is independent management, the second device reports the status of the first process to the first device. The deployment process of a certain model in the target model group can be identified by the first identification information and the identification information of this model, or the deployment process of this process can be identified by the identification information of a certain process.
[0257] In a possible implementation manner, the second device actively sends the third information, and the third information is used to report the status of all or part of the processes among multiple processes. The third information includes the first identification information, the identification information of each model corresponding to all or part of the processes among multiple processes, and the status of each of all or part of the processes; or, the identification information of all or part of the processes, and the status of each of all or part of the processes.
[0258] In another possible implementation manner, the second device can report the status of all or part of the processes among multiple processes to the first device in response to the second request of the first device. Specifically, before the second device sends the third information, the second device receives the second request from the first device, and the second request is used to request the status of all or part of the processes among multiple processes. The second request includes the first identification information, the identification information of each model corresponding to all or part of the processes among multiple processes, or the identification information of all or part of the processes.
[0259] Optionally, the first device can control multiple processes included in the first process based on a control instruction. In a possible implementation manner, the method further includes: the first device sends the fourth information, and the fourth information is used to control the deployment processes of all or part of the models among multiple models. The fourth information includes the first identification information, the identification information of each of all or part of the models, and the control instruction for the deployment process of each of all or part of the models; or, the identification information of the deployment process of each of all or part of the models, and the control instruction.
[0260] In the embodiments of this application, through the model deployment method of independent management, the first process can include multiple processes, and the multiple processes correspond one-to-one to the identification information of each model among multiple models. The multiple processes respectively correspond to the deployment processes of multiple models. In this way, the first device can respectively obtain and control the deployment status of each model in the target model group, and can reflect the different deployment statuses of each model. In the case where some models in the target model group fail to be deployed, it can be determined which models have failed to be deployed. Based on the identification information of the partially failed models, these models can be redeployed without redeploying the entire target model group, which is beneficial to improving the management accuracy and reducing unnecessary resource waste.
[0261] In a possible implementation, the deployment process management method of the target model group is the default one, which includes batch management or independent management, that is, no explicit instruction is required from either party. Optionally, the deployment process management method can be agreed upon by a protocol, which is not limited in this application.
[0262] In another possible implementation, the deployment process management method of the target model group is indicated by the first device. The implementation method can be that the first device sends the first indication information to the second device, and the first indication information is used to indicate the deployment process management method of the target model group.
[0263] Optionally, the first indication information can be represented by the value of the attribute groupBatchIndicator. In the above scenario one, the first indication information can be represented by the value of the attribute groupBatchIndicator of the MLEntityLoadingRequest object; in the above scenario two, the first indication information can be represented by the value of the attribute groupBatchIndicator of the MLEntityLoadingPolicy object, but this application does not limit this.
[0264] Optionally, the values of the first indication information groupBatchIndicator can include 0 and 1. When the value is 0, it indicates that the deployment process management method of the target model group is batch management, and when the value is 1, the deployment process management method of the target model group is independent management. This application does not specifically limit the values of groupBatchIndicator and their meanings.
[0265] Next, taking Figure 6 the first device in [ ] as NMS, the second device as EMS, and the third device that undertakes the model inference function as gNB as an example, for the above scenario one, scenario two, and the two deployment process management methods of batch management and independent management, combined with Figures 7 to 10 it will be described in detail in the following four cases.
[0266] The first case: Scenario one + Batch management
[0267] Figure 7 This is a schematic flowchart of a model deployment method 700 provided by an embodiment of this application. The method 700 includes:
[0268] S701. When NMS determines that the conditions for the deployment policy of the target model group are met, it sends an MLEntityLoadingRequest to the EMS. This MLEntityLoadingRequest contains the identification information of the target model group. Correspondingly, the EMS receives this MLEntityLoadingRequest.
[0269] Optionally, this MLEntityLoadingRequest also contains the identification information of each model in the target model group.
[0270] In a possible implementation manner, corresponding to the foregoing Design One, the identification information of the target model group can be indicated by the MLEntityCoordinationGroup associated with the MLEntityLoadingRequest. Optionally, the identification information of each model in the target model group can also be represented by the value of the role-related attribute memberMLEntityRefList of the MLEntityCoordinationGroup object.
[0271] In another possible implementation manner, corresponding to the foregoing Design Two, the identification information of the target model group can be indicated by the attribute mLEntityGroupInfo of the MLEntityLoadingRequest object. Optionally, the identification information of each model in the target model group can be correspondingly represented by multiple values of the role-related attribute mLEntityToLoadRef in this MLEntityLoadingRequest object.
[0272] S702. The EMS creates an MLEntityLoadingRequest instance and an MLEntityLoadingProcess instance.
[0273] It should be understood that an MLEntityLoadingProcess instance created in S702 corresponds to the overall deployment process of the target model group, that is, multiple models in the target model group are deployed and managed as a whole.
[0274] It should also be understood that the management method of the deployment process of the target model group can be the default or indicated by the NMS. In the case of being indicated by the NMS, this indication information can be indicated by the groupBatchIndicator included in the MLEntityLoadingRequest. In this embodiment, the groupBatchIndicator indicates that the management method of the deployment process of the target model group is batch management.
[0275] S703. The EMS sends a response message of the MLEntityLoadingRequest to the NMS. The response message includes the identifier of the MLEntityLoadingRequest instance and the identifier of the MLEntityLoadingProcess instance. Correspondingly, the NMS receives the response message.
[0276] S704. The NMS sends a request for the deployment process status of the target model group to the EMS, including the identifier of the MLEntityLoadingRequest instance. Correspondingly, the EMS receives the request.
[0277] In another implementation of S704, when the NMS sends a request for the deployment process status of the target model group to the EMS, it includes the identifier of the MLEntityLoadingProcess instance.
[0278] S705. The EMS reports the deployment process status of the target model group to the NMS, including the identifier of the MLEntityLoadingRequest instance + the progressStatus value corresponding to the MLEntityLoadingProcess instance. Correspondingly, the NMS receives the deployment process status of the target model group.
[0279] In another implementation of S705, the EMS reports the deployment process status of the target model group to the NMS, including the identifier of the MLEntityLoadingProcess instance + the progressStatus value corresponding to the MLEntityLoadingProcess instance.
[0280] S706. The NMS sends a control message for the deployment process of the target model group to the EMS, including the identifier of the MLEntityLoadingRequest instance + the control instruction. Correspondingly, the EMS executes the control instruction.
[0281] In another implementation of S706, the NMS sends a control message for the deployment process of the target model group to the EMS, including the identifier of the MLEntityLoadingProcess instance + the control instruction.
[0282] S707. The EMS sends an indication message to the gNB. The indication message is used to indicate the activation of the target model group. The indication message includes the identification information of the target model group. Correspondingly, the gNB receives the indication message.
[0283] It should be understood that S707 can be understood as a process in which the EMS executes model deployment and loads the target model group into the target gNB. Optionally, the number of target gNBs can be one or more, and this application does not limit this.
[0284] Optionally, the indication message can be the above-mentioned second indication message or third indication message, and this application does not limit this.
[0285] S708. The gNB activates the target model group based on the identification information of the target model group.
[0286] In a possible implementation manner, after S708, method 700 further includes: the gNB performs an inference function based on the target model group.
[0287] It should be understood that the above S704, S705, and S706 are all optional steps.
[0288] In the embodiments of this application, by including the identification information of the target model group in the MLEntityLoadingRequest, it is indicated that this model deployment is for a model group, so that during the deployment process, the gNB can obtain the description file of the target model group based on the identification information of the target model group, enabling multiple models in the activated target model group to perform joint inference. In addition, the EMS maintains an MLEntityLoadingProcess instance for this model group, and tracks the process of the gNB activating the target model group through this instance, which is beneficial to saving energy consumption and improving the model deployment efficiency.
[0289] The second case: Scenario 1 + independent management
[0290] Figure 8 This is a schematic flowchart of a model deployment method 800 provided by the embodiments of this application. The method 800 includes:
[0291] S801. When the NMS determines that the deployment policy of the target model group has been satisfied, it sends an MLEntityLoadingRequest to the EMS, and the MLEntityLoadingRequest contains the identification information of the target model group. Correspondingly, the EMS receives the MLEntityLoadingRequest.
[0292] Optionally, the MLEntityLoadingRequest further contains the identification information of each model in the target model group.
[0293] It should be understood that the representation methods of the identification information of the target model group and the identification information of each model in the target model group in method 800 can be the same as those in the above method 700, and will not be elaborated here.
[0294] S802. The EMS creates an MLEntityLoadingRequest instance and multiple MLEntityLoadingProcess instances.
[0295] It should be understood that the multiple MLEntityLoadingProcess instances correspond one-to-one to the multiple models in the target model group.
[0296] It should also be understood that the deployment process management method of the target model group can be the default or NMS-indicated. In the case of NMS indication, the indication information can be indicated by the groupBatchIndicator included in the MLEntityLoadingRequest. In this embodiment, the groupBatchIndicator indicates that the deployment process management method of the target model group is independent management.
[0297] S803. The EMS sends a response message of the MLEntityLoadingRequest to the NMS. The response message includes the identifier of the MLEntityLoadingRequest instance and the multiple identifiers corresponding to the multiple MLEntityLoadingProcess instances. Correspondingly, the NMS receives the response message.
[0298] S804. The NMS sends a deployment process status request for all or part of the models in the target model group to the EMS, including the identifier of the MLEntityLoadingRequest instance + the identifier information of all or part of the requested models. Correspondingly, the EMS receives the request.
[0299] In another implementation manner of S804, when the NMS sends a deployment process status request for the target model group to the EMS, it includes all or part of the identifier information of the MLEntityLoadingProcess instances.
[0300] S805. The EMS reports the deployment process status of all or part of the models in the target model group to the NMS, including the identifier of the MLEntityLoadingRequest instance + the progressStatus value of the MLEntityLoadingProcess instances corresponding to all or part of the models. Correspondingly, the NMS receives the deployment process status of the target model group.
[0301] In another implementation of S805, the EMS reports the deployment process status of all or part of the models in the target model group to the NMS, including the identifier of the MLEntityLoadingProcess instance + the progressStatus value of the MLEntityLoadingProcess instance corresponding to all or part of the models.
[0302] S806. The NMS sends control messages for the deployment process of all or part of the models in the target model group to the EMS, including the identifier of the MLEntityLoadingRequest instance + the identifier information of all or part of the models to be controlled + the control instructions for the MLEntityLoadingProcess instance corresponding to each model identifier. Correspondingly, the EMS executes the control instructions.
[0303] In another implementation of S806, the NMS sends control messages for the deployment process of the target model group to the EMS, including the identifier of the MLEntityLoadingProcess instance + the control instructions for each MLEntityLoadingProcess instance.
[0304] S807. The EMS sends an indication message to the gNB, and this indication message is used to indicate the activation of the target model group. This indication message includes the identifier information of the target model group. Correspondingly, the gNB receives this indication message.
[0305] S808. The gNB activates the target model group based on the identifier information of the target model group.
[0306] Among them, S804, S805, and S806 are all optional steps.
[0307] In the embodiments of the present application, through the independent management of the deployment process of the target model group, the NMS can obtain and control the deployment status of each model in the target model group. In the case where some models in the target model group fail to be deployed, it can be determined which models have failed to be deployed. Based on the identifier information of the partially failed models, these models can be redeployed without redeploying the entire target model group, which is beneficial to improving the management accuracy and reducing unnecessary resource waste.
[0308] The third case: Scenario Two + Batch Management
[0309] Figure 9 This is a schematic flowchart of a model deployment method 900 provided by the embodiments of the present application. The method 900 includes:
[0310] S901. The NMS sends an MLEntityLoadingPolicy to the EMS. The MLEntityLoadingPolicy contains the identification information of the target model group and the deployment policy of the target model group. Correspondingly, the EMS receives the MLEntityLoadingPolicy.
[0311] Optionally, the MLEntityLoadingPolicy further contains the identification information of each model in the target model group.
[0312] In a possible implementation manner, corresponding to the foregoing Design One, the identification information of the target model group can be indicated by the MLEntityCoordinationGroup associated with the MLEntityLoadingPolicy. Optionally, the identification information of each model in the target model group can be represented by the value of the role-related attribute memberMLEntityRefList of the MLEntityCoordinationGroup object; or, the identification information of multiple models in the target model group can also be correspondingly indicated by multiple values of the mLEntityId attribute of the MLEntityLoadingPolicy object.
[0313] In another possible implementation manner, corresponding to the foregoing Design Two, the identification information of the target model group can be indicated by the attribute mLEntityGroupInfo of the MLEntityLoadingPolicy object.
[0314] Optionally, the identification information of multiple models in the target model group can be correspondingly indicated by multiple values of the mLEntityId attribute of the MLEntityLoadingPolicy object.
[0315] S902. The EMS creates an MLEntityLoadingPolicy instance.
[0316] Under the condition that the EMS determines that the deployment policy of the target model group has been satisfied, the EMS executes S903.
[0317] S903. The EMS creates an MLEntityLoadingProcess instance.
[0318] It should be understood that an MLEntityLoadingProcess instance created by the EMS in S903 corresponds to the overall deployment process of the target model group, that is, multiple models in the target model group are deployed and managed as a whole.
[0319] It should also be understood that the deployment process management method of the target model group can be the default or NMS-indicated. In the case of NMS indication, the indication information can be indicated by the groupBatchIndicator included in the MLEntityLoadingPolicy. In this embodiment, the groupBatchIndicator indicates that the deployment process management method of the target model group is batch management.
[0320] S904. The EMS sends the identifier of the MLEntityLoadingPolicy instance and the identifier of the MLEntityLoadingProcess instance to the NMS. Correspondingly, the NMS receives the identifier of the MLEntityLoadingPolicy instance and the identifier of the MLEntityLoadingProcess instance.
[0321] Optionally, the EMS can send the identifier of the MLEntityLoadingPolicy instance to the NMS after executing S902, or can send the identifier of the MLEntityLoadingPolicy instance to the NMS after executing S903. The identifier of the MLEntityLoadingPolicy instance and the identifier of the MLEntityLoadingProcess instance can be carried by the same message or by different messages. This application does not make any limitations on this.
[0322] S905. The NMS sends a request for the deployment process status of the target model group to the EMS, including the identifier of the MLEntityLoadingPolicy instance. Correspondingly, the EMS receives this request.
[0323] In another implementation manner of S905, the request for the deployment process status of the target model group sent by the NMS to the EMS includes the identifier of the MLEntityLoadingProcess instance.
[0324] S906. The EMS reports the deployment process status of the target model group to the NMS, including the identifier of the MLEntityLoadingPolicy instance + the progressStatus value corresponding to the MLEntityLoadingProcess instance. Correspondingly, the NMS receives the deployment process status of the target model group.
[0325] In another implementation manner of S906, the EMS reports the deployment process status of the target model group to the NMS, including the identifier of the MLEntityLoadingProcess instance + the progressStatus value corresponding to the MLEntityLoadingProcess instance.
[0326] S907. The NMS sends a control message for the deployment process of the target model group to the EMS, including the identifier of the MLEntityLoadingPolicy instance + a control instruction. Correspondingly, the EMS executes the control instruction.
[0327] In another implementation manner of S907, the NMS sends a control message for the deployment process of the target model group to the EMS, including the identifier of the MLEntityLoadingProcess instance + a control instruction.
[0328] S908. The EMS sends an indication message to the gNB, and this indication message is used to indicate the activation of the target model group. This indication message includes the identification information of the target model group. Correspondingly, the gNB receives this indication message.
[0329] It should be understood that S908 can be understood as the process in which the EMS performs model deployment and loads the target model group into the target gNB. Optionally, the number of target gNBs can be one or more, and this application does not make any limitation in this regard.
[0330] Optionally, this indication message can be the above-mentioned second indication message or third indication message, and this application does not make any limitation in this regard.
[0331] S909. The gNB activates the target model group based on the identification information of the target model group.
[0332] In a possible implementation manner, after S909, method 900 further includes: the gNB performs an inference function based on the target model group.
[0333] It should be understood that the above S905, S906, and S907 are all optional steps.
[0334] The beneficial effects of the embodiments of this application are similar to those of the foregoing method 700 and will not be elaborated here.
[0335] The fourth case: Scenario 2 + independent management
[0336] Figure 10 This is a schematic flowchart of a model deployment method 1000 provided by an embodiment of this application. This method 1000 includes:
[0337] S1001. The NMS sends the MLEntityLoadingPolicy to the EMS, and this MLEntityLoadingPolicy contains the identification information of the target model group and the deployment policy of the target model group. Correspondingly, the EMS receives this MLEntityLoadingPolicy.
[0338] Optionally, the MLEntityLoadingPolicy further includes identification information of each model in the target model group.
[0339] It should be understood that the representation of the identification information of the target model group and the identification information of each model in the target model group in method 1000 may be the same as that in method 900 above, and will not be elaborated here.
[0340] S1002. The EMS creates an MLEntityLoadingPolicy instance.
[0341] Under the condition that the EMS determines that the deployment policy of the target model group has been satisfied, the EMS executes S1003.
[0342] S1003. The EMS creates multiple MLEntityLoadingProcess instances.
[0343] It should be understood that the multiple MLEntityLoadingProcess instances correspond one by one to the multiple models in the target model group.
[0344] It should also be understood that the management mode of the deployment process of the target model group can be the default or indicated by the NMS. In the case of being indicated by the NMS, the indication information can be indicated by the groupBatchIndicator included in the MLEntityLoadingPolicy. In this embodiment, the groupBatchIndicator indicates that the management mode of the deployment process of the target model group is independent management.
[0345] S1004. The EMS sends the identification of the MLEntityLoadingPolicy instance and the multiple identifications corresponding to the multiple MLEntityLoadingProcess instances to the NMS. Correspondingly, the NMS receives this message.
[0346] Optionally, the EMS may send the identification of the MLEntityLoadingPolicy instance to the NMS immediately after executing S1002, or may send the identification of the MLEntityLoadingPolicy instance to the NMS after executing S1003. The identification of the MLEntityLoadingPolicy instance and the multiple identifications corresponding to the multiple MLEntityLoadingProcess instances may be carried by the same message or by different messages. This application does not make any limitations in this regard.
[0347] S1005. The NMS sends a request for the deployment process status of all or part of the models in the target model group to the EMS, including the identifier of the MLEntityLoadingPolicy instance + the identifier information of all or part of the requested models. Correspondingly, the EMS receives this request.
[0348] In another implementation of S1005, when the NMS sends a request for the deployment process status of the target model group to the EMS, it includes all or part of the identifier information of the MLEntityLoadingProcess instance.
[0349] S1006. The EMS reports the deployment process status of all or part of the models in the target model group to the NMS, including the identifier of the MLEntityLoadingPolicy instance + the progressStatus value of the MLEntityLoadingProcess instance corresponding to all or part of the models. Correspondingly, the NMS receives the deployment process status of the target model group.
[0350] In another implementation of S1006, the EMS reports the deployment process status of all or part of the models in the target model group to the NMS, including the identifier of the MLEntityLoadingProcess instance + the progressStatus value of the MLEntityLoadingProcess instance corresponding to all or part of the models.
[0351] S1007. The NMS sends a control message for the deployment process of all or part of the models in the target model group to the EMS, including the identifier of the MLEntityLoadingPolicy instance + the identifier information of all or part of the models to be controlled + the control instructions for the MLEntityLoadingProcess instance corresponding to each model identifier. Correspondingly, the EMS executes this control instruction.
[0352] In another implementation of S1007, the NMS sends a control message for the deployment process of the target model group to the EMS, including the identifier of the MLEntityLoadingProcess instance + the control instructions for each MLEntityLoadingProcess instance.
[0353] S1008. The EMS sends an indication message to the gNB, and this indication message is used to indicate the activation of the target model group. This indication message includes the identifier information of the target model group. Correspondingly, the gNB receives this indication message.
[0354] S1009. The gNB activates the target model group based on the identifier information of the target model group.
[0355] Among them, S1005, S1006, and S1007 are all optional steps.
[0356] The beneficial effects of the embodiments of this application are similar to those of the foregoing method 800 and will not be elaborated here.
[0357] It should be understood that the magnitudes of the sequence numbers of the respective steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not impose any limitation on the implementation process of the embodiments of this application.
[0358] In the foregoing text, in combination with Figures 6 to 10 , the model deployment method of the embodiments of this application has been described in detail. Next, in combination with Figure 11 and Figure 12 , the communication device of the embodiments of this application will be described in detail.
[0359] Figure 11 FIG. 1100 shows a communication device 1100 provided by an embodiment of this application. The communication device 1100 includes: a transceiver module 1101 and a processing module 1102.
[0360] In a possible implementation manner, the communication device 1100 is used to implement the steps and processes corresponding to the foregoing second device.
[0361] Among them, the transceiver module 1101 is configured to: receive identification information of a target model group, where the target model group includes multiple models, and the multiple models are jointly trained and / or jointly tested; the processing module 1102 is configured to: deploy the target model group based on the identification information of the target model group.
[0362] Optionally, the transceiver module 1101 is further configured to: send first identification information and identification information of a first process, where the first identification information is associated with a message carrying the identification information of the target model group, and the first identification information has a mapping relationship with the first process, and the first process is used to manage the deployment process of the target model group.
[0363] Optionally, the transceiver module 1101 is further configured to: receive first indication information, where the first indication information is used to indicate a management method for the deployment process of the target model group.
[0364] Optionally, the management method for the deployment process of the target model group is batch management.
[0365] Optionally, the management method for the deployment process of the target model group is default to batch management.
[0366] Optionally, the first process corresponds to the overall deployment process of the target model group.
[0367] Optionally, the transceiver module 1101 is further configured to: send a first message, where the first message includes first identification information and the status of a first process, or the identification information of the first process and the status of the first process, and the status includes: running, canceling, canceled, paused, or ended.
[0368] Optionally, the transceiver module 1101 is further configured to: receive a second message, where the second message is used to control the overall deployment process of a target model group, and the second message includes first identification information and a control instruction for the first process, or the identification information of the first process and the control instruction, and the control instruction includes cancel, pause, or start.
[0369] Optionally, the default management method for the deployment process of the target model group is independent management.
[0370] Optionally, the management method for the deployment process of the target model group is independent management.
[0371] Optionally, the first process includes multiple processes, where the multiple processes are in one-to-one correspondence with the identification information of each model in multiple models, the multiple processes respectively correspond to the deployment processes of the multiple models, and the identification information of the first process includes the multiple identification information corresponding to the multiple processes.
[0372] Optionally, the transceiver module 1101 is further configured to: send a third message, where the third message includes first identification information, the identification information of each model corresponding to all or part of the multiple processes, and the status of all or part of the processes respectively; or,
[0373] the identification information of all or part of the processes, and the status of all or part of the processes respectively, and the status includes: running, canceling, canceled, paused, or ended.
[0374] Optionally, the transceiver module 1101 is further configured to: receive a fourth message, where the fourth message is used to control the deployment processes of all or part of the models in the multiple models, and the fourth message includes first identification information, the identification information of each model in all or part of the models, and control instructions for the deployment processes of each model in all or part of the models; or, the identification information of the deployment processes of each model in all or part of the models, and the control instructions.
[0375] Optionally, the identification information of the target model group is sent through a request message or a configuration message.
[0376] Optionally, the transceiver module 1101 is further configured to: receive the identification information of each model in the target model group.
[0377] In another possible implementation, the communication device 1100 is used to implement the steps and processes corresponding to the above first device.
[0378] Among them, the processing module 1102 is used to: determine the identification information of the target model group, where the target model group includes multiple models that have been jointly trained and / or jointly tested; the transceiver module 1101 is used to: send the identification information of the target model group.
[0379] Optionally, the transceiver module 1101 is further used to: receive the first identification information and the identification information of the first process, where the first identification information is associated with the message carrying the identification information of the target model group, and there is a mapping relationship between the first identification information and the first process, and the first process is used to manage the deployment process of the target model group.
[0380] Optionally, the transceiver module 1101 is further used to: send the first indication information, where the first indication information is used to indicate the management method of the deployment process of the target model group.
[0381] Optionally, the management method of the deployment process of the target model group is batch management.
[0382] Optionally, the management method of the deployment process of the target model group is default to batch management.
[0383] Optionally, the first process corresponds to the overall deployment process of the target model group.
[0384] Optionally, the transceiver module 1101 is further used to: receive the first information, where the first information includes the first identification information and the status of the first process, or the identification information of the first process and the status of the first process, and the status includes: running, canceling, canceled, paused or ended.
[0385] Optionally, the transceiver module 1101 is further used to: send the second information, where the second information is used to control the overall deployment process of the target model group, and the second information includes the first identification information and the control instruction for the first process, or the identification information of the first process and the control instruction, and the control instruction includes cancel, pause or start.
[0386] Optionally, the management method of the deployment process of the target model group is default to independent management.
[0387] Optionally, the management method of the deployment process of the target model group is independent management.
[0388] Optionally, the first process includes multiple processes, where the multiple processes correspond one-to-one to the identification information of each model in the multiple models, the multiple processes respectively correspond to the deployment processes of the multiple models, and the identification information of the first process includes the multiple identification information corresponding to the multiple processes.
[0389] Optionally, the transceiver module 1101 is further configured to: receive third information, where the third information includes first identification information, identification information of each model corresponding to all or part of multiple processes, and the states of all or part of the processes respectively; or, identification information of all or part of the processes, and the states of all or part of the processes respectively, where the states include: running, canceling, canceled, paused, or ended.
[0390] Optionally, the transceiver module 1101 is further configured to: send fourth information, where the fourth information is used to control the deployment processes of all or part of multiple models, and the fourth information includes first identification information, identification information of each of all or part of the models, and control instructions for the deployment processes of each of all or part of the models; or, identification information of the deployment processes of each of all or part of the models, and control instructions.
[0391] Optionally, the identification information of the target model group is sent through a request message or a configuration message.
[0392] Optionally, the transceiver module 1101 is further configured to: send the identification information of each model in the target model group.
[0393] It should be understood that the device 1100 here is embodied in the form of a functional module. The term "module" here may refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a proprietary processor, or a group of processors, etc.) for executing one or more software or firmware programs, and a memory, a combined logic circuit, and / or other suitable components that support the described functions. In an alternative example, those skilled in the art can understand that the device 1100 may specifically be the first device or the second device in the above embodiments, or, the functions described in the above embodiments may be integrated in the device 1100, and the device 1100 may be used to execute the respective processes and / or steps corresponding to the first device or the second device in the above method embodiments. To avoid repetition, it will not be elaborated here.
[0394] The above device 1100 has the function of implementing the corresponding steps executed by the first device or the second device in the above method; the above function may be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0395] In the embodiments of the present application, Figure 11 the device 1100 in
[0396] Figure 12 Fig. 0 shows a schematic block diagram of a communication device 1200 provided by an embodiment of the present application. The device 1200 includes a processor 1201, a transceiver 1202, and a memory 1203. Among them, the processor 1201, the transceiver 1202, and the memory 1203 communicate with each other through an internal connection path. The memory 1203 is used to store instructions, and the processor 1201 is used to execute the instructions stored in the memory 1203 to control the transceiver 1202 to send signals and / or receive signals.
[0397] It should be understood that the device 1200 may specifically be the first device or the second device in the above embodiments, and may be used to execute the respective steps and / or processes corresponding to the first device or the second device in the above method embodiments. Optionally, the memory 1203 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type. The processor 1201 may be used to execute the instructions stored in the memory, and when the processor 1201 executes the instructions stored in the memory, the processor 1201 is used to execute the respective steps and / or processes of the above method embodiments. The transceiver 1202 may include a transmitter and a receiver. The transmitter may be used to implement the respective steps and / or processes corresponding to the above transceiver for performing the sending action, and the receiver may be used to implement the respective steps and / or processes corresponding to the above transceiver for performing the receiving action.
[0398] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0399] In the implementation process, the respective steps of the above method may be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed and completed by the hardware processor, or executed and completed by a combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor executes the instructions in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0400] The present application also provides a computer-readable storage medium for storing a computer program, which is used to implement the method shown in the above method embodiments.
[0401] The present application also provides a computer program product, which includes computer program code (which can also be referred to as a computer program or instruction). When the computer program runs on a computer, the computer can execute the method shown in the above method embodiments.
[0402] The present application also provides a communication method, which is applied to a communication system including a first device and a second device, and the above method embodiments are executed by the second device and / or the second device.
[0403] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0404] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0405] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings, direct couplings, or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0406] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0407] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit.
[0408] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, part of the technical solution of the present application can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.< / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc>
Claims
1. A model deployment method, characterized in that, including: receiving identification information of a target model group, where the target model group includes multiple models, and the multiple models are jointly trained and / or jointly tested; deploying the target model group based on the identification information of the target model group.
2. The method according to claim 1, characterized in that, The method further includes: sending first identification information and identification information of a first process, where the first identification information is associated with a message carrying the identification information of the target model group, and there is a mapping relationship between the first identification information and the first process, and the first process is used to manage the deployment process of the target model group.
3. The method according to claim 2, wherein The deployment process of the target model group is managed in a batch manner.
4. The method according to claim 3, wherein The first process corresponds to the overall deployment process of the target model group.
5. The method according to claim 3 or 4, characterized in that, The method further includes: sending first information, where the first information includes the first identification information and the status of the first process, or the identification information of the first process and the status of the first process, and the status includes: running, canceling, canceled, paused, or ended.
6. The method according to any one of claims 3 to 5, characterized in that, The method further includes: receiving second information, where the second information is used to control the overall deployment process of the target model group, and the second information includes the first identification information and a control instruction for the first process, or the identification information of the first process and the control instruction, and the control instruction includes cancel, pause, or start.
7. The method according to claim 2, wherein The deployment process of the target model group is managed independently.
8. The method according to claim 7, characterized in that The first process includes multiple processes, and the multiple processes correspond one-to-one to the identification information of each of the multiple models. The multiple processes respectively correspond to the deployment processes of the multiple models, and the identification information of the first process includes the multiple identification information corresponding to the multiple processes.
9. The method according to claim 7 or 8, characterized in that, The method further includes: sending third information, where the third information includes the first identification information, the identification information of each of all or part of the processes corresponding to the multiple processes, and the status of each of the all or part of the processes; or the identification information of the all or part of the processes, and the status of each of the all or part of the processes, and the status includes: running, canceling, canceled, paused, or ended.
10. The method according to any one of claims 7 to 9, characterized in that The method further includes: receiving fourth information, where the fourth information is used to control the deployment processes of all or part of the multiple models, and the fourth information includes the first identification information, the identification information of each of the all or part of the models, and a control instruction for the deployment process of each of the all or part of the models; or the identification information of the deployment process of each of the all or part of the models, and the control instruction.
11. The method according to any one of claims 1 to 10, characterized in that, The method further includes: receiving first indication information, where the first indication information is used to indicate the management method of the deployment process of the target model group.
12. The method according to any one of claims 3 to 6, characterized in that The default management method of the deployment process of the target model group is batch management.
13. The method according to any one of claims 7 to 10, characterized in that, The default management method of the deployment process of the target model group is independent management.
14. The method according to any one of claims 1 to 13, characterized in that, The identification information of the target model group is sent through a request message or a configuration message.
15. The method according to any one of claims 1 to 14, characterized in that, The method further includes: receiving the identification information of each model in the target model group.
16. A model deployment method, characterized in that, including: Determine the identification information of the target model group, where the target model group includes multiple models, and the multiple models are jointly trained and / or jointly tested; Send the identification information of the target model group.
17. The method according to claim 16, wherein The method further includes: Receive the first identification information and the identification information of the first process, where the first identification information is associated with a message carrying the identification information of the target model group, and there is a mapping relationship between the first identification information and the first process, and the first process is used to manage the deployment process of the target model group.
18. The method according to claim 17, wherein The deployment process management method of the target model group is batch management.
19. The method according to claim 18, wherein The first process corresponds to the overall deployment process of the target model group.
20. The method according to claim 18 or 19, characterized in that, The method further includes: Receive the first information, where the first information includes the first identification information and the status of the first process, or the identification information of the first process and the status of the first process, and the status includes: running, canceling, canceled, paused or ended.
21. The method according to any one of claims 17 to 19, characterized in that, The method further includes: Send the second information, where the second information is used to control the overall deployment process of the target model group, and the second information includes the first identification information and a control instruction for the first process, or the identification information of the first process and the control instruction, and the control instruction includes cancel, pause or start.
22. The method according to claim 17, wherein The deployment process management method of the target model group is independent management.
23. The method according to claim 22, wherein The first process includes multiple processes, and the multiple processes correspond one-to-one to the identification information of each model in the multiple models. The multiple processes respectively correspond to the deployment processes of the multiple models, and the identification information of the first process includes the multiple identification information corresponding to the multiple processes.
24. The method according to claim 22 or 23, characterized in that The method further includes: Receive the third information, where the third information includes the first identification information, the identification information of each model corresponding to all or part of the multiple processes, and the respective statuses of all or part of the processes; or, The identification information of all or part of the processes, and the respective statuses of all or part of the processes, and the status includes: running, canceling, canceled, paused or ended.
25. The method according to any one of claims 22 to 24, characterized in that, The method further includes: Send the fourth information, where the fourth information is used to control the deployment processes of all or part of the multiple models, and the fourth information includes the first identification information, the identification information of each model in all or part of the models, and a control instruction for the deployment process of each model in all or part of the models; or, The identification information of the deployment process of each model in all or part of the models, and the control instruction.
26. The method according to any one of claims 16 to 25, characterized in that, The method further includes: Send the first indication information, where the first indication information is used to indicate the deployment process management method of the target model group.
27. The method according to any one of claims 18 to 21, characterized in that, The default deployment process management method of the target model group is batch management.
28. The method according to any one of claims 22 to 25, characterized in that The default deployment process management method of the target model group is independent management.
29. The method according to any one of claims 16 to 28, characterized in that, The identification information of the target model group is sent through a request message or a configuration message.
30. The method according to any one of claims 16 to 29, characterized in that The method further includes: Send the identification information of each model in the target model group.
31. A communication device, characterized in that, Includes: A module for performing the method according to any one of claims 1 to 15, or a module for performing the method according to any one of claims 16 to 30.
32. A communication device, characterized in that, Comprising: A processor, the processor being coupled to a memory, the memory storing computer-executable instructions, the processor executing the computer-executable instructions stored in the memory such that the processor performs the method according to any one of claims 1 to 15, or performs the method according to any one of claims 16 to 30.
33. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program comprising instructions for implementing the method according to any one of claims 1 to 15, or instructions for performing the method according to any one of claims 16 to 30.
34. A computer program product, characterized in that, The computer program product includes computer program code which, when run on a computer, causes the computer to implement the method according to any one of claims 1 to 15, or perform the method according to any one of claims 16 to 30.
35. A communication system, characterized in that, Comprising: A device for performing the method according to any one of claims 1 to 15, and / or a device for performing the method according to any one of claims 16 to 30.
36. A communication method, characterized in that, Applied to a communication system comprising a first device and a second device, the method comprising: performing the method according to any one of claims 1 to 15 by the second device, and performing the method according to any one of claims 16 to 30 by the first device.
Citation Information
Cited By
Model deployment method and communication apparatus
EP4815388A1
Model deployment method and communication apparatus
WO2025130683A1