Model transmission method and device, storage medium and program product
By using model deployment request and response messages between network devices and terminals, the problem of unsuccessful model deployment after transmission is solved, and successful model deployment and availability control on the terminal are achieved.
Patent Information
- Application Number
- CN202410686814.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-29
- Publication Date
- 2025-12-02
AI Technical Summary
In existing technologies, successful deployment on the receiving device cannot be guaranteed after model transmission, resulting in low availability of model transmission.
By receiving and sending model deployment request messages, network devices and terminals confirm and control the model deployment status, ensuring that the model is successfully deployed on the terminal.
This improves the availability of model transmission, enabling network devices to know whether the model has been successfully deployed on the terminal, and achieves unified scheduling and control of terminal model deployment.
Smart Images

Figure CN121056293A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of communication technology, and in particular to a model transmission method, apparatus, storage medium, and program product. Background Technology
[0002] In the field of communication technology, in order to deploy models on different devices or to integrate with other systems or applications, it is sometimes necessary to transfer models from one device or environment to another.
[0003] In related technologies, the model to be transmitted is usually sent to the receiving device as a whole. However, this method can only achieve the transmission of the model and cannot guarantee that the transmitted model will be usable after being deployed on the receiving device, resulting in low availability of model transmission. Summary of the Invention
[0004] This disclosure provides a model transmission method, apparatus, storage medium, and program product, which can control the deployment of models on terminals through network devices, making the models to be deployed on the terminals available and improving the availability of model transmission.
[0005] On the one hand, a model transmission method is provided for application in a terminal, including:
[0006] Receive model deployment request messages sent by network devices. The model deployment request messages include model information of the model to be deployed and model inference requirements.
[0007] Based on the model deployment request message, a first model deployment response message is sent to the network device. The first model deployment response message is used to indicate whether the model to be deployed has been successfully deployed on the terminal.
[0008] On the other hand, a model transmission method is provided for application in network devices, including:
[0009] Send a model deployment request message to the terminal. The model deployment request message includes the model information of the model to be deployed and the model inference requirements.
[0010] The system receives a first model deployment response message from the terminal, which indicates whether the model to be deployed has been successfully deployed on the terminal.
[0011] On the other hand, a model transmission method is provided for use on a server, including:
[0012] Receive a model deployment request message, which includes model information of the model to be deployed and model inference requirements;
[0013] Based on the model deployment request message, a second model deployment response message is sent, which indicates whether the model to be deployed can be successfully deployed on the terminal.
[0014] In another aspect, a communication device is provided, comprising: a processor and a memory for storing processor-executable instructions; the processor is configured to execute the instructions such that the communication device implements any of the model transmission methods provided in the embodiments of this disclosure.
[0015] In another aspect, a computer-readable storage medium is provided, on which computer program instructions are stored, which, when executed on a computer, cause the computer to implement any of the model transmission methods provided in the embodiments of this disclosure.
[0016] In another aspect, a computer program product is provided, which includes computer program instructions that, when executed on a computer, cause the computer to implement any of the model transfer methods provided in the embodiments of this disclosure.
[0017] The model transmission method provided in this disclosure has several advantages. First, it enables network devices to control model deployment on terminals via model inference requests sent by network devices, facilitating unified scheduling of model deployment on terminals. Second, it allows network devices to know whether a model to be deployed has been successfully deployed on a terminal. Third, since the successful deployment of a model on a terminal is related to the model inference requests and model information sent by the network devices, models that can be successfully deployed on terminals are usable. Attached Figure Description
[0018] Figure 1 This is one of the structural diagrams of a communication system according to some embodiments;
[0019] Figure 2 This is a second structural diagram of a communication system according to some embodiments;
[0020] Figure 3 This is one of the flowcharts for a model transfer method according to some embodiments;
[0021] Figure 4 This is a second flowchart of a model transmission method according to some embodiments;
[0022] Figure 5 This is a flowchart of a model transmission method according to some embodiments;
[0023] Figure 6 This is a flowchart of a model transmission method according to some embodiments;
[0024] Figure 7 This is the fifth flowchart of a model transmission method according to some embodiments;
[0025] Figure 8 This is a flowchart of a model transmission method according to some embodiments;
[0026] Figure 9 This is the seventh flowchart of a model transmission method according to some embodiments;
[0027] Figure 10 This is the eighth flowchart of a model transmission method according to some embodiments;
[0028] Figure 11 This is the ninth flowchart of a model transmission method according to some embodiments;
[0029] Figure 12 This is flowchart ten of a model transmission method according to some embodiments;
[0030] Figure 13 This is a structural diagram of a terminal according to some embodiments;
[0031] Figure 14 This is a structural diagram of a network device according to some embodiments;
[0032] Figure 15 This is a structural diagram of a server according to some embodiments;
[0033] Figure 16 This is a structural diagram of a communication device according to some embodiments. Detailed Implementation
[0034] The technical solutions in the embodiments of this disclosure will now be clearly and completely described with reference to the accompanying drawings.
[0035] In the description of this disclosure, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone.
[0036] Furthermore, "at least one" refers to one or more, and "more than one" refers to two or more. To facilitate a clear description of the technical solutions of the embodiments of this disclosure, the terms "first" and "second" are used in the embodiments of this disclosure to distinguish identical or similar items with substantially the same function and effect. It should be understood that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.
[0037] Furthermore, in this disclosure, the words "exemplarily" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplarily" or "for example" in this disclosure should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.
[0038] In the field of communication technology, in order to deploy models on different devices or to integrate with other systems or applications, it is sometimes necessary to transfer models from one device or environment to another.
[0039] In related technologies, the model to be transmitted is usually sent to the receiving device as a whole. However, this method can only achieve the transmission of the model and cannot guarantee that the transmitted model will be usable after being deployed on the receiving device, resulting in low reliability of model transmission.
[0040] Taking artificial intelligence / machine learning (AI / ML) scenarios as an example, AI / ML technology is a technique that uses computers to simulate and realize human intelligence, and it has been widely used in core networks, network management and optimization, and access networks. In particular, air interface transmission technology based on AI / ML has also made significant progress in recent years. For example, by enhancing air interface characteristics to support AI / ML algorithms, communication performance can be improved, system complexity reduced, or communication overhead lowered. Specifically, models based on AI / ML technology can be applied in at least one of the following three scenarios:
[0041] (1) Feedback of channel state information (CSI);
[0042] For example, it can be applied to compressed feedback of CSI and / or time-domain prediction of CSI.
[0043] (2) Beam management (BM);
[0044] For example, it can be applied to beam prediction that includes both time and space.
[0045] (3) Positioning;
[0046] For example, it can be applied to AI / ML localization and / or AI / ML assisted localization.
[0047] Among them, the output of model inference based on AI / ML technology can be new measurements and / or enhancements to existing measurements.
[0048] Typically, AI / ML models can be deployed at one end of a communication link (e.g., on a network device or terminal) or at both ends of the communication link simultaneously. In such cases, it is sometimes necessary to transfer the model from one device or environment to another.
[0049] Taking CSI compression feedback as an example, it is necessary to deploy AI / ML models on both the terminal side and network devices simultaneously, that is, to deploy a two-sided model. The models on the terminal and network devices are used to implement CSI compression and decompression, respectively.
[0050] For this two-sided model, one possible implementation in related technologies is that the network device directly sends the trained model to the terminal via a model transfer mechanism. However, this approach typically only achieves model transfer and cannot guarantee that the model sent by the network device will be usable after deployment on the terminal. Related technologies have not yet proposed a corresponding solution for model transfer.
[0051] To address this issue, this disclosure provides a model transmission method. The method includes: receiving a model deployment request message sent by a network device, the model deployment request message including model information of the model to be deployed and model inference requirements; and based on the model deployment request message, sending a first model deployment response message to the network device, the first model deployment response message indicating whether the model to be deployed has been successfully deployed on the terminal. This allows for control over model deployment on the terminal through the network device, ensuring that models successfully deployed on the terminal are available, thus improving the availability of model transmission.
[0052] The model transmission method provided in this disclosure can be applied to systems with various communication standards. For example, applicable systems include, but are not limited to, Long Term Evolution (LTE) systems, various versions based on LTE evolution, 5th generation (5G) systems, New Radio (NR) systems, 5G NR systems, 5th generation new radio-advanced (5G-advanced) systems, and 6th generation (6G) systems, among other next-generation communication systems. Furthermore, the model transmission method provided in this disclosure can also be applied to future-oriented communication technologies.
[0053] To more clearly illustrate the solution provided in this disclosure, Figure 1 The diagram illustrates a communication system related to this disclosure. (See reference...) Figure 1 The communication system includes network equipment 110 and terminal 120.
[0054] There is a communication connection between network device 110 and terminal 120.
[0055] Network device 110 can be a next-generation node B (gNB), a transceiver point (TRP), an evolved node B (eNB), a radio access point (AP), or an evolved node base station (eNB), a base station in a 5G network, or a base station in a 6G network, etc. Furthermore, network device 110 can also be understood as a telecommunications operator, etc. It should be understood that this disclosure does not specifically limit the scope of the embodiments.
[0056] In some embodiments, network device 110 may implement any of the model transmission methods provided in the embodiments of this disclosure. For example, network device 110 may send a model deployment request message to terminal 120 and receive a first model deployment response message sent by terminal 120; wherein, the model deployment request message includes model information of the model to be deployed and model inference requirements, and the first model deployment response message is used to indicate whether the model to be deployed has been successfully deployed on terminal 120.
[0057] Furthermore, if the first model deployment response message indicates that the model to be deployed has failed to be deployed on terminal 120, network device 110 may also resend an updated model deployment request message to terminal 120.
[0058] Terminal 120 may be user equipment (UE), handheld device with various communication functions, vehicle-mounted device, wearable device, computer, smart home device or smart office device, etc., and this disclosure does not specifically limit it.
[0059] In some embodiments, terminal 120 may implement any of the model transmission methods provided in the embodiments of this disclosure. For example, terminal 120 may receive a model deployment request message sent by a network device, and send first model deployment response information to network device 110 based on the model deployment request message, wherein the concepts of model deployment request message and first model deployment response information can be referred to the description above.
[0060] Furthermore, if the first model deployment response message indicates that the model to be deployed has failed to be deployed on terminal 120, an updated model deployment request message resent by network device 110 can also be received.
[0061] Figure 2 Another communication system related to this disclosure is shown. (Refer to...) Figure 2 The communication system includes network equipment 110, terminal 120 and server 130.
[0062] Among them, network device 110 and terminal 120 can refer to the above-mentioned... Figure 1 The description will not be elaborated here.
[0063] There is a communication connection between server 130 and terminal 120, and / or, there is a communication connection between server 130 and network device 110. Server 130 has at least one of the following functions: compiling model information; optimizing and / or lightweighting model information, such as retraining the model based on the model information; determining whether the model information can be successfully deployed on a certain terminal.
[0064] In a specific example, server 130 is associated with terminal 120 to provide services to terminal 120.
[0065] In a specific example, server 130 can be an OTT (Over-The-Top) server, which can also be referred to as an OTT service platform; this disclosure does not specifically limit this. OTT is a service method, typically referring to a service service built on top of a network provided by a network operator. These services do not depend on the operator's physical network but are provided by a third party (such as a terminal manufacturer). It should be understood that server 130 can also be other servers with the above-mentioned functions; this disclosure does not specifically limit this.
[0066] Taking server 130 as an OTT server as an example, then Figure 2 The communication system shown can be applied to Internet services such as OTT / IPTV (Internet Protocol Television), enabling users to watch videos, download applications, browse web pages, and receive services on various terminal devices.
[0067] In a specific example, when server 130 is an OTT server: network device 110 can enable communication between terminal 120 and server 130; terminal 120 can directly face users and provide mainstream OTT services, such as providing mainstream OTT services including video, application download, web browsing, and service applications; server 130 is an OTT server and can carry specific services, such as providing support for various services provided on terminal 120.
[0068] It should be understood that the examples of the communication systems described above are merely for illustrating the technical solutions of this disclosure more clearly and do not constitute a limitation of this disclosure. Those skilled in the art will recognize that, with the evolution of network architecture and the emergence of new service scenarios, the technical solutions provided in this disclosure are equally applicable to similar technical problems.
[0069] To illustrate the solution more clearly, the model transmission method provided in this disclosure is described below with reference to the accompanying drawings. It should be noted that the various embodiments of this disclosure can be mutually referenced or understood; for example, identical or similar steps, method embodiments, and device embodiments can be mutually referenced without limitation.
[0070] like Figure 3 As shown, this disclosure provides a model transmission method applied to a terminal, the method comprising:
[0071] S101. The terminal receives a model deployment request message sent by the network device. The model deployment request message includes model information of the model to be deployed and model inference requirements.
[0072] Among them, model inference requirements can also be called model application requirements, which are used to indicate the requirements that the model to be deployed needs to meet in actual operation.
[0073] In some embodiments, model inference requirements include at least one of the following: inference latency requirements, inference sample requirements, and inference frequency requirements. Inference latency requirements indicate the maximum allowed time for a single model inference operation, and include at least one of the following: one or more microseconds, one or more milliseconds, one or more time-domain symbols, one or more time slots, and one or more radio frames, etc. Inference sample requirements include at least the number of samples required for a single model inference operation. Inference frequency requirements include at least the number of model inferences per unit time. Based on this, network devices can control the model deployment of terminals through model inference requirements.
[0074] In some embodiments, the model inference requirement is determined by the network device. For example, the model inference requirement is determined based on the actual model inference requirement of a first model, which is a model already deployed on the network device corresponding to the model to be deployed.
[0075] In a specific example, the model to be deployed is a CSI compressed model, which corresponds to a CSI decompressed model already deployed on the network device. Furthermore, the model inference requirements are determined by the network device based on the actual model inference requirements of the CSI decompressed model. Therefore, the network device can control the model deployment on the terminal through these model inference requirements.
[0076] In some embodiments, the model deployment request message also includes the priority of the model to be deployed.
[0077] As an example, the model deployment request message only includes the model priority of the model to be deployed, and the terminal stores the priority of at least one deployed model.
[0078] As another example, the model deployment request message also includes the priority of at least one deployed model on the terminal.
[0079] In some embodiments, the model information includes at least one of the following: model parameters, model structure, preprocessing information of the model input data, and model download address. It should be noted that the method for transmitting the model information can be referred to the description in the method embodiments below, and will not be repeated here.
[0080] In one example, the model information includes the model download address. It should be understood that the model download address contains less data than the model structure or model parameters; therefore, using the model download address as model information is beneficial for improving the model transmission rate.
[0081] In one example, the model information includes preprocessing information for the model input data. This preprocessing information includes at least one preprocessing type and specific processing methods corresponding to each preprocessing type. The preprocessing type indicates the type of preprocessing operation required for the model input data, including but not limited to at least one of the following: data normalization, data scaling, and data biasing. Specifically, data normalization normalizes the maximum value of the modulus of the model input data so that the modulus of the element with the maximum modulus is 1 after normalization. Data scaling multiplies the normalized data by a scaling factor S; in a specific example, the scaling factor S is 1 / 2. Data biasing adds a bias factor O to the scaled data; in a specific example, O is 1 / 2. It should be understood that the above preprocessing types are merely examples, and other preprocessing types, such as data cleaning, may exist. This disclosure does not impose specific limitations on these types.
[0082] In some embodiments, it should be noted that the model information of the model to be deployed can also be understood as a data processing method or an input-output mapping relationship, etc. Correspondingly, when the model information of the model to be deployed is understood as a data processing method, the model structure of the model to be deployed can be understood as a set of parameters used by that data processing method.
[0083] In some embodiments, the model information of the model to be deployed includes model information of one or more models to be deployed. Where the model information of the model to be deployed includes model information of multiple models to be deployed, the model inference requirement includes model inference requirements of multiple models to be deployed, and the model information of each model to be deployed corresponds to one model inference requirement of a model to be deployed.
[0084] In some embodiments, the model information of the model to be deployed is model information compiled by the server; or, the model information of the model to be deployed is model information not compiled by the server.
[0085] The model information that has not been compiled by the server can also be referred to as open-format model information, or first-format model information, etc., and this disclosure does not specifically limit the name of this concept. In a specific example, the model information that has not been compiled by the server includes the initial model information of the model to be deployed generated by the network device.
[0086] Model information compiled by the server can also be referred to as proprietary format model information, or second format model information, etc., and this disclosure does not specifically limit the name of this concept. It should be noted that, compared with uncompiled model information, server-compiled model information is more suitable for deployment on the terminal. For example, server-compiled model information may be optimized and / or lightweight model information. Furthermore, server-compiled model information may be model information specifically optimized for the terminal's AI / ML accelerator hardware platform. Therefore, server-compiled model information can be better adapted to the terminal's hardware platform, thereby improving the performance of model inference or model execution on the terminal.
[0087] S102. Based on the model deployment request message, the terminal sends a first model deployment response message to the network device. The first model deployment response message is used to indicate whether the model to be deployed has been successfully deployed on the terminal.
[0088] In some embodiments, the first model deployment response message indicates that the model to be deployed has been successfully deployed on the terminal. Further, in response to the first model deployment response message indicating that the model to be deployed has been successfully deployed on the terminal, the first model deployment response message may include the activation time of the model to be deployed.
[0089] In a specific example, the activation time of the model to be deployed includes the earliest time slot and / or the index of the time slot at which the model can be activated on the terminal. For example, the activation time of the model to be deployed includes the index of the Nth time slot after the time slot in which the terminal sends the first model deployment response information, where N is a natural number; that is, the terminal can activate the model to be deployed in the Nth time slot after the time slot in which the terminal sends the first model deployment response information. Correspondingly, the network device can determine whether to activate the model to be deployed based on the activation time of the model to be deployed, and / or, the network device can determine the specific time slot for activating the model to be deployed based on the activation time of the model to be deployed. It should be noted that the above-mentioned form of activation time is only an example. In addition to the form of time slot or time slot index, the above-mentioned form of activation time can also be time domain symbol, microsecond, etc., and this disclosure does not specifically limit it in this way.
[0090] Furthermore, in response to the first model deployment response message indicating that the model to be deployed has been successfully deployed on the terminal, the first model deployment response message may also include the model's activation duration and / or deactivation time, etc. The model deactivation time may also be referred to as the model end time, etc., and this disclosure does not impose specific limitations on it.
[0091] It should be noted that the first model deployment response message may not include the model's activation duration and deactivation time. For example, after the model to be deployed is successfully deployed on the terminal, the terminal receives model operation information sent by the network device; based on this model operation information, the terminal runs the model to be deployed; this model operation information includes: the model's activation duration and / or deactivation time. In this way, the network device can control the model operation on the terminal.
[0092] Based on this, network devices can know the model activation time of the model to be deployed. Especially when the network device has also deployed other models corresponding to the model to be deployed, the other models of the network device and the model to be deployed on the terminal can be activated at the same time, so as to achieve model alignment between the other models of the network device and the model to be deployed, thereby achieving accurate deployment of the two-sided model.
[0093] In other embodiments, the first model deployment response message indicates that the model to be deployed failed to deploy on the terminal. Further, in response to the first model deployment response message indicating that the model to be deployed failed to deploy on the terminal, the first model deployment response message may include a reason for the failure.
[0094] In a specific example, the failure reason may include at least one of the following: the terminal does not support the deployment of the model to be deployed; the model to be deployed does not meet the model inference requirements when running on the terminal; the terminal does not support the model structure of the model to be deployed; the terminal has insufficient computing power; the terminal has insufficient memory; and other implementation-related reasons that prevent the model to be deployed from being successfully deployed on the terminal. It should be understood that the above-mentioned first model deployment response message may also not include a failure reason, and this disclosure does not specifically limit this.
[0095] It should be noted that the specific implementation of step S102 can be found in the description in the method embodiments below, and will not be detailed here.
[0096] The model transmission method provided in this disclosure has several advantages. First, it enables network devices to control model deployment on terminals via model inference requests sent by network devices, facilitating unified scheduling of model deployment on terminals. Second, it allows network devices to know whether the model to be deployed has been successfully deployed on the terminal. Third, since the successful deployment of a model on a terminal is related to the model inference requests and model information sent by the network devices, models that can be successfully deployed on terminals are available.
[0097] In some embodiments, step S102 includes: the terminal determining the deployment result of the model to be deployed based on the model deployment request message; and the terminal sending a first model deployment response message to the network device based on the deployment result.
[0098] In this way, it is possible to determine whether the model to be deployed is suitable for deployment on the terminal based on the model deployment request message sent by the network device, and to inform the network device of the deployment status of the model to be deployed through the first model deployment response information.
[0099] In one example, the above determination of the deployment result of the model to be deployed based on the model deployment request message includes the following steps Sa1 to Sa2:
[0100] Sa1: Based on the model information of the model to be deployed, the terminal determines whether the terminal supports the deployment of the model to be deployed, and whether the model to be deployed meets the model inference requirements when running on the terminal.
[0101] For example, step Sa1 includes: the terminal deploys the model to be deployed on the terminal based on the model information of the model to be deployed, determines whether the terminal supports the deployment of the model to be deployed, and determines whether the model to be deployed meets the model inference requirements when running on the terminal.
[0102] For example, step Sa1 includes: the terminal determines whether it supports deploying the model based on the model information of the model to be deployed; in response to the terminal supporting the deployment of the model to be deployed, the terminal deploys the model to be deployed on the terminal based on the model information of the model to be deployed, and determines whether the model to be deployed meets the model inference requirements when running on the terminal.
[0103] In a specific example, deploying the model to be deployed on the terminal based on the model information of the model to be deployed includes: the terminal deploying the model to be deployed in the actual operating environment of the terminal based on the model information of the model to be deployed; or, the terminal deploying the model to be deployed in the simulated operating environment of the terminal based on the model information of the model to be deployed. The actual operating environment may also be referred to as the physical operating environment, etc., and the simulated operating environment may also be referred to as the virtual operating environment, etc., and this disclosure does not specifically limit it in this way.
[0104] Sa2. In response to the terminal supporting the deployment of the model to be deployed and the model to be deployed meeting the model inference requirements when running on the terminal, the terminal determines that the model to be deployed has been successfully deployed on the terminal; or, in response to the terminal not supporting the deployment of the model to be deployed and / or the model to be deployed not meeting the model inference requirements when running on the terminal, the terminal determines that the model to be deployed has failed to be deployed on the terminal.
[0105] Based on this, it is possible to ensure that a model successfully deployed on a terminal is supported by the terminal and meets the model inference requirements when running on the terminal; that is, to ensure that a model successfully deployed on a terminal is usable. Furthermore, the network device can control the model deployment on the terminal through model inference requests sent by the network device.
[0106] In some embodiments, the model deployment request message also includes the priority of the model to be deployed; before step S102, the following steps are also included:
[0107] Sb1. The terminal determines whether a reference model exists in the terminal based on the priority of the model to be deployed. If the reference model has been deployed in the terminal and its priority is lower than that of the model to be deployed.
[0108] Sb2, In response to the existence of a reference model in the terminal, the terminal disables the reference model.
[0109] In one example, the terminal disables the reference model in the terminal's actual operating environment.
[0110] In another example, the terminal disables the reference model on the terminal's simulated runtime environment.
[0111] Furthermore, in response to the terminal disabling the reference model, the terminal executes step S102; or, in response to the absence of a reference model in the terminal, the terminal directly executes step S102.
[0112] Based on this, if the priority of the model to be deployed is high, the reference model that has already been deployed in the terminal and has a lower priority can be deactivated first, and then the deployment result of the model to be deployed can be determined, so as to ensure that the terminal's resources are given priority to the model to be deployed with higher priority.
[0113] In some embodiments, the model information of the model to be deployed is model information compiled by the server; or, the model information of the model to be deployed is model information not compiled by the server. The concepts of server-compiled model information and server-uncompiled model information can be found in the description of step S101, and will not be repeated here.
[0114] In a specific example, the server is an OTT server; that is, the model information of the model to be deployed is model information compiled by the OTT server; or, the model information of the model to be deployed is model information that has not been compiled by the OTT server.
[0115] In some embodiments, the model information of the model to be deployed is model information that has not been compiled by the server, based on... Figure 3 The illustrated embodiments, such as Figure 4 As shown, after step S101, the model transfer method further includes the following steps:
[0116] S103. The terminal sends the model information of the model to be deployed to the server.
[0117] S104. The terminal receives the target model information sent by the server.
[0118] The target model information is the model information compiled by the server, and the target model information is obtained by the server compiling the model information of the model to be deployed.
[0119] In one example, before step S104, the terminal also sends a model inference request to the server so that the server can determine whether the model to be deployed can be successfully deployed on the terminal; step S104 includes: if the server determines that the model to be deployed can be successfully deployed on the terminal, the terminal receives the target model information sent by the server.
[0120] S105. The terminal updates the model information to be deployed with the target model information.
[0121] Based on this, on the one hand, the terminal can determine whether the model to be deployed has been successfully deployed on the terminal based on the model information compiled by the server; since the model information compiled by the server is more applicable to the terminal, the success rate of model deployment can be improved. On the other hand, since the terminal itself can find the corresponding server to compile the model information, there is no need for the network device to interact with the server, reducing the workload of the network device; in addition, since there is no need for the network device to interact with the server, the terminal does not need to inform the network device of the address of the corresponding server, saving communication resources.
[0122] In some embodiments, the model information of the model to be deployed is model information compiled by the server, based on... Figure 3 The illustrated embodiments, such as Figure 5 As shown, before step S101, the procedure further includes:
[0123] S106. The terminal sends the address of the server corresponding to the terminal to the network device. The server is used to generate model information for the model to be deployed.
[0124] Based on this, the terminal can inform the network device of the address of the server corresponding to the terminal, so that the network device can find the corresponding server; thereby, it can assist the network device in obtaining the model information of the model to be deployed, which has been compiled by the server.
[0125] In some embodiments, the terminal may omit step S106. For example, if the terminal has already connected to the network device, the network device already stores the address of the server corresponding to the terminal. Therefore, even if the terminal does not disclose the address of the server corresponding to the terminal to the network device, the network device can accurately locate the corresponding server.
[0126] In some embodiments, based on Figure 3 The illustrated embodiments, such as Figure 6 As shown, step S102 includes:
[0127] S1021. The terminal sends a model deployment request message to the server.
[0128] The model deployment request message can be referred to in the description above.
[0129] S1022. The terminal receives the second model deployment response information sent by the server. The second model deployment response message is used to indicate whether the model to be deployed can be successfully deployed on the terminal.
[0130] S1023. The terminal sends the first model deployment response message to the network device based on the second model deployment response information.
[0131] In one example, step S1023 includes: in response to a second model deployment response message indicating that the model to be deployed cannot be successfully deployed on the terminal, the terminal sends a first model deployment response message to the network device, the first model deployment response message indicating that the model to be deployed has failed to be deployed on the terminal.
[0132] In another example, step S1023 includes: in response to a second model deployment response message indicating that the model to be deployed can be successfully deployed on the terminal, the terminal sends a first model deployment response message to the network device, the first model deployment response message indicating that the model to be deployed is successfully deployed on the terminal.
[0133] In another example, step S1023 includes: in response to the second model deployment response information indicating that the model to be deployed can be successfully deployed on the terminal, the terminal sends a first model deployment response message to the network device based on the model deployment request message. Based on this, the model to be deployed can be judged twice, improving the accuracy of the deployment judgment.
[0134] The model transmission method provided in this disclosure can send the model information and model inference requirements of the model to be deployed to the server, so that the server can also determine whether the model to be deployed can be successfully deployed on the terminal. On the one hand, this method can provide more possibilities or multiple guarantees for the model deployment judgment, thereby improving the accuracy of the model deployment judgment. On the other hand, it can also avoid consuming the relevant resources of the terminal device, thereby reducing power consumption.
[0135] In some embodiments, based on Figure 3 The illustrated embodiments, such as Figure 7 As shown, the procedure before step S101 also includes:
[0136] S107. The terminal sends its model processing capability to the network device. The model processing capability is used to assist in generating model information for the model to be deployed.
[0137] The specific details of the model processing capabilities and transmission methods can be found in the description below, and will not be elaborated here.
[0138] In a specific example, the model processing capacity refers to the terminal's idle model processing capacity, which can also be called the terminal's unused model processing capacity. Thus, since the model information of the model to be deployed is generated with the assistance of the terminal's idle model processing capacity, the deployment and runtime of the model on the terminal will not exceed the terminal's idle model processing capacity.
[0139] In another specific example, the model processing capacity is the entire model processing capacity of the terminal. Based on this, when generating model information for the model to be deployed, the already used model processing capacity can be disregarded, and instead, the entire model processing capacity of the terminal can be prioritized for the model to be deployed; furthermore, this ensures that the deployment and runtime of the model on the terminal will not exceed the terminal's entire model processing capacity.
[0140] For example, if the model to be deployed has a high priority, even if there are several reference models in the terminal with lower priority than the model to be deployed that occupy part of the model processing capacity, the model processing capacity occupied by the reference models can be disregarded, and all the model processing capacity of the terminal can be provided to the network device first to assist in generating the model information of the model to be deployed.
[0141] In a specific example, the terminal's model processing capability is the AI / ML model processing capability.
[0142] In a specific example, the terminal sends its model processing capabilities to the network device. Correspondingly, the network device performs a preliminary screening of model information based on these capabilities, obtaining the preliminary screening results, which include model information for the models to be deployed. This allows the network device to perform a preliminary screening of model information before obtaining the model information for deployment, improving the success rate of model deployment.
[0143] The model transmission method provided in this disclosure improves the success rate of model deployment because the model information of the model to be deployed is determined based on the model processing capability of the terminal. This ensures that the model to be deployed will not exceed the model processing capability of the terminal when it is deployed and run, thus improving the success rate of model deployment.
[0144] It should be noted that although step S107 provides the aforementioned beneficial effects, it is not a mandatory step. Even without step S107 in the model transmission method, the model transmission method provided by this disclosure can still control the deployment of the model on the terminal through the network device, enabling the network device to know whether the model to be deployed has been successfully deployed on the terminal, and that a successfully deployed model on the terminal is usable, thereby improving the availability of model transmission. Furthermore, omitting step S107 in the model transmission method can also reduce information interaction between the terminal and the network device, saving communication resources.
[0145] In some embodiments, the model processing capabilities of the terminal include at least one of the following: data type-related capability information, computing capability-related information, storage capability-related information, memory read / write capability information, model structure-related capability information, multi-instance processing capability, and processor type-related information.
[0146] The data type-related capability information includes at least the data types supported by the terminal. In a specific example, the data types supported by the terminal include, but are not limited to, at least one of the following: whether it supports single-precision floating-point numbers FP32, whether it supports half-precision floating-point numbers FP16, whether it supports Bfloat16 floating-point numbers BF16, and whether it supports 8-bit integer numbers INT8.
[0147] Computational capability information includes, but is not limited to, at least one of the following: computational capability corresponding to various data types, general computational capability, and hardware acceleration module computational capability. In a specific example, hardware acceleration module computational capability includes the computational capability of the module used for matrix calculations. In a specific example, the unit of computational capability information is trillion floating-point operations per second (TFLOPs) or trillion operations per second (TOPs).
[0148] Storage capacity information includes at least the amount of memory configured on the terminal. In a specific example, the unit of memory size is gigabytes (GB).
[0149] Memory read / write capability includes at least memory read / write bandwidth. In a specific example, the unit of memory read / write bandwidth is gigabytes per second (GB / s).
[0150] The capability information related to model structure includes, but is not limited to: the types of model structures supported by the terminal, and / or, the types of activation functions supported by the terminal.
[0151] Multi-instance capabilities include, but are not limited to: the number of instances the terminal can handle, and / or the instance memory size the terminal can support. In a specific example, the number of instances the terminal can handle is the maximum number of instances the terminal can process in parallel.
[0152] Processor type information includes, but is not limited to, information about at least one of the following components: central processing unit (CPU), graphics processing unit (GPU), neural network processing unit (NPU), and tensor processing unit (TPU).
[0153] It should be noted that the above-mentioned model processing capabilities are merely examples, and this disclosure does not impose any specific limitations on them.
[0154] In some embodiments, when the terminal's model processing capabilities include at least one of data type-related capability information, computing capability-related information, and memory read / write capability information, the aforementioned data type-related capability information, computing capability-related information, and memory read / write capability information can be transmitted based on a hierarchical reporting method.
[0155] For example, the terminal stores multiple preset capability levels, each capability level corresponding to a capability range. Correspondingly, the network device also stores the aforementioned multiple preset capability levels, each capability level corresponding to a capability range. This allows the network device to determine the capability range corresponding to the capability level sent by the terminal based on the multiple preset capability levels and the capability level sent by the terminal.
[0156] Taking conventional floating-point computing power as an example, as shown in Table 1 below, it can be divided into 5 levels of conventional floating-point computing power.
[0157] Table 1
[0158] Standard floating-point computing power level The capabilities of conventional floating-point calculations 0 Less than 10 TFLOPs 1 Greater than or equal to 10 TFLOPs and less than 20 TFLOPs 2 Greater than or equal to 20 TFLOPs and less than 30 TFLOPs 3 Greater than or equal to 30 TFLOPs and less than 40 TFLOPs 4 40 TFLOPs or more
[0159] In a specific example, if the terminal's conventional floating-point computing capability is less than 10 TFLOPs, then the terminal's conventional floating-point computing capability level is 0. Therefore, step S107 can be implemented as follows: sending the terminal's conventional floating-point computing capability level, which is 0, to the network device. This can be repeated for other cases.
[0160] It should be understood that compared to the transmission capacity range, the transmission capacity level can use fewer bits, thus increasing the communication rate and saving communication resources.
[0161] In some embodiments, the terminal's model processing capabilities include model structure-related capability information. In a specific example, where the model structure-related capability information includes the types of model structures or activation functions supported by the terminal, it can be reported using a bitmap, where each bit in the bitmap corresponds to a model structure type or activation function type. For example, a bit of 0 indicates that the terminal does not support it; a bit of 1 indicates that the terminal supports it. As another example, a bit of 0 indicates that the terminal supports it; a bit of 1 indicates that the terminal does not support it.
[0162] Correspondingly, network devices store bitmap meaning information so that they can determine the terminal's model structure-related capability information based on the bitmap meaning information and the bitmap sent by the terminal.
[0163] In a specific example, the type of model architecture includes, but is not limited to, at least one of the following: convolutional neural network (CNN), long short-term memory network (LSTM), multilayer perceptron (MLP), and Transformer network.
[0164] In a specific example, the activation function type includes, but is not limited to, at least one of the following: ReLU activation function, PReLU (parametric ReLU) activation function, sigmoid activation function, softmax activation function, and tanh activation function.
[0165] Taking the terminal's model processing capabilities, including the types of model structures supported by the terminal, as an example, Table 2 shows a bitmap used to indicate the types of model structures supported by the terminal.
[0166] Table 2
[0167] serial number meaning 0 Does it support MLP networks? 1 Does it support CNN networks? 2 Does it support LSTM networks? 3 Does it support Transformer networks? 4 reserve 5 reserve 6 reserve 7 reserve
[0168] Referring to Table 2, it can be seen that the length of this bitmap is 8 bits. Bit number 0 in the bitmap indicates whether the terminal supports MLP networks; bit number 1 indicates whether the terminal supports CNN networks; bit number 2 indicates whether the terminal supports LSTM networks; and bit number 3 indicates whether the terminal supports Transformer networks. Bits 4-7 in the bitmap are reserved bits and can be configured for other functions based on user needs; this disclosure does not specifically limit their use. For example, when supporting additional model structures is required, any of the reserved bits numbered 4-7 can be used, or a larger indication overhead can be introduced.
[0169] In a specific example, the bitmap sent by the terminal is 01100000, where a bit of 0 indicates that the terminal does not support it, and a bit of 1 indicates that the terminal supports it. The meaning of the bitmap 01100000 is: the terminal does not support MLP networks and Transformer networks, but the terminal supports CNN networks and LSTM networks.
[0170] The model transmission method provided in this disclosure can transmit terminal model structure-related capability information through a bitmap. It can be seen that this method can transmit model processing capability information using fewer bits, thus improving communication speed and saving communication resources.
[0171] In some embodiments, the model information includes model parameters and / or model structure. In the transmission of model information (e.g., a network device sends it to a terminal, or a terminal sends it to a server), the model information can be sent in any of the following ways: directly transmitting the source code of the model parameters and / or model structure; or, converting the source code of the model parameters and / or model structure into other formats before sending it.
[0172] In a specific example, this other format includes the Open Neural Network Exchange (ONNX) format, which is a standard for representing deep learning models and enabling model transfer between different frameworks.
[0173] In some embodiments, the model information includes preprocessing information of the model input data. For example, a number of bits can be used to indicate the preprocessing information of the input data. As a specific example, 3 bits can be used to indicate the preprocessing information of the input data. Table 3 shows, for instance, a form of indication information for indicating the preprocessing information of the model input data.
[0174] Table 3
[0175] serial number Indication information for preprocessing model input data 0 No pretreatment 1 Data normalization processing 2 Data normalization and data scaling 3 Data normalization + data scaling + data bias processing 4 reserve 5 reserve 6 reserve 7 reserve
[0176] Referring to Table 3, when the indication information is numbered 0, it indicates that the model input data has no preprocessing; when the indication information is numbered 1, it indicates that the preprocessing of the model input data includes data normalization; when the indication information is numbered 2, it indicates that the preprocessing of the model input data includes data normalization and data scaling; and when the indication information is numbered 3, it indicates that the preprocessing of the model input data includes data normalization, data scaling, and data biasing. Bits numbered 4-7 in this bitmap are reserved bits and can be configured for other functions based on user needs. This disclosure does not specifically limit their use. For example, when additional preprocessing information (e.g., additional preprocessing types and corresponding specific processing methods) is required, any of the reserved bits numbered 4-7 can be used, or greater indication overhead can be introduced.
[0177] The model transmission method provided in this disclosure can transmit model information using a number of bits of indication information. It can be seen that this method can use fewer bits, thus increasing communication speed and saving communication resources.
[0178] As can be seen, the above description mainly focuses on the model transmission method provided in this disclosure from the perspective of the terminal side. This disclosure also provides a model transmission method applied to a network device. It should be understood that the relevant content of the model transmission method applied to a network device can be referred to in conjunction with the model transmission methods applied to terminals and servers provided in this disclosure.
[0179] like Figure 8 As shown in the embodiments of this disclosure, a model transmission method is provided and applied to a network device. The method includes:
[0180] S201. The network device sends a model deployment request message to the terminal. The model deployment request message includes the model information of the model to be deployed and the model inference requirements.
[0181] S202. The network device receives a first model deployment response message sent by the terminal. The first model deployment response message is used to indicate whether the model to be deployed has been successfully deployed on the terminal.
[0182] The details of the model deployment request message and the first model deployment response message can be found above and will not be repeated here.
[0183] In some embodiments, model inference requirements include at least one of the following: inference latency requirements, inference sample requirements, and inference frequency requirements.
[0184] In some embodiments, the model deployment request message may also include the priority of the model to be deployed.
[0185] In some embodiments, in response to a first model deployment response message indicating that the model to be deployed has been successfully deployed on the terminal, the first model deployment response message includes the activation time of the model to be deployed; or, in response to a first model deployment response message indicating that the model to be deployed has failed to be deployed on the terminal, the first model deployment response message includes the reason for the failure.
[0186] In some embodiments, in response to a first model deployment response message indicating that the model to be deployed has failed to be deployed on the terminal, the network device updates the model deployment request message, obtains the updated model deployment request message, and resends the updated model deployment request message to the terminal.
[0187] In some embodiments, the model information of the model to be deployed is model information compiled by the server; or, the model information of the model to be deployed is model information that has not been compiled by the server. The concepts of server-compiled model information and server-uncompiled model information are as described above and will not be repeated here.
[0188] The model transmission method provided in this disclosure has several advantages. First, it enables network devices to control model deployment on terminals via model inference requests sent by network devices, facilitating unified scheduling of model deployment on terminals. Second, it allows network devices to know whether a model to be deployed has been successfully deployed on a terminal. Third, since the successful deployment of a model on a terminal is related to the model inference requests and model information sent by the network devices, models that can be successfully deployed on terminals are usable.
[0189] In some embodiments, the model information of the model to be deployed is model information compiled by the server. Before sending the model deployment request message to the terminal, the method further includes: the network device generating initial model information of the model to be deployed, wherein the initial model information of the model to be deployed is model information that has not been compiled by the server; the network device sending the initial model information of the model to be deployed to the server corresponding to the terminal; the network device receiving target model information sent by the server; wherein the target model information is model information compiled by the server, and the target model information is obtained by the server compiling the initial model information of the model to be deployed; the network device determining the target model information as the model information of the model to be deployed.
[0190] Based on this, on the one hand, network devices can transform the generated initial model information into server-compiled model information, so that the model information to be deployed sent by the network device to the terminal is server-compiled model information; since server-compiled model information is better suited for the terminal, the success rate of model deployment can be improved. On the other hand, it reduces the workload of the terminal, as the terminal does not need to go to the server to compile the model information.
[0191] Furthermore, the network device can also receive the address of the server corresponding to the terminal sent by the terminal, and the server uses this address to generate model information for the model to be deployed. The network device sending the initial model information for the model to be deployed to the server corresponding to the terminal includes: the network device sending the initial model information for the model to be deployed to the server corresponding to the terminal based on the address of the server corresponding to the terminal.
[0192] Based on this, the network device can obtain the address of the server corresponding to the terminal, thereby enabling the compilation of model information for the model to be deployed. It should be understood that in some cases, the network device itself does not know the server corresponding to the terminal. For example, if the terminal has signed a contract with a server vendor, the network device is unaware of the server's address. Therefore, the network device can obtain the address of the server corresponding to the terminal by receiving the address sent by the terminal. In other cases, the network device itself knows the terminal's server. For example, if the terminal has already connected to the network device, the network device already stores the address of the server corresponding to the terminal. Therefore, the steps related to receiving the server address described above may not be necessary.
[0193] In some embodiments, the network device also sends a model deployment request message to the server, the model deployment request information including model information of the model to be deployed and model inference requirements; the network device receives a second model deployment response message sent by the server, the second model deployment response message being used to indicate whether the model to be deployed can be successfully deployed on the terminal.
[0194] In one example, if the second model deployment response message indicates that the model to be deployed cannot be successfully deployed on the terminal, the network device can regenerate the first model information and update the model information of the model to be deployed with the first model information; furthermore, the network device can resend the model deployment request message to the server and receive the second model deployment response information sent by the server.
[0195] In another example, the second model deployment response message indicates that the model to be deployed can be successfully deployed on the terminal; further, step S201 includes: when the second model deployment response message indicates that the model to be deployed can be successfully deployed on the terminal, the network device sends a model deployment request message to the terminal.
[0196] Based on this, the model information and inference requirements of the model to be deployed can be sent to the server, so that the server can also determine whether the model to be deployed can be successfully deployed on the terminal. This approach can provide more possibilities or multiple guarantees for the model deployment judgment, thereby improving the accuracy of the model deployment judgment. On the other hand, it can also avoid consuming the relevant resources of the terminal device, thereby reducing power consumption.
[0197] In some embodiments, the network device may also receive the model processing capability of the terminal sent by the terminal, and determine the model information of the model to be deployed based on the model processing capability of the terminal; the model processing capability is used to assist in generating the model information of the model to be deployed.
[0198] In a specific example, the network device receives the model processing capability of the terminal sent by the terminal; the network device generates initial model information of the model to be deployed based on the model processing capability of the terminal, the initial model information being model information that has not been compiled by the server; the network device sends the initial model information of the model to be deployed to the server corresponding to the terminal; the network device receives target model information sent by the server, the target model information being model information compiled by the server; the network device determines the target model information as the model information of the model to be deployed.
[0199] In another specific example, the network device generates initial model information of the model to be deployed; the network device sends the initial model information of the model to be deployed to the server corresponding to the terminal; the network device receives the target model information sent by the server; the network device determines the target model information as the model information of the model to be deployed; the network device determines whether the model to be deployed can be successfully deployed on the terminal based on the model processing capability and model inference requirements of the terminal; step S201 includes: in response to the network device determining that the model to be deployed can be successfully deployed on the terminal, the network device sends a model deployment request message to the terminal.
[0200] The model transmission method provided in this embodiment ensures that the model information of the model to be deployed is determined based on the model processing capability of the terminal, so that the model to be deployed will not exceed the model processing capability of the terminal when it is deployed and run. In addition, it also allows the network device to make a determination on the deployment status, thus improving the success rate of model deployment.
[0201] As can be seen, the above description mainly focuses on the model transmission method provided in this disclosure from the perspectives of the terminal side and the network device side. This disclosure also provides a model transmission method applied to a server. It should be understood that the relevant content of the model transmission method applied to the server can be referred to in conjunction with the model transmission methods applied to the terminal and the model transmission methods applied to the network device provided in this disclosure.
[0202] like Figure 9As shown, this disclosure provides a model transmission method applied to a server, the method comprising:
[0203] S301. The server receives a model deployment request message, which includes model information of the model to be deployed and model inference requirements.
[0204] In some embodiments, the model deployment request message may also include the priority of the model to be deployed.
[0205] In some embodiments, the server is an OTT server.
[0206] In some embodiments, the model information of the model to be deployed is model information that has not been compiled by the server, such as the initial model information of the model to be deployed; or model information that has been compiled by the server.
[0207] S302. Based on the model deployment request message, the server sends a second model deployment response message, which indicates whether the model to be deployed can be successfully deployed on the terminal.
[0208] In some embodiments, the server receives deployment request information sent by the terminal; based on the model deployment request message, it sends a second model deployment response message to the terminal.
[0209] In other embodiments, the server receives deployment request information sent by the network device; based on the model deployment request message, it sends a second model deployment response message to the network device.
[0210] The model transmission method provided in this disclosure can send the model information and model inference requirements of the model to be deployed to the server, so that the server can also determine whether the model to be deployed can be successfully deployed on the terminal. On the one hand, this method can provide more possibilities or multiple guarantees for the model deployment judgment, thereby improving the accuracy of the model deployment judgment. On the other hand, it can also avoid consuming the relevant resources of the terminal device, thereby reducing power consumption.
[0211] In some embodiments, the server sends a second model deployment response message based on the model deployment request message, including: the server determining the deployment result of the model to be deployed based on the model deployment request message; and the server sending the second model deployment response message based on the deployment result.
[0212] In one example, the server determines the deployment result of the model to be deployed based on the model deployment request message, including the following steps:
[0213] Sc1: Based on the model information of the model to be deployed, the server determines whether the terminal supports the deployment of the model to be deployed, and whether the model to be deployed meets the model inference requirements when running on the terminal.
[0214] For example, step Sc1 includes: the server deploys the model to be deployed on the terminal's simulated runtime environment based on the model information of the model to be deployed; determines whether the terminal's simulated runtime environment supports the deployment of the model to be deployed; and determines whether the model to be deployed meets the model inference requirements when running on the terminal's simulated runtime environment.
[0215] For example, step Sc1 includes: the server determines whether the terminal's simulation environment supports the deployment of the model based on the model information of the model to be deployed; in response to the terminal's simulation environment supporting the deployment of the model, the server deploys the model to be deployed in the terminal's simulation environment based on the model information of the model to be deployed, and determines whether the model to be deployed meets the model inference requirements when running in the terminal's simulation environment.
[0216] It should be understood that the server can also determine the simulated operating environment of the terminal based on the terminal's operating status information, so as to determine whether the terminal supports the deployment of the model to be deployed, and whether the model to be deployed meets the model inference requirements when running on the terminal.
[0217] The model information of the model to be deployed is model information that has not been compiled by the server, such as the initial model information of the model to be deployed, or the model information of the model to be deployed is model information that has not been compiled by the server.
[0218] In a specific example, the model information of the model to be deployed is the model information compiled by the server; step Sc1 includes: the server directly determines whether the terminal supports the deployment of the model to be deployed based on the model information of the model to be deployed, and determines whether the model to be deployed meets the model inference requirements when running on the terminal.
[0219] In another specific example, the model information of the model to be deployed is model information that has not been compiled by the server; step Sc1 includes: the server compiles the model information of the model to be deployed to obtain target model information, which is model information compiled by the server; based on the target model information, the server determines whether the terminal supports the deployment of the model to be deployed, and determines whether the model to be deployed meets the model inference requirements when running on the terminal.
[0220] Sc2. If the terminal supports the deployment of the model to be deployed and the model to be deployed meets the model inference requirements when running on the terminal, the server determines that the model to be deployed has been successfully deployed on the terminal; or, if the terminal does not support the deployment of the model to be deployed and / or the model to be deployed does not meet the model inference requirements when running on the terminal, the server determines that the model to be deployed has failed to be deployed on the terminal.
[0221] The model transmission method provided in this disclosure enables a model to be successfully deployed on a terminal to be supported by the terminal and to meet the model inference requirements when running on the terminal. In other words, it ensures that a model successfully deployed on the terminal is usable. Furthermore, the network device can control the model deployment on the terminal through model inference requirements sent by the network device.
[0222] In some embodiments, the model deployment request message also includes the priority of the model to be deployed; prior to step S302 above, it further includes:
[0223] Sd1. The server determines whether a reference model exists in the terminal based on the priority of the model to be deployed. The reference model is already deployed in the terminal and its priority is lower than that of the model to be deployed.
[0224] Sd2. In response to the existence of a reference model in the terminal, the server disables the reference model in the terminal simulation environment.
[0225] In one example, the server can acquire terminal runtime status information in real time or periodically. This runtime status information is used to determine the terminal's simulated runtime environment. Based on this, the server can determine the terminal's simulated runtime environment to make decisions regarding model deployment.
[0226] Furthermore, in response to disabling the reference model in the terminal simulation environment, the server executes step S302; or, in response to the absence of a reference model in the terminal, the server directly executes step S302.
[0227] Based on this, if the priority of the model to be deployed is high, the reference model that has already been deployed in the terminal and has a lower priority can be deactivated first, and then the deployment result of the model to be deployed can be determined, so as to ensure that the terminal's resources are given priority to the model to be deployed with higher priority.
[0228] In some embodiments, the model information of the model to be deployed is model information that has not been compiled by the server. The server further performs the following steps: the server compiles the model information of the model to be deployed to obtain target model information, which is model information compiled by the server; in response to the second model deployment response message indicating that the model to be deployed can be successfully deployed on the terminal, the server sends the target model information.
[0229] In one example, the server sends the target model information and the second model deployment response information separately.
[0230] In another example, the server sends a second model deployment response message, which includes the target model information.
[0231] For example, the model information of the model to be deployed is model information that has not been compiled by the server; step S302 includes: the server compiles the model information of the model to be deployed to obtain target model information; in response to the second model deployment response message indicating that the model to be deployed can be successfully deployed on the terminal, the server sends the second model deployment response message; the second model deployment response message indicates that the model to be deployed can be successfully deployed on the terminal, and the second model deployment response message includes target model information.
[0232] In one example, when the model deployment request message is sent by the terminal, the server sending the target model information includes: the server sending the target model information to the terminal.
[0233] In another example, where the model deployment request message is sent by a network device, the server sending the target model information includes: the server sending the target model information to the network device.
[0234] As can be seen, the above description of the model transmission method provided by this disclosure is from the perspectives of terminals, network devices, and servers. To further illustrate the model transmission method provided by this disclosure, some embodiments are also provided in the above-mentioned terminal, network device, and server interaction scenarios. It should be understood that the following embodiments are only some combinations of the above-mentioned model transmission method, not all combinations, and should not constitute a specific limitation on this disclosure. Other possible combinations are also within the protection scope of this disclosure.
[0235] In one embodiment, such as Figure 10 As shown, the model transfer method includes:
[0236] Step 1: The network device sends a model deployment request message to the terminal; correspondingly, the terminal receives the model deployment request message sent by the network device.
[0237] The model deployment request message includes model information of the model to be deployed and model inference requirements.
[0238] In some examples, the model deployment request message also includes the priority of the model to be deployed.
[0239] Step 2: Based on the model deployment request message, the terminal sends a first model deployment response message to the network device. The first model deployment response message is used to indicate whether the model to be deployed has been successfully deployed on the terminal. Correspondingly, the network device receives the first model deployment response message sent by the terminal.
[0240] In some examples, the response to the model deployment request message also includes the priority of the model to be deployed, based on... Figure 10 Before step two, as shown, the following also applies:
[0241] The terminal determines whether a reference model exists in the terminal based on the priority of the model to be deployed. If the reference model has been deployed in the terminal and its priority is lower than that of the model to be deployed;
[0242] In response to the presence of a reference model in the terminal, the terminal disables the reference model; in response to the terminal disabling the reference model, the terminal executes the above-mentioned... Figure 10 Step two as shown; or,
[0243] In response to the absence of a reference model in the terminal, the terminal directly executes the above-mentioned method based on... Figure 10 Step two is shown.
[0244] The model transmission method provided in this disclosure, on the one hand, enables the network device to control the deployment of models on the terminal through model inference requests sent by the network device; on the other hand, it allows the terminal to inform the network device whether the model to be deployed has been successfully deployed on the terminal based on the model information and model inference requests of the model to be deployed. Furthermore, since the successful deployment of a model on the terminal is related to the model inference requests and model information sent by the network device, a model that can be successfully deployed on the terminal is guaranteed to be usable.
[0245] In one embodiment, based on Figure 10 The illustrated embodiments, such as Figure 11 As shown, the model transfer method includes:
[0246] Step 1: The network device sends a model deployment request message to the terminal; correspondingly, the terminal receives the model deployment request message sent by the network device.
[0247] The model deployment request message includes the model information of the model to be deployed and the model inference requirements, and the model information of the model to be deployed is the model information that has not been compiled by the server.
[0248] In some examples, the model deployment request message also includes the priority of the model to be deployed.
[0249] Step 2: The terminal sends the model information of the model to be deployed to the server; correspondingly, the server receives the model information of the model to be deployed sent by the terminal; furthermore, the server can also compile the model information of the model to be deployed to obtain the target model information, which is the model information compiled by the server.
[0250] In some examples, the terminal only sends model information of the model to be deployed to the server, and correspondingly, the server only receives the model information of the model to be deployed sent by the terminal. For example... Figure 11 As shown in (a) in the figure.
[0251] In other examples, the terminal sends a model deployment request message to the server, which includes model information about the model to be deployed and model inference requirements; correspondingly, the server receives the model deployment request message sent by the terminal. For example... Figure 11 As shown in (b) of the diagram.
[0252] Furthermore, the server can also use the model deployment request message to determine whether the model to be deployed can be successfully deployed on the terminal; in response to the server determining that the model to be deployed can be successfully deployed on the terminal, it can continue to execute subsequent steps.
[0253] Step 3: The server sends the target model information to the terminal; correspondingly, the terminal receives the target model information; furthermore, the terminal also updates the model information to be deployed with the target model information.
[0254] In some examples, the server only sends the target model information to the terminal. For example... Figure 11 As shown in (a) in the figure.
[0255] In other examples, the server sends a second model deployment response to the terminal; correspondingly, the terminal receives the second model deployment response from the server. For example... Figure 11 As shown in (b) of the diagram.
[0256] The second model deployment response information is used to indicate whether the model to be deployed can be successfully deployed on the terminal. In response to the second model deployment response information indicating that the model to be deployed can be successfully deployed on the terminal, the second model deployment response information includes target model information.
[0257] It should be understood that, Figure 11 In the scenario shown in (b), if the second model deployment response information indicates that the model to be deployed cannot be successfully deployed on the terminal, the second model deployment response information returned by the server may not include the target model information.
[0258] Step 4: The terminal sends a first model deployment response message to the network device; the first model deployment response message is used to indicate whether the model to be deployed has been successfully deployed on the terminal; correspondingly, the network device receives the first model deployment response message sent by the terminal.
[0259] In some examples, the terminal determines the deployment result of the model to be deployed based on the model deployment request message; based on the deployment result, it sends a first model deployment response message to the network device. Figure 11 (a) in the middle.
[0260] In other examples, the terminal sends a first model deployment response message to the network device based on the second model deployment response information, as detailed in the description of step S1023 in the above embodiments. Figure 11 (b) in the middle.
[0261] based on Figure 11 The illustrated embodiment, in Figure 11 Before step one, the following steps may also be included: the terminal sends its model processing capability to the network device, which is used to assist in generating model information of the model to be deployed; correspondingly, the network device receives the terminal's model processing capability sent by the terminal; furthermore, the network device may also determine the model information of the model to be deployed based on the terminal's model processing capability.
[0262] In one embodiment, based on Figure 10 The illustrated embodiments, such as Figure 12 As shown, the model transfer method includes:
[0263] Step 1: The network device sends the initial model information of the model to be deployed to the server corresponding to the terminal; correspondingly, the server receives the initial model information of the model to be deployed sent by the network device; wherein, the initial model information of the model to be deployed is model information that has not been compiled by the server.
[0264] In some examples, the network device only sends the initial model information of the model to be deployed to the server corresponding to the terminal; correspondingly, the server only receives the initial model information of the model to be deployed. For example... Figure 12 As shown in (a) in the figure.
[0265] In other examples, the network device sends a model deployment request message to the server corresponding to the terminal. This message includes initial model information and model inference requirements for the model to be deployed; correspondingly, the server receives the model deployment request message. Optionally, the model deployment request message may also include the priority of the model to be deployed. For example... Figure 12 As shown in (b) of the diagram.
[0266] Step 2: The server compiles the initial model information of the model to be deployed to obtain the target model information, which is the model information compiled by the server; the server sends the target model information to the network device; correspondingly, the network device receives the target model information and identifies it as the model information of the model to be deployed.
[0267] In some examples, the server only sends the target model information to the network device; correspondingly, the network device receives the target model information and identifies it as the model information for the model to be deployed. For example... Figure 12 As shown in (a) in the figure.
[0268] In other examples, the server sends a second model deployment response to the network device. This second model deployment response indicates whether the model to be deployed can be successfully deployed on the terminal. In response to the second model deployment response indicating successful deployment on the terminal, the second model deployment response includes target model information. Correspondingly, the network device receives the second model deployment response sent by the server. For example... Figure 12 As shown in (b) above. It should be understood that in this scenario, if the second model deployment response information indicates that the model to be deployed cannot be successfully deployed on the terminal, the second model deployment response information returned by the server may not include the target model information. It should be understood that if the second model deployment response information indicates that the model to be deployed cannot be successfully deployed on the terminal, the network device can regenerate the initial model information of the model to be deployed and re-execute based on... Figure 12 Steps one and two of the embodiment.
[0269] Step 3: The network device sends a model deployment request message to the terminal; correspondingly, the terminal receives the model deployment request message sent by the network device. The model deployment request message includes model information of the model to be deployed and model inference requirements.
[0270] In some examples, the model deployment request message also includes the priority of the model to be deployed.
[0271] Step 4: Based on the model deployment request message, the terminal sends a first model deployment response message to the network device. The first model deployment response message is used to indicate whether the model to be deployed has been successfully deployed on the terminal; correspondingly, the network device receives the first model deployment response message sent by the terminal.
[0272] based on Figure 12 The illustrated embodiment, in Figure 12 Before step one shown, the following steps may also be included: the terminal sends its model processing capability and / or the address of the server corresponding to the terminal to the network device; correspondingly, the network device receives the terminal's model processing capability and / or the address of the server corresponding to the terminal sent by the terminal.
[0273] Furthermore, the network device can also determine the model information of the model to be deployed based on the terminal's model processing capabilities. Additionally, the network device can also execute commands based on the address of the server corresponding to the terminal. Figure 12 Step one is shown.
[0274] As can be seen, the above mainly describes the solutions provided by the embodiments of this disclosure from a methodological perspective. To achieve the above functions, the embodiments of this disclosure provide corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the modules and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of this disclosure can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0275] The related apparatus provided in this disclosure will now be described. It should be understood that the communication apparatus described below can be referred to in correspondence with the model transmission method described above.
[0276] like Figure 13 As shown in the figure, this disclosure also provides a structural diagram of a terminal, including a first communication module 1310. In some embodiments, it further includes a model deployment module 1320.
[0277] The first communication module 1310 is used to receive a model deployment request message sent by a network device. The model deployment request message includes model information of the model to be deployed and model inference requirements.
[0278] The first communication module 1310 is used to send a first model deployment response message to the network device based on the model deployment request message. The first model deployment response message is used to indicate whether the model to be deployed has been successfully deployed on the terminal.
[0279] In some embodiments, the model deployment module 1320 is used to determine the deployment result of the model to be deployed based on the model deployment request message; and to send a first model deployment response message to the network device based on the deployment result.
[0280] In some embodiments, the model deployment module 1320 is specifically used to determine, based on the model information of the model to be deployed, whether the terminal supports the deployment of the model to be deployed, and whether the model to be deployed meets the model inference requirements when running on the terminal; in response to the terminal supporting the deployment of the model to be deployed and the model to be deployed meeting the model inference requirements when running on the terminal, determine that the model to be deployed has been successfully deployed on the terminal; or, in response to the terminal not supporting the deployment of the model to be deployed and / or the model to be deployed not meeting the model inference requirements when running on the terminal, determine that the model to be deployed has failed to be deployed on the terminal.
[0281] In some embodiments, the model deployment request message also includes the priority of the model to be deployed. The model deployment module 1320 is specifically used to determine whether a reference model exists in the terminal based on the priority of the model to be deployed before sending a first model deployment response message to the network device based on the model deployment request message. If the reference model has been deployed in the terminal and its priority is lower than that of the model to be deployed, the reference model is disabled in response to the existence of the reference model in the terminal.
[0282] In some embodiments, the model information of the model to be deployed is model information that has not been compiled by the server. The first communication module 1310 is also used to send the model information of the model to be deployed to the server and receive the target model information sent by the server. The target model information is model information that has been compiled by the server and is obtained by the server compiling the model information of the model to be deployed. The model deployment module 1320 is also used to update the model information to be deployed with the target model information.
[0283] In some embodiments, the first communication module 1310 is further configured to send the address of the server corresponding to the terminal to the network device, and the server is configured to generate model information of the model to be deployed.
[0284] In some embodiments, the first communication module 1310 is further configured to send a model deployment request message to the server; receive a second model deployment response message sent by the server, the second model deployment response message being used to indicate whether the model to be deployed can be successfully deployed on the terminal; and send a first model deployment response message to the network device based on the second model deployment response message.
[0285] In some embodiments, the first communication module 1310 is specifically configured to: in response to a second model deployment response message indicating that the model to be deployed cannot be successfully deployed on the terminal, send a first model deployment response message to the network device, the first model deployment response message indicating that the model to be deployed has failed to be deployed on the terminal; or, in response to a second model deployment response message indicating that the model to be deployed can be successfully deployed on the terminal, send a first model deployment response message to the network device, the first model deployment response message indicating that the model to be deployed can be successfully deployed on the terminal; or, in response to a second model deployment response message indicating that the model to be deployed can be successfully deployed on the terminal, send a first model deployment response message to the network device based on the model deployment request message.
[0286] In some embodiments, the first communication module 1310 is further configured to send the terminal's model processing capability to the network device, the model processing capability being used to assist in generating model information for a model to be deployed.
[0287] like Figure 14As shown in the figure, this disclosure also provides a structural diagram of a network device, including a second communication module 1410. In some embodiments, the network device further includes a model control module 1420.
[0288] The second communication module 1410 is used to send a model deployment request message to the terminal. The model deployment request message includes model information of the model to be deployed and model inference requirements.
[0289] The second communication module 1410 is also used to receive a first model deployment response message sent by the terminal, which indicates whether the model to be deployed has been successfully deployed on the terminal.
[0290] In some embodiments, the model information of the model to be deployed is model information compiled by the server. The model control module 1420 is used to generate initial model information of the model to be deployed, which is model information not compiled by the server. The second communication module 1410 is also used to send the initial model information of the model to be deployed to the server corresponding to the terminal and receive target model information sent by the server. The target model information is model information compiled by the server and is obtained by the server compiling the initial model information of the model to be deployed. The model control module 1420 is also used to determine the target model information as the model information of the model to be deployed.
[0291] In some embodiments, the second communication module 1410 is further configured to send a model deployment request message to the server; receive a second model deployment response message sent by the server, the second model deployment response message being used to indicate whether the model to be deployed can be successfully deployed on the terminal; and, if the second model deployment response message indicates that the model to be deployed can be successfully deployed on the terminal, send a model deployment request message to the terminal.
[0292] like Figure 15 As shown in the figure, this disclosure also provides a structural diagram of a server, including a third communication module 1510. In some embodiments, the server further includes a model processing module 1520.
[0293] The third communication module 1510 is used to receive model deployment request messages, which include model information of the model to be deployed and model inference requirements.
[0294] The third communication module 1510 is also used to send a second model deployment response message based on the model deployment request message. The second model deployment response message is used to indicate whether the model to be deployed can be successfully deployed on the terminal.
[0295] In some embodiments, the model processing module 1520 is used to determine the deployment result of the model to be deployed based on the model deployment request message; the third communication module 1510 is specifically used to send a second model deployment response message based on the deployment result.
[0296] In some embodiments, the model processing module 1520 is specifically used to determine, based on the model information of the model to be deployed, whether the terminal supports the deployment of the model to be deployed, and whether the model to be deployed meets the model inference requirements when running on the terminal; in response to the terminal supporting the deployment of the model to be deployed and the model to be deployed meeting the model inference requirements when running on the terminal, determine that the model to be deployed has been successfully deployed on the terminal; or, in response to the terminal not supporting the deployment of the model to be deployed and / or the model to be deployed not meeting the model inference requirements when running on the terminal, determine that the model to be deployed has failed to be deployed on the terminal.
[0297] In some embodiments, the model information of the model to be deployed is model information that has not been compiled by the server. The model processing module 1520 is also used to compile the model information of the model to be deployed to obtain target model information, which is model information compiled by the server. The third communication module 1510 is specifically used to send the target model information in response to the second model deployment response message indicating that the model to be deployed can be successfully deployed on the terminal.
[0298] It should be noted that, Figures 13 to 15 The module division shown is illustrative and represents only one logical functional division; in actual implementation, other division methods are possible. For example, two or more functions can be integrated into a single processing module. These integrated modules can be implemented either in hardware or as software functional modules.
[0299] In implementing the functions of the integrated modules described above in hardware, this disclosure also provides a possible structure for a communication device used to execute the model transmission method provided in this disclosure. Similarly, the communication device and the model transmission method described above can be referred to in correspondence.
[0300] like Figure 16 As shown, the communication device includes a processor 1602 and a communication interface 1603. In some examples, the communication device may also include at least one of a bus 1604 and a memory 1601.
[0301] Processor 1602 may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with embodiments of this disclosure. Processor 1602 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with embodiments of this disclosure. Processor 1602 may also be a combination of computing functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0302] The communication interface 1603 is used to connect to other devices via a communication network. This communication network can be Ethernet, wireless access network, wireless local area network (WLAN), etc.
[0303] The memory 1601 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.
[0304] As one possible implementation, the memory 1601 can exist independently of the processor 1602. The memory 1601 can be connected to the processor 1602 via a bus 1604 and is used to store instructions or program code executable by the processor 1602, such as computer program instructions. When the processor 1602 calls and executes the instructions or program code stored in the memory 1601, it can implement the model transfer method provided in this embodiment.
[0305] In another possible implementation, the memory 1601 can also be integrated with the processor 1602.
[0306] The 1604 bus can be an extended industry standard architecture (EISA) bus, etc. The 1604 bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 16The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0307] Some embodiments of this disclosure provide a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium) storing computer program instructions that, when executed on a computer (e.g., the aforementioned communication device, base station, first terminal, second terminal, and their processor, etc.), cause the computer to perform the model transmission method as described in any of the above embodiments. It should be understood that this disclosure does not limit the specific form of the computer.
[0308] In some examples, the aforementioned computer-readable storage media may include, but are not limited to: magnetic storage devices (e.g., hard disks, floppy disks, or magnetic tapes), optical discs (e.g., compact disks (CDs), digital versatile disks (DVDs), etc.), smart cards, and flash memory devices (e.g., erasable programmable read-only memory (EPROMs), cards, sticks, or key drives, etc.). The various computer-readable storage media described in this disclosure may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage media" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.
[0309] This disclosure provides a computer program product containing instructions that, when run on a computer, cause the computer to execute the model transfer method described in any of the above embodiments.
[0310] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any changes or substitutions within the technical scope disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A model transfer method, characterized in that, Applied to a terminal, the method includes: Receive a model deployment request message sent by a network device, the model deployment request message including model information of the model to be deployed and model inference requirements; Based on the model deployment request message, a first model deployment response message is sent to the network device. The first model deployment response message is used to indicate whether the model to be deployed has been successfully deployed on the terminal.
2. The method according to claim 1, characterized in that, The model inference requirements include at least one of the following: inference latency requirements, inference sample requirements, and inference frequency requirements.
3. The method according to claim 1, characterized in that, In response to the first model deployment response message indicating that the model to be deployed has been successfully deployed on the terminal, the first model deployment response message includes the activation time of the model to be deployed; or, In response to the first model deployment response message indicating that the model to be deployed has failed to be deployed on the terminal, the first model deployment response message includes the reason for the failure.
4. The method according to claim 1, characterized in that, The step of sending a first model deployment response message to the network device based on the model deployment request message includes: Based on the model deployment request message, determine the deployment result of the model to be deployed; Based on the deployment results, a first model deployment response message is sent to the network device.
5. The method according to claim 4, characterized in that, The step of determining the deployment result of the model to be deployed based on the model deployment request message includes: Based on the model information of the model to be deployed, determine whether the terminal supports the deployment of the model to be deployed, and determine whether the model to be deployed meets the model inference requirements when running on the terminal; In response to the terminal supporting the deployment of the model to be deployed and the model to be deployed meeting the model inference requirements when running on the terminal, it is determined that the model to be deployed has been successfully deployed on the terminal; or... In response to the terminal not supporting the deployment of the model to be deployed and / or the model to be deployed not meeting the model inference requirements when running on the terminal, it is determined that the deployment of the model to be deployed on the terminal has failed.
6. The method according to claim 1, characterized in that, The model deployment request message also includes the priority of the model to be deployed; before sending a first model deployment response message to the network device based on the model deployment request message, the method further includes: Based on the priority of the model to be deployed, determine whether there is a reference model in the terminal, wherein the reference model has been deployed in the terminal and the priority of the reference model is lower than the priority of the model to be deployed; In response to the presence of the reference model in the terminal, the reference model is disabled.
7. The method according to claim 1, characterized in that, The model information of the model to be deployed is model information that has been compiled by the server; or, the model information of the model to be deployed is model information that has not been compiled by the server.
8. The method according to claim 7, characterized in that, The model information of the model to be deployed is model information that has not been compiled by the server, and the method further includes: Send the model information of the model to be deployed to the server; The server receives target model information; wherein the target model information is obtained by the server compiling the model information of the model to be deployed. The model information of the model to be deployed is updated with the target model information.
9. The method according to claim 7, characterized in that, The model information of the model to be deployed is model information compiled by the server. Before receiving the model deployment request message sent by the network device, the method further includes: The address of the server corresponding to the terminal is sent to the network device, and the server is used to generate the model information of the model to be deployed.
10. The method according to claim 1, characterized in that, The step of sending a first model deployment response message to the network device based on the model deployment request message includes: Send the model deployment request information to the server; The server receives a second model deployment response message, which indicates whether the model to be deployed can be successfully deployed on the terminal. Based on the second model deployment response information, a first model deployment response message is sent to the network device.
11. The method according to claim 10, characterized in that, The step of sending a first model deployment response message to the network device based on the second model deployment response includes: In response to the second model deployment response information indicating that the model to be deployed cannot be successfully deployed on the terminal, a first model deployment response message is sent to the network device, the first model deployment response message indicating that the model to be deployed has failed to deploy on the terminal; or... In response to the second model deployment response information indicating that the model to be deployed can be successfully deployed on the terminal, a first model deployment response message is sent to the network device, the first model deployment response message indicating that the model to be deployed can be successfully deployed on the terminal; or... In response to the second model deployment response information indicating that the model to be deployed can be successfully deployed on the terminal, a first model deployment response message is sent to the network device based on the model deployment request message.
12. The method according to claim 1, characterized in that, The method further includes: The model processing capability of the terminal is sent to the network device, and the model processing capability is used to assist in generating model information of the model to be deployed.
13. The method according to claim 1, characterized in that, The model information includes at least one of the following: model parameters, model structure, preprocessing information of model input data, and model download address.
14. A model transfer method, characterized in that, Applied to network devices, the method includes: Send a model deployment request message to the terminal. The model deployment request message includes model information of the model to be deployed and model inference requirements. The system receives a first model deployment response message sent by the terminal, which indicates whether the model to be deployed has been successfully deployed on the terminal.
15. The method according to claim 14, characterized in that, The model information of the model to be deployed is model information that has been compiled by the server; or, the model information of the model to be deployed is model information that has not been compiled by the server.
16. The method according to claim 15, characterized in that, Before sending the model deployment request message to the terminal, the method further includes: Generate initial model information for the model to be deployed, wherein the initial model information for the model to be deployed is model information that has not been compiled by the server; Send the initial model information of the model to be deployed to the server corresponding to the terminal; The server receives target model information; wherein the target model information is obtained by the server compiling the initial model information of the model to be deployed. The target model information is determined to be the model information of the model to be deployed.
17. The method according to claim 14, characterized in that, The method further includes: Send the model deployment request message to the server; The server receives a second model deployment response message, which indicates whether the model to be deployed can be successfully deployed on the terminal. Sending the model deployment request message to the terminal includes: If the second model deployment response information indicates that the model to be deployed can be successfully deployed on the terminal, a model deployment request message is sent to the terminal.
18. A model transfer method, characterized in that, Applied to a server, the method includes: Receive a model deployment request message, which includes model information of the model to be deployed and model inference requirements; Based on the model deployment request message, a second model deployment response message is sent, which indicates whether the model to be deployed can be successfully deployed on the terminal.
19. The method according to claim 18, characterized in that, The step of sending a second model deployment response message based on the model deployment request message includes: Based on the model deployment request message, determine the deployment result of the model to be deployed; Based on the deployment results, a second model deployment response message is sent.
20. The method according to claim 19, characterized in that, The step of determining the deployment result of the model to be deployed based on the model deployment request message includes: Based on the model information of the model to be deployed, determine whether the terminal supports the deployment of the model to be deployed, and determine whether the model to be deployed meets the model inference requirements when running on the terminal; In response to the terminal supporting the deployment of the model to be deployed and the model to be deployed meeting the model inference requirements when running on the terminal, it is determined that the model to be deployed has been successfully deployed on the terminal; or... In response to the terminal not supporting the deployment of the model to be deployed and / or the model to be deployed not meeting the model inference requirements when running on the terminal, it is determined that the deployment of the model to be deployed on the terminal has failed.
21. The method according to claim 18, characterized in that, The model information of the model to be deployed is model information that has not been compiled by the server, and the method further includes: The model information of the model to be deployed is compiled to obtain target model information, which is model information compiled by the server; In response to the second model deployment response message indicating that the model to be deployed can be successfully deployed on the terminal, the target model information is sent.
22. A communication device, characterized in that, include: A processor and a memory for storing processor-executable instructions; The processor is configured to execute the instructions, causing the communication device to perform the model transmission method as described in any one of claims 1-21.
23. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed on a computer, cause the computer to perform the model transfer method as described in any one of claims 1-21.
24. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed on a computer, cause the computer to perform the model transfer method as described in any one of claims 1-21.