Model reasoning and training method, device and system

CN120234467APending Publication Date: 2025-07-01HUAWEI TECH CO LTD +1
0 Cites 0 Cited by

Patent Information

Application Number
CN202311863964.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

Smart Images

  • Figure CN120234467A_ABST
    Figure CN120234467A_ABST
Patent Text Reader

Abstract

The invention provides a model reasoning and training method, device and system, which are applied to the technical field of artificial intelligence, such as machine learning. The method is applied to a model training system and an inference system, in particular to a recommendation system, after a recommendation request is obtained from a client, content related to the recommendation request is recalled, the recommendation request is inferred to obtain an intermediate inference result, and then the intermediate inference result is sent to the client for further inference. And outputting a final reasoning result, analyzing the result and the content related to the recalled recommendation request, and outputting a final recommendation result. In the recommendation system, part of model parameters of the server side are trained at the client side, extreme resources of the server side can be saved, meanwhile, after the server side completes part of reasoning, the client side continues to conduct reasoning, computing resources of the server side and the client side can be reasonably utilized, and the recommendation result can be more optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular, to a method, apparatus, and system for model inference and training. Background Art

[0002] Currently, in order to implement model inference and training, in the current prior art, there are two methods. The first method is that the cloud side and the edge side use a distributed framework for model training. The edge-side model is trained in real time based on Delta-tuning. The cloud side collects parameters, and the edge-side model performs inference. This method can reduce the computational load on the cloud side. The second method is to use the edge-side machine for model training and perform aggregation on the edge side, that is, the edge-side data does not need to be uploaded. The edge-side data is maintained locally and does not need to be uploaded to the cloud side. After the edge-side model training is completed, only the parameters of the edge-side model are uploaded to the cloud side, and the parameters are updated on the cloud side based on the model trained on the edge side. This method can reduce the training cost on the cloud side. The problem with these methods is that inference or training is all completed on the edge side. There is a large performance gap between the edge-side machine and the cloud-side machine. It is very difficult to place the entire model on the edge side and a lot of adaptation work needs to be done. Based on this, this application provides a method for model inference and training to solve the above problems. Summary of the Invention

[0003] In a first aspect, the present invention provides a recommendation system. The recommendation system includes a server and a client. The server deploys a first model, and the client deploys a second model. The client is configured to send a recommendation request to the server; the server is configured to obtain a plurality of first recommendation results based on the recommendation request; the server is configured to obtain a first vector through the first model based on the recommendation request; the client is configured to input the first vector from the server into the second model to obtain a second vector; the client is configured to evaluate the plurality of first recommendation results from the server based on the second vector to output at least one second recommendation result.

[0004] It should be noted that the first vector and the second vector are the results of different stages of the inference of the recommended product, such as the picture data, link data, etc. of a certain product.

[0005] It can be seen that after the server obtains multiple first recommendation results based on the recommendation request, the recommendation request is input into the first model for inference to obtain a first vector, and then the first vector is input into the second model of the client for further inference to obtain a second vector. In this way, through the collaborative participation of the server and the client in the inference, part of the inference task on the server side is assigned to the client, which can reduce the computing cost on the cloud side and also make reasonable use of the computing resources on the client side. At the same time, the client evaluates the multiple first recommendation results based on the second vector to output at least one second recommendation result, that is, after the server model obtains a preliminary inference result, the client model performs further inference to optimize the inference result, that is, the server and the client jointly recommend, making the recommendation result more accurate.

[0006] In a possible implementation manner, the second model includes a first sub-model and a second sub-model. The first sub-model is trained in the client, and the second sub-model is trained in the server.

[0007] It can be seen that some sub-models in the client's second model are trained by the server, which can reduce the computing burden on the client side and lower the training cost of the client-side model.

[0008] In a possible implementation manner, the client is further configured to train the first sub-model based on the user's feedback on the at least one second recommendation result.

[0009] It can be seen that the training of the first sub-model can be achieved through the feedback of the at least one second recommendation result, which reduces the training burden on the client and lowers the training cost of the client.

[0010] In a possible implementation manner, the recommendation request includes at least one of the following parameters: user identification, geographical location information, Wi-Fi information, and time information.

[0011] It should be noted that the user identification includes user ID, user nickname, device number, etc. The present application does not limit the expression form of the user identification. Here, Wi-Fi, also known as "Wireless Network", is a trademark of the Wi-Fi Alliance and a wireless local area network technology based on the IEEE 802.11 standard.

[0012] It can be seen that by obtaining at least one of the above parameters, multiple first recommendation results can be obtained, that is, without limiting the obtained parameters, multiple recommendation results can also be obtained, making the recommendation more flexible.

[0013] Second aspect, the present invention provides a recommendation method, which is applied to a client, and a second model is deployed on the client. It is characterized in that the client is used to send a recommendation request to the server; the client is used to input the first vector from the server into the second model to obtain a second vector; the client is used to evaluate the multiple first recommendation results from the server based on the second vector to output at least one second recommendation result.

[0014] It can be seen from this that the client uses the model updated by the client and the model continuously updated by the server, without directly requesting the server, and directly performs inference on the client side. That is, after the server completes part of the inference, the result is transmitted to the client in the form of a vector, so that it is not necessary to perform the difficult work of adapting to the large model on the client side, and the inference result can also be optimized.

[0015] In a possible implementation manner, the second model includes a first sub-model and a second sub-model. The first sub-model is trained on the client, and the second sub-model is trained on the server.

[0016] It can be seen from this that some sub-models in the second model of the client are trained by the server, which can reduce the computing burden on the client side and reduce the training cost of the client-side model.

[0017] In a possible implementation manner, the client is further used to train the first sub-model based on the user's feedback on the at least one second recommendation result.

[0018] It can be seen from this that the training of the first sub-model can be achieved through the feedback of at least one second recommendation result, which reduces the training burden of the client and reduces the training cost of the client.

[0019] In a possible implementation manner, the recommendation request includes at least one of the following parameters: user identification, geographical location information, Wi-Fi information, time information.

[0020] It should be noted that the user identification includes user ID, user nickname, device number, etc. The present application does not limit the expression form of the user identification. Here, Wi-Fi, also known as "Wireless Network", is a trademark of the Wi-Fi Alliance and a wireless local area network technology based on the IEEE 802.11 standard.

[0021] It can be seen from this that by obtaining at least one of the above parameters, multiple first recommendation results can be obtained, that is, without limiting the obtained parameters, multiple recommendation results can also be obtained, making the recommendation more flexible.

[0022] In a third aspect, the present invention provides a recommendation method, which is applied to a server. The server is deployed with a first model. It is characterized in that the server is used to obtain a plurality of first recommendation results based on the recommendation request; the server is used to obtain a first vector through the first model based on the recommendation request;

[0023] In a fourth aspect, the present invention provides a model training system, which includes a server and a client. It is characterized in that

[0024] The server is deployed with a model system, which includes a first model, a second model, a third model, and a copy of the second model. The output of the first model corresponds to the input of the second model. The third model corresponds to the historical state of the first model. The first model and the third model have the same input. The output of the third model corresponds to the input of the copy; the client is used to send data and model parameters to the server, and the data indicates the behavior of the user; the server is used to receive the data and the model parameters from the client; the server is used to update the second model based on the model parameters and keep the copy unchanged; the server is used to calculate a first loss using the data, the first model, and the second model; the server is used to calculate a second loss using the third model and the copy; the server is used to update the model system based on the first loss and the second loss to obtain the trained first model.

[0025] It should be noted that the historical state here is referenced to the trained first model, and the historical state is the state of the first model before training. The copy of the second model here can be understood as the second model before training.

[0026] From this, it can be seen that when the server performs training and update, it only trains through the models on the server side, rather than relying on the models on the client side for training, thus avoiding major errors. That is, it is not necessary to update the server-side model in real time with the client's parameters, and the training effect of the server-side model can still be optimized with only training on the server side.

[0027] In a possible implementation, the second model includes a first sub-model and a second sub-model. The first sub-model is trained on the client side, and the second sub-model is trained on the server side.

[0028] From this, it can be seen that some sub-models of the second model on the client side are trained by the server, which can reduce the computational burden on the client side and lower the training cost of the client-side model.

[0029] In a possible implementation, the server is further configured to obtain a first vector based on a recommendation request through the first model; the server is further configured to send the first vector.

[0030] It can be seen from this that the server can obtain the first vector only through the first model based on the recommendation request, and at the same time can also send the first vector. This can reduce the computational burden on the server, save the server's computing resources, and also output the vector for external transmission, generating results and sending training results under the optimal conditions of the server's resources.

[0031] In a possible implementation, the client is configured to train the second model based on the first vector and the data of the client to update the first sub-model, and the data of the client indicates the behavior of the user.

[0032] It can be seen from this that by combining the first vector sent by the server with the data of the client to train the second model, the first sub-model in the second model is updated, and the second sub-model does not need to be updated, reducing the computational burden on the client and saving the computational cost of the client.

[0033] In a possible implementation, the server is configured to update the model system based on the first loss and the second loss, and is further configured to obtain the trained second sub-model.

[0034] It can be seen from this that after the server uses the most recent server model and the server model in the historical state for training respectively, two loss results are obtained. Based on these two loss results, the server's model is updated, so that training can be performed without relying on the second model for parameter update, ensuring the ability to perform independent training, and at the same time, the training effect is not affected.

[0035] In a possible implementation, the client is further configured to update the second model based on the trained second sub-model from the server.

[0036] It can be seen from this that a part of the client's model is updated with the parameters trained by a part of the server's model, which can reduce the training burden of the client's model and save the computing resources of the client.

[0037] Fifth aspect, the present invention provides a training method, which is applied to a server. It is characterized in that the server is deployed with a model system, and the model system includes a first model, a second model, a third model, and a copy of the second model. The output of the first model corresponds to the input of the second model. The third model corresponds to the historical state of the first model. The first model and the third model have the same input. The output of the third model corresponds to the input of the copy. The server is used to receive the data and the model parameters from the client. The server is used to update the second model based on the model parameters and keep the copy unchanged. The server is used to calculate a first loss using the data, the first model, and the second model. The server is used to calculate a second loss using the third model and the copy. The server is used to update the model system based on the first loss and the second loss to obtain the trained first model.

[0038] It can be seen that when the server performs training and update, it only trains through the models on the server side, rather than relying on the models of the client for training to avoid major errors. That is, it does not require the parameters of the client to update the models of the server in real time, and it can also optimize the training effect of the models of the server with only training on the server side.

[0039] In a possible implementation, the second model includes a first sub-model and a second sub-model. The first sub-model is trained on the client side, and the second sub-model is trained on the server side.

[0040] It can be seen that some sub-models in the second model of the client are trained by the server, which can reduce the computing burden on the client side and lower the training cost of the client-side model.

[0041] In a possible implementation, the server is further used to obtain a first vector through the first model based on a recommendation request. The server is further used to send the first vector.

[0042] It can be seen that the server can obtain the first vector only through the first model based on the recommendation request, and at the same time can send the first vector. This can reduce the computing burden on the server, save the computing resources of the server, and also output and send the vector. The result is generated and the training result is sent under the condition of optimizing the server resources.

[0043] In a possible implementation, the server is used to update the model system based on the first loss and the second loss, and is further used to obtain the trained second sub-model.

[0044] It can be seen from this that after the server uses the most recent server model and the server model of the historical state for training respectively, two loss results are obtained. Based on these two loss results, the model of the server is updated, so that training can be carried out without relying on the second model for parameter update, ensuring the ability to perform independent training, and at the same time, the training effect is not affected.

[0045] In a sixth aspect, the present invention provides a model training device deployed on a server. The device includes a receiving module, a storage module, a sending module, and a training module. The receiving module is used to receive the data and the model parameters from the client; the storage module is used to deploy a model system; the model system includes a first model, a second model, a third model, and a copy of the second model; the training module is used to calculate a first loss using the data, the first model, and the second model; calculate a second loss using the third model and the copy; update the model system based on the first loss and the second loss to obtain the trained first model; update the model system based on the first loss and the second loss to obtain the trained second sub-model. The sending module is used to send the first vector.

[0046] In a seventh aspect, the present invention provides a model inference device deployed on a server. The device includes a sending and receiving module, a storage module, and an inference module. The sending module is used to send a recommendation request to the server; the storage module is used to deploy a first model and a second model; the inference module is used to obtain a plurality of first recommendation results based on the recommendation request; is also used to obtain a first vector through the first model based on the recommendation request; is also used to input the first vector from the server into the second model to obtain a second vector; is also used to evaluate the plurality of first recommendation results from the server based on the second vector to output at least one second recommendation result. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic framework diagram of a model inference and training system provided by the present invention.

[0048] Figure 2 It is a schematic module diagram of a model inference and training system provided by the present invention.

[0049] Figure 3 It is a schematic structural diagram of a model on the terminal side provided by the present invention.

[0050] Figure 4a It is a schematic diagram of a model recommendation system provided by the present invention.

[0051] Figure 4bA system diagram of cloud-side model training provided by the present invention.

[0052] Figure 5 A system diagram of a model inference system provided by the present invention.

[0053] Figure 6a A schematic diagram of a model training system provided by the present invention

[0054] Figure 6b A schematic diagram of the cloud-side model training process provided by the present invention.

[0055] Figure 7 A schematic diagram of the edge-side model training process provided by the present invention.

[0056] Figure 8 A schematic diagram of the model inference process provided by the present invention.

[0057] Figure 9 A diagram showing an implementation form of a model inference and training method provided by the present invention.

[0058] Figure 10 A core flowchart of model training provided by the present invention.

[0059] Figure 11 A core flowchart of model inference provided by the present invention.

[0060] Figure 12 A flowchart of few-parameter fine-tuning for an edge-side model provided by the present invention.

[0061] Figure 13 A flowchart of cloud-side model training based on the continual learning method provided by the present invention. Detailed implementation manners

[0062] First, some technical terms appearing in this application are briefly described.

[0063] Few-parameter fine-tuning (delta-tuning): Compared with standard full-parameter fine-tuning, this method only fine-tunes a part of the model parameters, and the parameters of the rest of the model remain unchanged, which can greatly reduce the computational and storage costs, and at the same time has performance comparable to full-parameter fine-tuning.

[0064] Continual learning: It is a machine learning method that can alleviate catastrophic forgetting in deep learning models. It aims to continuously expand the adaptability of the model so that the model can learn knowledge of different tasks at different times. Currently, continual learning algorithms are mainly divided into 4 major aspects, namely regularization methods, memory replay methods, parameter isolation methods, and comprehensive methods. Different times can be time periods such as one day or one week apart for training.

[0065] Vector: A geometric object that has both magnitude and direction and satisfies the parallelogram law. In this application, it refers to some intermediate results calculated by the model.

[0066] Loss: The loss function is a function that maps the values of a random event or its associated random variables to non - negative real numbers to represent the "risk" or "loss" of the random event. In applications, the loss function is usually associated with the learning criterion and the optimization problem, that is, by minimizing the loss function to solve and evaluate the model. In machine learning, it is used to measure the degree of inconsistency between the predicted value f(x) of the model and the true value Y.

[0067] Label: It is an identifier that indicates the user to perform an operation. Taking the product recommendation on the home page of a shopping APP as an example, the label is that after the recommended product is displayed, the user clicks on the recommended product.

[0068] Adapter: It refers to some small - amount parameters inserted into the model. When fine - tuning for a downstream task, only these parameters are trained while keeping the original parameters of the pre - trained model unchanged, achieving the same or even better effect as fine - tuning the entire model.

[0069] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0070] To better understand the solution provided in the embodiments of the present application, first, the application scenario of the solution provided in the embodiments of the present application is introduced. The method provided in the embodiments of the present application is applied to a terminal. The terminal in the embodiments of the present application may include a mobile phone, a tablet computer, a notebook, etc. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of the present application are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product or device.

[0071] There are three main methods commonly used in the industry for end-cloud collaborative reasoning and training technology. The first is that both reasoning and training are performed on the cloud-side model. This technology places both the reasoning and training tasks of the model on the cloud-side model, requiring the cloud-side model to complete the reasoning of a huge number of end-side models, greatly increasing the reasoning cost.

[0072] The second method is to use the cloud-side model to complete the training task and the end-side model to complete the inference task. The overall process of this method is as follows: first, the problem is defined. Through demand analysis, business problems are identified and broken down into AI algorithm problems; second, data collection, cloud model design, training, compression and conversion are performed; finally, the model is deployed on the end-side. When the business is applied, the algorithm model is loaded for inference. In this case, the inference is completely placed on the end-side model. The large model cannot be directly placed on the end-side for inference. The end-side environment is complex, and a lot of additional adaptation and lightweight work needs to be done.

[0073] The third method is to complete the training task on the cloud-side model, complete part of the inference on the cloud-side model, and complete the rest of the inference task on the client-side model. When initiating a request, the client-side model completes some lightweight feature processing and part of the model inference task, transmits the inference result to the cloud-side model, and then the cloud-side model performs further inference and finally returns the result. This method still places the training task entirely on the cloud-side model, which is very costly and the effect of model training is easily affected.

[0074] In the prior art, the end-side model uses the Adapter method in Delta-tuning for training. When there is a certain amount of data on the end-side model, the end-side model uses the Adapter to perform training with a few parameter fine-tuning. The cloud side regularly uploads the parameters of the Adapter part of the end side and then stores them. The models updated by training on the end side and the models continuously updated by the cloud side do not need to directly request the cloud side model to perform reasoning first, but can be directly reasoned on the end side. This also brings some problems. The reasoning of the model is all completed on the end side. Although the computing tasks on the cloud side are reduced, it is relatively difficult to deploy large models to the end side, and a lot of development and adaptation work needs to be completed.

[0075] In another existing technology, the end-side model maintains the data locally without uploading it to the cloud-side model. At the same time, simulation training is performed based on the data on the end-side, and the parameters of the model after the end-side training are uploaded to the cloud-side model. The cloud-side model updates the parameters uploaded by the end-side, and after the parameters are aggregated, they are distributed to the end-side for further training. After the model is trained, reasoning is performed directly on the end-side. The problem with this method is that it is difficult for the end-side model to store and train a model that is too large.

[0076] To solve the above technical problems, the present application provides a method for model inference and training. This model inference and training method realizes end-cloud collaborative inference and training based on Delta-tuning. Although there are similar methods of introducing Delta-tuning in current end-cloud collaborative training, it does not consider how to save the computing cost on the cloud side and make good use of the computing resources on the edge side.

[0077] This method has the following beneficial technical effects: Combining with the Delta-tuning technology, in the model, the large base model performs inference and training on the cloud side, and a small part of the Adapter parameters perform inference and training on the edge side. The edge side realizes the Delta-tuning process based on the current user's request, thus realizing end-cloud collaboration based on the Delta-tuning large model.

[0078] At the same time, end-cloud independent updates are realized based on the method of continuous learning. For the update of the large model on the cloud side in the conventional method, it often needs to be accompanied by the simultaneous update of the edge-side model to ensure that there are no major errors in model inference, that is, when the large model is updated, the model parameters on the edge side are forgotten and thus the model parameters on the edge side need to be updated. The present application proposes a method of continuous learning to realize the update of the cloud-side model, without the need for real-time parameter update on the edge side, and at the same time, it can also ensure that the model effect is optimized with re-training.

[0079] It should be noted that the above first model corresponds to the eighth subnet in the following text, the above second model corresponds to the ninth subnet, the tenth subnet and the eleventh subnet in the cloud-side server in the following text, where the first sub-model in the second model corresponds to the ninth subnet, and the second sub-model corresponds to the tenth subnet and the eleventh subnet; the replicas of the second model correspond to the second subnet, the third subnet and the fourth subnet in the following text, and the above second model also corresponds to the fifth subnet, the sixth subnet and the seventh subnet of the edge-side device in the following text, where the first sub-model in the second model corresponds to the fifth subnet, and the second sub-model corresponds to the sixth subnet and the seventh subnet; the above third model corresponds to the first subnet in the following text.

[0080] It should be further noted that the above server corresponds to the cloud-side server in the following text, and the above client corresponds to the edge-side device in the following text.

[0081] As Figure 1 shown, the embodiment of the present application provides a framework schematic diagram of a model inference and training system. The system includes a cloud-side model and an edge-side model.

[0082] The cloud - side model extracts feature information from the user features uploaded by the edge - side model and inputs it into the cloud - side model. Combining the model and the user's label information uploaded by the edge - side model, it trains the cloud - side model, updates the model parameters of the cloud - side model sent to the edge - side and deployed on the cloud - side, and periodically sends the updated model parameters to the edge - side model; the edge - side model trains based on the sub - model on the edge - side and the model sent by the cloud - side, and updates the sub - model on the edge - side.

[0083] As Figure 2 shown, the embodiment of the present application provides a schematic diagram of a device for a model inference and training system. The device includes a cloud - side model and an edge - side model. Among them, the cloud - side model includes a model generation module, a training module, an aggregation module, an inference module, and a communication module. The cloud - side model is used for model inference and training. The edge - side model includes a training module, an inference module, and a communication module, and the edge - side model is used for model inference and training.

[0084] Among them, the model generation module of the cloud - side model is used to generate the same model with the parameter generation of all sub - models of the cloud - side model and deploy it to the cloud - side model. The communication module of the cloud - side model is used to receive the edge - side sub - model parameters uploaded by the communication module of the edge - side model and send them to the training module of the cloud - side model. At the same time, it sends the sub - model parameters of the training module of the cloud - side model to the communication module of the edge - side, and is also responsible for sending the vector generated by the inference module to the communication module of the edge - side model. The aggregation module of the cloud - side model aggregates the model parameters uploaded by multiple edge - side models received, and then sends the processed model parameters uploaded by multiple edge - side models to the training module and the inference module of the cloud - side model. The training module of the cloud - side model is used to train the input feature information and update the sub - model parameters of the cloud - side model. The inference module of the cloud - side model is used to infer the label information input by the edge - side model.

[0085] The communication module of the edge - side model is used to receive the sub - model parameters sent by the communication module of the cloud - side model and send them to the training module of the edge - side model. At the same time, it receives the vector generated by the inference module of the cloud - side model sent by the communication module of the cloud - side model, and sends it to the inference module and the training module of the edge - side model, and also uploads the sub - model parameters of the edge - side model to the communication device of the cloud - side model. The training module of the edge - side model is used to train the vector generated by the inference of the cloud - side model transmitted by the communication module of the edge - side model and update the sub - model parameters of the edge - side model. The inference module of the edge - side is used to infer the vector generated by the inference of the cloud - side model transmitted by the communication module of the edge - side model.

[0086] As Figure 3As shown in the figure, it is a schematic structural diagram of an edge device provided by an embodiment of the present application. An edge device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0087] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the edge device. In other embodiments of the present application, the edge device may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0088] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0089] The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.

[0090] A memory can also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can hold the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the said memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0091] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0092] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple groups of I2C buses. The processor 110 can be respectively coupled to the touch sensor 180K, the charger, the flash, the camera 193, etc. through different I2C bus interfaces. For example: The processor 110 can be coupled to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface to implement the touch function of the end-side device.

[0093] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple groups of I2S buses. The processor 110 can be coupled to the audio module 170 through the I2S bus to implement communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit an audio signal to the wireless communication module 160 through the I2S interface to implement the function of answering a phone call through a Bluetooth headset.

[0094] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled through the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 through the PCM interface to implement the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0095] The UART interface is a general-purpose serial data bus for asynchronous communication. This bus can be a two-way communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 through the UART interface to implement the Bluetooth function. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 through the UART interface to implement the function of playing music through a Bluetooth headset.

[0096] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 110 and the camera 193 communicate through the CSI interface to implement the shooting function of the end-side device. The processor 110 and the display screen 194 communicate through the DSI interface to implement the display function of the end-side device.

[0097] The GPIO interface can be configured through software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0098] The USB interface 130 is an interface that conforms to the USB standard specification, and can specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the end-side device, and can also be used for data transmission between the end-side device and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as AR devices, etc.

[0099] It should be understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are only illustrative descriptions and do not constitute a structural limitation on the terminal device. In other embodiments of the present application, the terminal device may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0100] The wireless communication function of the terminal device can be implemented by antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modulation and demodulation processor, baseband processor, etc.

[0101] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the terminal device can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: Antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0102] The mobile communication module 150 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the terminal device. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves through antenna 1, filter, amplify, etc. the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 150 may be provided in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be provided in the same device.

[0103] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to speaker 170A, receiver 170B, etc.), or displays an image or video through the display screen 194. In some embodiments, the modulation and demodulation processor may be an independent device. In other embodiments, the modulation and demodulation processor may be independent of the processor 110 and be provided in the same device as the mobile communication module 150 or other functional modules.

[0104] In some embodiments, the edge device can establish a communication connection with the operator's mobile communication network through the mobile communication module 105 and access the Internet through the mobile communication network. For example, the edge device communicates with the game server 104 through the mobile communication module 105, communicates with the live broadcast server 101 through the mobile communication module 105, and can even perform peer-to-peer communication with the edge device through the mobile communication module 105. For another example, the edge device can communicate with the live broadcast server 101 through the mobile communication module 105.

[0105] The wireless communication module 160 can provide solutions for wireless communications applied to the edge device, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signals to be sent from the processor 110, perform frequency modulation on them, amplify them, and convert them into electromagnetic waves through the antenna 2 for radiation.

[0106] In some embodiments, the edge device can establish a communication connection with the Internet through the wireless communication module 160 to access the Internet. For example, the edge device communicates with the server 101 through the wireless communication module 160, and can also perform peer-to-peer communication with the edge device through the wireless communication module 160. For another example, the edge device can communicate with the server 101 through the wireless communication module 160.

[0107] The NPU is a neural-network (NN) computing processor. By learning from the biological neural network structure, for example, learning from the transmission mode between human brain neurons, it can quickly process the input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the edge device can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.

[0108] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the terminal device. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.

[0109] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the terminal device (such as audio data, phone book, etc.). In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the terminal device by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.

[0110] As Figure 4a shown, it is a model recommendation system diagram provided by this embodiment. The recommendation system includes a server and a client. The first model is deployed on the server, and the second model is deployed on the client. The client sends a product recommendation request to the server. The server obtains multiple recommended product results based on the product recommendation request sent by the client, and obtains a first vector through the first model based on the recommendation request. The first vector is input into the client, and the client inputs the first vector into the second model to obtain a second vector. The client is used to evaluate multiple first recommendation results from the server based on the second vector to output at least one second recommendation result, and sort and display the second recommendation result.

[0111] As Figure 4bAs shown in the figure, it is a diagram of a model training system provided by this embodiment. The model training system is used to train the cloud-side model and the edge-side model. The model training system consists of the first subnet, the eighth subnet, the second subnet, the ninth subnet, the third subnet, the tenth subnet, the fourth subnet, and the eleventh subnet of the cloud-side model, and the fifth subnet, the sixth subnet, and the seventh subnet of the edge-side model. The first subnet of the cloud-side model is obtained by initializing the cloud-side model. The second subnet of the cloud-side model corresponds to the fifth subnet of the edge-side model, and is generated and uploaded from the parameters of the fifth subnet of the edge-side model. In the subsequent training phase, the edge-side model uploads the updated parameters of the fifth subnet to update the ninth subnet. The third subnet and the fourth subnet of the cloud-side model are obtained by initializing and training the cloud-side model in the cold start phase. The eighth subnet, the ninth subnet, the tenth subnet, and the eleventh subnet of the cloud-side model are respectively generated from the first subnet, the second subnet, the third subnet, and the fourth subnet of the cloud-side model in the initialization phase. The fifth subnet of the edge-side model is obtained by initializing the edge-side model in the cold start phase, and the parameters are updated based on the training of the edge-side model in each subsequent training. In the subsequent training, the edge-side model regularly uploads the parameters of the fifth subnet to the cloud-side model. The sixth subnet of the edge-side model corresponds to the third subnet of the cloud-side model, and the seventh subnet corresponds to the fourth subnet of the cloud-side model, and is obtained by the cloud-side model sending the parameters of the third subnet and the fourth subnet to the edge-side model. In the subsequent training phase, after the cloud-side model updates the parameters of the tenth subnet and the eleventh subnet, it regularly updates the sixth subnet and the seventh subnet of the edge-side model.

[0112] As Figure 5 shown in the figure, an embodiment of the present application provides a diagram of a model inference system. The model inference module includes the first subnet of the cloud-side model and the fifth subnet, the sixth subnet, and the seventh subnet of the edge-side model. The first subnet is generated in the initial stage. The fifth subnet is generated by the edge-side model in the initialization stage. The sixth subnet and the seventh subnet are respectively generated by sending the parameters of the third subnet and the fourth subnet generated by the cloud-side model in the initialization stage to the edge-side model. In the inference process, the vector inferred by the first subnet of the cloud-side model is input to the edge-side model for inference.

[0113] As Figure 6aAs shown in the figure, a schematic diagram of a model training system provided by an embodiment of the present application. The model training system includes a server and a client. The server is deployed with a model system. The model system includes a first model, a second model, a third model, and a copy of the second model. The output of the first model corresponds to the input of the second model. The third model corresponds to the historical state of the first model. The first model and the third model have the same input. The output of the third model corresponds to the input of the copy. The client is used to send data and model parameters to the server, and the data indicates the behavior of the user. The server is used to receive the data and model parameters from the client. The server is used to update the second model based on the model parameters and keep the copy unchanged. The server is used to calculate a first loss using the data, the first model, and the second model. The server is used to calculate a second loss using the third model and the copy. The server is used to update the model system based on the first loss and the second loss to obtain the trained first model.

[0114] As Figure 6b shown in the figure, a schematic diagram of a cloud-side model training process provided by an embodiment of the present application. During the training process of the cloud-side model, the cloud-side model aggregates the fifth subnet parameters after less-parameter fine-tuning sent by the edge-side model, the feature information obtained by feature extraction of user features, and the label information. The cloud-side model inputs the feature information into the eighth subnet for training to update the parameters of the eighth subnet, inputs the trained vector into the ninth subnet and the tenth subnet, and together generates a training vector for output. During this process, the parameters of the tenth subnet are updated. The output vector is input into the eleventh subnet for further training, and the trained vector is input into the loss. At the same time, the cloud-side model inputs the label into the loss1. The trained vector is simultaneously input into the continuous learning loss1. Similarly, during the above training process, the cloud-side model synchronously inputs the feature information into the first subnet and subsequent subnets, and performs the same steps, and inputs the finally trained vector into the continuous learning loss.

[0115] As Figure 7 shown in the figure, it is a training system diagram of an edge-side model provided by this embodiment.

[0116] The edge-side model reports the user request to the cloud-side model. The cloud-side model performs rough recall and then completes feature extraction, and performs inference through the eighth subnet. The vector after the inference of the eighth subnet and the corresponding information obtained after rough recall are sent to the edge-side model. The edge-side model inputs the sent vector into the fifth subnet and the sixth subnet for inference, and inputs the inference vectors generated by the fifth subnet and the sixth subnet into the seventh subnet for inference again. The seventh subnet outputs the vector after inference.

[0117] In a possible implementation, during the model inference process, the user opens the APP of the edge model and requests recommended products. When the user makes a request, after the edge model obtains the relevant information of the user, it will recall products based on the user's information. At the same time, the edge model sends the relevant information of the user to the cloud model. After the cloud model extracts features, it sends the feature information to the first subnet. After inference, it generates recommended product information and vectors. The cloud model sends the vector to the edge model. The edge model inputs the vector into the fourth model and the fifth subnet. After inference by the fourth model and the fifth subnet, it generates the parameters of the fifth subnet after inference. The edge side scores the products, and finally sorts and displays the recommended products in the APP.

[0118] As Figure 8 shown, it is a model inference system diagram provided by this embodiment. During the operation of the inference system, the edge model reports the user request to the cloud model. The cloud model completes feature extraction after rough recall and performs inference through the first subnet. The vector after the first subnet's inference and the corresponding information obtained after rough recall are sent to the edge model. The edge model inputs the sent vector into the fifth subnet and the sixth subnet for inference, and inputs the inference vectors generated by the fifth subnet and the sixth subnet into the seventh subnet for inference again. The seventh subnet outputs the inferred vector, which can be output to the application of the edge model (such as APP) for display.

[0119] As Figure 9 shown, it is an implementation form diagram of a model inference and training method provided by an embodiment of this application. The location of the product implementation form of the present invention includes, but is not limited to, the following situations. It can be included in edge models such as mobile APPs and PDAs, or in the cloud deep learning platform software; the form of existence of the product can be program code deployed on the edge model, or program code or model parameters on the cloud model hardware. On the edge model, the training code for Delta-tuning is stored for training the edge sub-model, and there is also the inference code for some parameters for inferring the edge model. On the cloud side, there is also training code and relevant inference code for inferring and training the cloud model.

[0120] When the cloud model is trained, the model parameters on the cloud side are manifested as a piece of training code. The training code is input into the model training module. During the training process, the training code updates the parameters in the training module. After updating the parameters of some sub-models on the cloud side, the cloud sub-model performs inference. At this time, the inference code is deployed to the inference module on the cloud side for inference.

[0121] When the edge - side model is trained, the model parameters on the edge - side are represented as a piece of training code. The training code is input into the model training module. During the training process, the training code updates the parameters in the training module. After updating the parameters of some sub - models on the edge - side, the edge - side sub - model performs inference. At this time, the inference code is deployed to the inference module on the edge - side for inference. At the same time, the code of some model parameters of the cloud - side server is regularly sent to the edge - side device to update the code of the model parameters on the edge - side device.

[0122] As Figure 10 shown, it is a core flowchart of model training provided by an embodiment of the present application. During the training process of the cloud - side model, the cloud - side model aggregates the fifth subnet parameters after few - parameter fine - tuning sent by the edge - side model, the feature information after feature extraction of user features, and the label information. The cloud - side model inputs the feature information into the eighth subnet for training to update the parameters of the eighth subnet. At the same time, training vectors are generated. The vectors after training the eighth subnet are input into the ninth subnet and the tenth subnet for training and then output. During this process, the parameters of the tenth subnet are updated. The vectors after training the ninth subnet and the tenth subnet are input into the eleventh subnet for further training. After training the eleventh subnet, the vectors are input into loss. At the same time, the cloud - side model inputs the label into loss to converge together, thereby optimizing the training. The vectors after training the eleventh subnet are also input into the continual learning loss. Similarly, during the above - mentioned training process, the cloud - side model synchronously inputs the feature information into the first subnet and performs the same steps. Finally, corresponding training vectors are generated through the fourth subnet. The vectors after training the fourth subnet are input into the continual learning loss and not into loss. For the output results after passing the two vectors after training the pre - and post - generation models through the continual learning loss and the results after convergence after passing through loss, they are together input into the final given loss, and the Distillation Loss is used to calculate the difference between the two distributions to obtain the newly trained model.

[0123] During the training process of the edge - side model, the edge - side model inputs the vectors sent by the cloud - side model after inference into the fifth subnet and the sixth subnet of the edge - side model for training. At the same time, the parameters of the fifth subnet are updated. The generated vectors after training are together input into the seventh subnet for further training. The vectors after training are input into the loss of the edge - side model and converge with the label input into the edge - side model to optimize the model parameters.

[0124] As Figure 11As shown in the figure, it is a core flowchart of model inference provided by an embodiment of the present application. The edge model reports the user request to the cloud model. After the cloud model performs rough recall, it completes feature extraction and performs inference through the eighth subnet or the first subnet. The vectors after inference by the eighth subnet or the first subnet and the corresponding information obtained after rough recall are sent to the edge model. The edge model inputs the sent vectors into the fifth subnet and the sixth subnet for inference, and inputs the inference vectors generated by the fifth subnet and the sixth subnet into the seventh subnet for inference again. The seventh subnet outputs the vectors after inference, which can be output to the application (such as APP) of the edge model for display.

[0125] In a possible inference method, for example, in a commodity recommendation scenario, when the user opens the APP home page, the edge model will request the cloud model to obtain a commodity list for display. The edge model will encapsulate relevant information into the cloud model. Here, the relevant information includes but is not limited to the user's ID information, geographical location information, current Wi-Fi information, and time. After the edge model obtains the relevant information of the user, it will perform commodity recall based on the user's information, such as commodities that the user is interested in, commodities of the type that the user is interested in, etc. Some attribute information of the commodities will also be obtained when recalling the commodities.

[0126] When the first model of the cloud model performs inference, the trained model is accelerated and deployed on the server. In the case where the edge model transmits the user's relevant information, the relevant confidence of the user is extracted and input into the first subnet for inference to generate an inference vector and the corresponding commodity information, which is compressed and encrypted by the output vector compression and encryption module and then transmitted to the edge model. When the edge model performs inference, the edge model inputs the obtained vector into the fifth subnet and the sixth subnet of the edge model for inference, and sends the inference results after inference by the fifth subnet and the sixth subnet to the seventh subnet for inference again. Finally, the inference result is output to score the commodities, and finally the commodities are sorted and displayed.

[0127] As Figure 12 shown in the figure, it is a flowchart of the edge model for performing few-parameter fine-tuning provided by an embodiment of the present application.

[0128] In the case of a cloud-side model and an edge-side model, a large model consists of a first cloud-side model, a second cloud-side model, and an edge-side model. The first cloud-side model or the second cloud-side model consists of a base subnet generated by initialization, a subnet uploaded from the edge side, and two subnets sent from the cloud side to the edge-side model and deployed to the cloud side. The edge-side model consists of a subnet generated by edge-side initialization and two subnets sent from the cloud side to the edge side and deployed to the cloud side. In the case of a cloud-side model and multiple edge-side models, a large model consists of a first cloud-side model, a second cloud-side model, and multiple edge-side models. The first cloud-side model or the second cloud-side model consists of a base subnet generated by initialization, a subnet uploaded from the edge side, and two subnets sent from the cloud side to multiple edge-side models and deployed to the cloud side. The multiple edge-side models consist of models generated by each edge-side model initialization and two subnets sent from the cloud-side model to each edge-side model and deployed to the cloud side.

[0129] In this embodiment, only the process of real-time training of the edge-side model is described.

[0130] The subnet generated by the initialization of the edge-side model in this embodiment can be a Detal Objects network or multiple Detal Object network layers. The edge-side model is based on the vector after the inference of the cloud-side model obtained from the cloud-side model, combined with the label generated by the edge-side model according to the user request, and uses the delta-tuning technology to train the Detal Objects network.

[0131] The specific training process is as follows. The edge-side model inputs the vector sent from the cloud side into the edge-side subnet and the first subnet sent from the cloud side to the edge-side model at the same time. After training respectively, the vectors after their respective training are input into the second subnet sent from the cloud side to the edge-side model for further training, and finally the vectors after training are input into the loss module. The loss module converges the input vectors, thereby optimizing the training.

[0132] During the training process of the edge-side model, only the DeltaObject network is trained and its parameters are updated. The first subnet and the second subnet sent from the cloud-side model to the edge-side model participate in the training but do not update their parameters.

[0133] Specifically, the parameters are generated by the cloud-side model and sent to the edge-side model when the user first downloads the APP. During the subsequent training process, the edge-side model sends the parameters of the fifth subnet to the ninth subnet of the cloud-side model for update and training, and then updates the parameters. The cloud-side model sends the updated parameters of the tenth and eleventh subnets to the sixth and seventh subnets of the edge-side model. The edge-side model updates the sixth and seventh subnets of the edge-side model with the updated parameters of the tenth and eleventh subnets sent down, and then the edge-side model performs inference. During the training process of the edge-side model, the parameters of the sixth and seventh subnets are not updated. The edge-side model only updates the parameters of the fifth subnet based on the vectors generated by the inference of the cloud-side model input, so as to realize the adjustment of only a part of the parameters of the fifth subnet of the edge-side model, that is, the small-parameter fine-tuning of the edge-side model, but the training effect can reach the effect of full-parameter adjustment.

[0134] After the edge side sends the parameters of the fifth subnet to the ninth subnet of the cloud-side model for parameter update during the initialization phase, the cloud-side model trains based on this, combined with the eighth, tenth, and eleventh subnets of the cloud-side model. During the subsequent training phase, the edge-side model regularly uploads the parameters of the fifth subnet after any update to the ninth subnet of the cloud-side model, so that the ninth subnet of the cloud-side model is updated regularly.

[0135] Before the regular update, the cloud-side model completes multiple trainings of the model based on the ninth subnet after any parameter update, without relying on the real-time upload of the parameters of the fifth subnet of the edge-side model to the cloud-side model for update and then training, realizing the independent training of the cloud-side model and the edge-side model.

[0136] As shown in Table 1, the results of training four different training datasets using different low-parameter fine-tuning methods based on MLM and Distil are presented. Adapter, Prompt, and LoRa in the table are different low-parameter fine-tuning methods, namely, methods of Delta tuning. WIKI, CS, BIOMED, and REALNEWS are four different training datasets. Model in the table represents two model training methods, MLM and Adapter. MLM is a method of directly training based on the base model, which is a commonly used model training method currently, that is, only training the model based on a single base model, different from Disti used in this application. MLM does not generate the same model and deploy it to the cloud-side model, and then train the models before and after generation respectively to obtain the vectors of the two trainings, and then use the continual learning loss for convergence. The data in the table represents the training effect of the model, and the larger the value, the better the training effect. It can be seen that under the exemplary low-parameter fine-tuning methods of Adapter, Prompt, and LoRa, the data of model training using Distil is significantly larger than that using MLM, that is, the effect of model training using Distil is better than that using MLM.

[0137] Table 1

[0138]

[0139] As Figure 13 shown, it is a flowchart of training a cloud-side model based on the continual learning method provided by an embodiment of this application. In this embodiment, in the cloud-side model, a first cloud-side model is generated in the initialization stage. This cloud-side model is jointly composed of a base subnet, a subnet uploaded from the edge side to the cloud side, and two subnets sent from the cloud side to the edge side. Here, the subnet uploaded from the edge side to the cloud side is a subnet generated by the edge-side model in the initialization stage. After the edge-side model is initialized, generated, and deployed to the edge-side model, it is simultaneously generated and uploaded to the cloud-side model. Here, the subnets sent from the cloud side to the edge side are that while the cloud-side service sends the two subnets generated in the initialization to the edge side, the cloud-side model also deploys these two subnets in the cloud-side model. The parameters of the two subnets are the same. For the convenience of subsequent expression, this model is described as two subnets sent from the cloud side to the edge side and deployed to the cloud side. To distinguish these two subnets, they are described as the first subnet sent from the cloud side to the edge side and deployed to the cloud side and the second subnet sent from the cloud side to the edge side and deployed to the cloud side.

[0140] In the preparation phase before training, the cloud-side model generates an identical model to the first cloud-side model and deploys it on the cloud-side model. In the training phase, the cloud-side model inputs the feature information extracted by the feature extraction module to the first cloud-side model and the generated second cloud-side model for training.

[0141] During training, the training process of the first cloud-side model is described as an example. The base subnet of the first cloud-side model first trains the feature information to generate a trained vector. The cloud-side model inputs the trained vector to the subnet uploaded by the end side and the first subnet sent from the cloud side to the end side and deployed to the cloud side for training. The trained vector of the subnet uploaded by the end side and the trained vector of the first subnet sent from the cloud side to the end side and deployed to the cloud side are simultaneously input to the second subnet sent from the cloud side to the end side and deployed to the cloud side for training. After the second subnet sent from the cloud side to the end side and deployed to the cloud side is trained, the trained vector is output to the continuous learning loss. Continuous learning loss

[0142] Similarly, when the cloud-side model inputs the feature information to the second cloud-side model for the same training steps, it also outputs a vector. Unlike the first cloud-side model, the vector finally output by the second cloud-side model training is not only input into the continuous learning loss, but also into the loss.

[0143] It should be noted that during the training process, the base model parameters in the first cloud-side model are fixed, that is, only the feature information input to the first cloud-side model is trained and its own parameters are not updated. The base model parameters in the second cloud-side model can be learned, that is, while the feature information input to the second cloud-side model is trained, its own parameters are also updated.

[0144] After generating the final output vectors of the two models after training, both are input into the continuous learning loss. The result after convergence using the continuous learning loss is input into the final loss together with the result after loss convergence. Distillation Loss is used to calculate the difference between the two distributions to obtain the latest trained model.

[0145] In order to enable the model to continuously learn and train, in one possible implementation, the cloud-side model will use the Distillation Loss to obtain the latest trained model through continuous learning loss, which will be used as the new first cloud-side model before the next training. The subsequent training process repeats the previous training process, thereby achieving continuous training of the cloud-side model based on the continuous learning method.

[0146] In this way, without changing the model parameters of the end-side model uploaded to the cloud-side model, only the parameters of the cloud-side model are continuously updated, and the overall effect does not change significantly compared with the effect of the cloud-side model relying on the end-side to update parameters in real time, realizing the separate training of the cloud-side model.

Claims

1. A recommendation system, the recommendation system comprising a server and a client, the server being deployed with a first model, and the client being deployed with a second model, characterized in that: The client is configured to send a recommendation request to the server; The server is configured to obtain a plurality of first recommendation results based on the recommendation request; The server is configured to obtain a first vector through the first model based on the recommendation request; The client is configured to input the first vector from the server into the second model to obtain a second vector; The client is configured to evaluate the plurality of first recommendation results from the server based on the second vector to output at least one second recommendation result.

2. The system according to claim 1, characterized in that, The second model includes a first sub-model and a second sub-model, the first sub-model is trained on the client, and the second sub-model is trained on the server.

3. The system according to claim 2, wherein The client is further configured to train the first sub-model based on the user's feedback on the at least one second recommendation result.

4. The system according to claim 1, wherein The recommendation request includes at least one of the following parameters: user identification, geographical location information, Wi-Fi information, time information.

5. A recommendation method, the recommendation method being applied to a client, the client being deployed with a second model, characterized in that The client is configured to send a recommendation request to the server; The client is configured to input the first vector from the server into the second model to obtain a second vector; The client is configured to evaluate the plurality of first recommendation results from the server based on the second vector to output at least one second recommendation result.

6. The method according to claim 5, wherein The second model includes a first sub-model and a second sub-model, the first sub-model is trained on the client, and the second sub-model is trained on the server.

7. According to the method described in 6, it is characterized in that, The client is further configured to train the first sub-model based on the user's feedback on the at least one second recommendation result.

8. The system according to claim 1, wherein The recommendation request includes at least one of the following parameters: user identification, geographical location information, Wi-Fi information, time information.

9. A recommendation method, the recommendation method being applied to a server, the server being deployed with a first model, characterized in that The server is configured to obtain a plurality of first recommendation results based on the recommendation request; The server is configured to obtain a first vector through the first model based on the recommendation request.

10. A model training system, the model training system comprising a server and a client, characterized in that The server is deployed with a model system, the model system including a first model, a second model, a third model, and a copy of the second model, the output of the first model corresponding to the input of the second model, the third model corresponding to the historical state of the first model, the first model and the third model having the same input, and the output of the third model corresponding to the input of the copy; The client is configured to send data and model parameters to the server, the data indicating the user's behavior; The server is configured to receive the data and the model parameters from the client; The server is used to update the second model based on the model parameters and keep the copy unchanged; The server is used to calculate a first loss using the data, the first model, and the second model; The server is used to calculate a second loss using the third model and the copy; The server is used to update the model system based on the first loss and the second loss to obtain the trained first model.

11. The system according to claim 10, wherein The second model includes a first sub-model and a second sub-model. The first sub-model is trained on the client side, and the second sub-model is trained on the server side.

12. The system according to claim 10, wherein The server is further used to obtain a first vector through the first model based on a recommendation request; the server is further used to send the first vector.

13. The system according to claim 10, wherein The client is used to train the second model based on the first vector and the data of the client, so as to update the first sub-model. The data of the client indicates the behavior of the user.

14. The system according to claim 10, wherein The server is used to update the model system based on the first loss and the second loss, and is also used to obtain the trained second sub-model.

15. The system according to claim 14, wherein The client is further used to update the second model based on the trained second sub-model from the server.

16. A training method, which is applied to a server, and is characterized in that The server deploys a model system, and the model system includes a first model, a second model, a third model, and a copy of the second model. The output of the first model corresponds to the input of the second model. The third model corresponds to the historical state of the first model. The first model and the third model have the same input. The output of the third model corresponds to the input of the copy; The server is used to receive the data and the model parameters from the client; The server is used to update the second model based on the model parameters and keep the copy unchanged; The server is used to calculate a first loss using the data, the first model, and the second model; The server is used to calculate a second loss using the third model and the copy; The server is used to update the model system based on the first loss and the second loss to obtain the trained first model.

17. The system according to claim 16, wherein The second model includes a first sub-model and a second sub-model. The first sub-model is trained on the client side, and the second sub-model is trained on the server side.

18. The system according to claim 16, wherein The server is further used to obtain a first vector through the first model based on a recommendation request; the server is further used to send the first vector.

19. The system according to claim 16, wherein The server is used to update the model system based on the first loss and the second loss, and is also used to obtain the trained second sub-model.

20. A model training device, which is deployed on a server. The device includes a receiving module, a storage module, a sending module, and a training module, and is characterized in that The receiving module is used to receive the data and the model parameters from the client; The storage module is used to deploy a model system; the model system includes a first model, a second model, a third model, and a copy of the second model; The training module is used to calculate a first loss using the data, the first model, and the second model; calculate a second loss using the third model and the copy; update the model system based on the first loss and the second loss to obtain the trained first model; update the model system based on the first loss and the second loss to obtain the trained second sub-model; The sending module is used to send the first vector.

21. A model inference device, the model inference device is deployed on the server side, and the device includes a receiving and sending block, a storage module, and an inference module, wherein The sending module is used to send a recommendation request to the server side; The storage module is used to deploy a first model and a second model; The inference module is used to obtain a plurality of first recommendation results based on the recommendation request; is further used to obtain a first vector through the first model based on the recommendation request; is further used to input the first vector from the server side into the second model to obtain a second vector; is further used to evaluate the plurality of first recommendation results from the server side based on the second vector to output at least one second recommendation result.