Execution method and device of remote server model task, storage medium and electronic equipment

By receiving model training requests on the local terminal and using the model calling interface on the remote server to perform model training tasks, the problem of low training efficiency of remote server model is solved, and an efficient model training process is realized.

CN120045348APending Publication Date: 2025-05-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510021942.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, remote server model training efficiency is low, and users need to build a training environment for models on their own on the local terminal, resulting in low training efficiency.

Method used

By receiving a model training request initiated by the training account in the local terminal, filtering out the corresponding reference model calling interface from the model calling interface on the remote server, sending the target training data set and instructions to the interface, performing the model training task, and returning the result to the local terminal.

Benefits of technology

It realizes efficient model training without building a training environment on the local terminal, and improves the efficiency of remote server model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045348A_ABST
    Figure CN120045348A_ABST
Patent Text Reader

Abstract

The invention discloses an execution method and device of a remote server model task, a storage medium and electronic equipment. The execution method of the remote server model task comprises the steps of receiving a model training request initiated by a training account; in response to the model training request, screening out a reference model calling interface corresponding to the reference model from a plurality of model calling interfaces on a remote server; sending a target calling request carrying the target training data set and the target training instruction to a reference model calling interface; and receiving a target training result returned by the model calling interface in response to the target calling request as an execution result of the model training task, by adopting the technical scheme, the problems of low training efficiency of the remote server model and the like in related technologies are solved, and the technical effect of improving the training efficiency of the remote server model is further achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computers, and more specifically, to a method and device for executing a remote server model task, a storage medium, and an electronic device. Background Art

[0002] In the related art, the training and inference tasks of an Artificial Intelligence (AI) model usually need to be completed in a server environment. Specifically, a user needs to prepare a specific model and a dataset to be used on a local terminal. Then, directly run the AI model training program on the local terminal. This program reads the dataset stored locally and learns and processes the input dataset according to the algorithm structure of the AI model, and finally completes the training of the AI model so that the AI model has specific inference capabilities.

[0003] This traditional training method has some limitations. First, the efficiency of AI model training highly depends on the computing resources of the local terminal. If the local terminal is busy or unavailable, the user cannot perform model training in a timely manner. Second, if a user wants to train a model that is not deployed on the local terminal on their own dataset, they need to build a model training environment on the local terminal by themselves. This process is not only complex but also affects the efficiency of model training.

[0004] In view of the problems such as low efficiency of remote server model training in the related art, no effective solution has been proposed yet. Summary of the Invention

[0005] The embodiments of the present application provide a method and device for executing a remote server model task, a storage medium, and an electronic device, so as to at least solve the problems such as low efficiency of remote server model training in the related art.

[0006] According to an embodiment of the embodiments of the present application, a method for executing a remote server model task is provided, which is applied to a local terminal. The target training dataset of a reference model is stored in the local terminal, and multiple models are deployed on a remote server. The remote server has generated at least one model call interface for each model in advance, and the multiple models include the reference model. The method includes:

[0007] Receiving a model training request initiated by a training account, where the model training request is used to request to perform a model training task on the reference model;

[0008] Responding to the model training request, and screening out a reference model call interface corresponding to the reference model from multiple model call interfaces on the remote server;

[0009] Send a target call request carrying the target training dataset and the target training instruction to the reference model call interface, where the target call request is used to request the reference model call interface to call the reference model to execute a model training task according to the target training dataset and the target training instruction, and return the training result of the model training task to the local terminal;

[0010] Receive the target training result returned by the model call interface in response to the target call request as the execution result of the model training task.

[0011] Optionally, sending a target call request carrying the target training dataset and the target training instruction to the reference model call interface includes:

[0012] Obtain the target interface address corresponding to the reference model call interface;

[0013] Add the target training dataset and the target training instruction to the initial call request to obtain the target call request;

[0014] Send the target call request to the reference model call interface according to the target interface address.

[0015] Optionally, when the reference model is a vision-language model, before adding the target training dataset and the target training instruction to the initial call request to obtain the target call request, the method further includes:

[0016] Obtain an image caption dataset, a visual question answering dataset, and a visual grounding dataset, and determine at least one of the image caption dataset, the visual question answering dataset, and the visual grounding dataset as the target training dataset, where the image caption dataset includes multiple images and descriptive text for each image, and the image caption dataset is used to train the performance of the vision-language model in image description or caption generation tasks, the visual question answering dataset includes multiple images and questions and corresponding answers for the content of each image, and the visual question answering dataset is used to train the performance of the vision-language model in visual question answering tasks, and the visual grounding dataset includes multiple images and descriptive text for local regions of each image, and the visual grounding dataset is used to train the performance of the vision-language model to recognize the content of local regions of images;

[0017] Obtain the reference model parameters corresponding to the reference model; generate a parameter adjustment instruction corresponding to the reference model parameters, where the parameter adjustment instruction is used to indicate adjusting the model parameters of the reference model to the reference model parameters; determine the parameter adjustment instruction as the target training instruction.

[0018] Optionally, the reference model call interface calls the reference model to execute the model training task through the following steps:

[0019] Adjust the model parameters of the reference model to the reference model parameters according to the parameter adjustment instruction to obtain the reference model to be trained;

[0020] Perform multiple rounds of training on the reference model to be trained through the image caption dataset, visual question answering dataset, and visual foundation dataset until the model parameters of the reference model to be trained converge, and obtain the trained reference model;

[0021] Return the training result indicating the successful training of the model training task to the local terminal.

[0022] Optionally, before screening out the reference model call interface corresponding to the reference model from multiple model call interfaces on the remote server, the method further includes:

[0023] The remote server obtains the terminal sending parameter and terminal receiving parameter of the local terminal, as well as the model call parameter of the reference model, where the terminal sending parameter is used to indicate the data sending method of the local terminal, the terminal receiving parameter is used to indicate the data receiving method allowed by the local terminal, and the model call parameter is the parameter required to call the reference model;

[0024] The remote server generates a reference model call interface according to the terminal sending parameter, terminal receiving parameter, and model call parameter, where the reference model call interface is used to receive data that meets the data sending method indicated by the terminal sending parameter, call the reference model according to the model call parameter, and send data to the local terminal according to the data receiving method indicated by the terminal receiving parameter.

[0025] According to an embodiment of the present application, a method for executing a remote server model task is provided, which is applied to a local terminal. A plurality of models are deployed on the remote server, and the remote server has pre-generated at least one model call interface for each of the models. The plurality of models include a trained target model. The method includes:

[0026] Receive a model inference request initiated by an inference account, where the model inference request is used to request the target model to execute a model inference task;

[0027] In response to the model inference request, screen out the target model call interface corresponding to the target model from multiple model call interfaces on the remote server;

[0028] Send a target inference request carrying model input data and a target inference instruction to the target model call interface, where the target inference request is used to request the target model call interface to call the target model to execute a model inference task according to the model input data and the target inference instruction, and return the inference result of the model inference task to the local terminal;

[0029] Receive the target inference result returned by the receiving model call interface in response to the target inference request as the execution result of the model inference task.

[0030] Optionally, the target model call interface calls the target model to execute the model inference task through the following steps:

[0031] Input the model input data into the target model;

[0032] Control the target model to perform the inference operation indicated by the target inference instruction according to the input model input data;

[0033] Return the inference result output by the target model when performing the inference operation to the local terminal.

[0034] According to another embodiment of the embodiments of the present application, there is also provided an execution device for a remote server model task, which is applied to a local terminal. The target training data set of the reference model is stored in the local terminal, and multiple models are deployed on the remote server. The remote server has pre-generated at least one model call interface for each model. The multiple models include the reference model. The execution device for the remote server model task includes:

[0035] A first receiving module, configured to receive a model training request initiated by a training account, where the model training request is used to request to perform a model training task on the reference model;

[0036] A first response module, configured to respond to the model training request and filter out the reference model call interface corresponding to the reference model from multiple model call interfaces on the remote server;

[0037] A first sending module, configured to send a target call request carrying the target training data set and the target training instruction to the reference model call interface, where the target call request is used to request the reference model call interface to call the reference model to perform the model training task according to the target training data set and the target training instruction, and return the training result of the model training task to the local terminal;

[0038] A second receiving module, configured to receive the target training result returned by the model call interface in response to the target call request as the execution result of the model training task.

[0039] According to another embodiment of the embodiments of the present application, there is also provided another execution device for a remote server model task, which is applied to a local terminal. Multiple models are deployed on the remote server. The remote server has pre-generated at least one model call interface for each model. The multiple models include the trained target model. The other execution device for the remote server model task includes:

[0040] A third receiving module, configured to receive a model inference request initiated by an inference account, where the model inference request is used to request a target model to perform a model inference task;

[0041] A second response module, configured to respond to the model inference request and screen out a target model call interface corresponding to the target model from multiple model call interfaces on a remote server;

[0042] A second sending module, configured to send a target inference request carrying model input data and a target inference instruction to the target model call interface, where the target inference request is used to request the reference model call interface to call the target model to perform a model inference task according to the model input data and the target inference instruction, and return the inference result of the model inference task to a local terminal;

[0043] A fourth receiving module, configured to receive the target inference result returned by the model call interface in response to the target inference request as the execution result of the model inference task.

[0044] According to another embodiment of the present application, there is also provided a computer program product, including a computer program, where the computer program is executed by a processor to perform the steps in any one of the above method embodiments.

[0045] According to another embodiment of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0046] According to another embodiment of the present application, there is also provided an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to perform the steps in any one of the above method embodiments.

[0047] In an embodiment of the present application, a method for executing a remote server model task is proposed. The target training dataset of the reference model is stored in the local terminal, and multiple models are deployed on the remote server. The remote server has pre-generated at least one model call interface for each model. The multiple models include the reference model. The local terminal first receives a model training request initiated by a training account. Among them, the model training request can request to execute a model training task on the reference model. Then, in response to the model training request, the reference model call interface corresponding to the reference model is filtered out from the multiple model call interfaces on the remote server. Then, a target call request carrying the target training dataset and the target training instruction is sent to the reference model call interface. Among them, the target call request can request the reference model call interface to call the reference model to execute the model training task according to the target training dataset and the target training instruction, and return the training result of the model training task to the local terminal. Finally, the target training result returned by the model call interface in response to the target call request is received as the execution result of the model training task. That is, the remote server sets a model call interface for the model, and the model call interface can call the corresponding model to execute the model task. Through the above method, when the user needs to execute a training task on the model, even if the model and the dataset to be trained are deployed on different servers respectively, such as the model is deployed on the remote server and the dataset to be trained is stored in the local terminal, the user can directly realize the call of the model through the data interaction between the local terminal and the model call interface on the remote server, and then complete the model training task, and view the training result returned by the model call interface. By adopting the above technical solution, the problems of low model training efficiency in the related technology are solved, and the technical effect of improving the model training efficiency of the remote server is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a hardware structure block diagram of a computer device for a method for executing a remote server model task according to an embodiment of the present application;

[0049] Figure 2 is a flowchart of a method for executing a remote server model task according to an embodiment of the present application;

[0050] Figure 3 is a schematic diagram of the data exchange format in the communication between the local terminal and the remote server according to an embodiment of the present application;

[0051] Figure 4 is a flowchart of another method for executing a remote server model task according to an embodiment of the present application;

[0052] Figure 5 is a logical architecture diagram of the input of a vision language model according to an embodiment of the present application;

[0053] Figure 6 Schematic diagram of a Restful interface of a vision - language model according to an embodiment of the present application;

[0054] Figure 7 Schematic diagram of a method for executing a remote server model task according to an embodiment of the present application;

[0055] Figure 8 Block diagram of the structure of another apparatus for executing a remote server model task according to an embodiment of the present application;

[0056] Figure 9 Block diagram of the structure of an apparatus for executing a remote server model task according to an embodiment of the present application. Detailed implementation manners

[0057] In the following, embodiments of the present application will be described in detail with reference to the drawings and in conjunction with the embodiments.

[0058] It should be noted that the terms "first", "second", etc. in the description and claims of the present application and the above - mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence.

[0059] The method embodiments provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 1 Hardware structure block diagram of a computer device for a method of executing a remote server model task according to an embodiment of the present application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in Figure 1 a processor 102 (the processor 102 may include, but is not limited to, a processing device such as a micro - processor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above - mentioned server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that, Figure 1 the structure shown in Figure 1 is only schematic, and it does not limit the structure of the above - mentioned server device. For example, the server device may further include more or fewer components than those shown in

[0060] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the execution method of the remote server model task in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0061] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the server device. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0062] In this embodiment, an execution method for a remote server model task is provided, which is applied to a local terminal. The local terminal stores a target training data set of a reference model, and multiple models are deployed on a remote server. The remote server has pre-generated at least one model call interface for each model, and the multiple models include a reference model. Figure 2 It is a flowchart of an execution method for a remote server model task according to an embodiment of the present application, as Figure 2 shown. The process includes the following steps:

[0063] Step S12, receiving a model training request initiated by a training account, where the model training request is used to request to perform a model training task on the reference model;

[0064] Step S14, in response to the model training request, screening out the reference model call interface corresponding to the reference model from multiple model call interfaces on the remote server;

[0065] Step S16, sending a target call request carrying a target training data set and a target training instruction to the reference model call interface, wherein the target call request is used to request the reference model call interface to call the reference model to perform a model training task according to the target training data set and the target training instruction, and return the training result of the model training task to the local terminal;

[0066] Step S18, receiving the target training result returned by the model calling interface in response to the target calling request as the execution result of the model training task.

[0067] Optionally, in this embodiment, the server model task may be, but is not limited to, operations performed on the AI ​​model, including training, reasoning, model evaluation, and hyperparameter tuning.

[0068] Optionally, in this embodiment, the local terminal is connected to the remote server via a network. The local terminal is a terminal that can be directly operated by the user, and can be but not limited to a local server. The local terminal stores a data set (i.e., a target training data set) of the model (i.e., a reference model) for which the user wants to perform a training task. The remote server can be but is not limited to a server that can perform server model tasks, and multiple AI models are deployed on the remote server, that is, a training environment for each AI model is built, and one or more model call interfaces are generated in advance for each AI model, wherein the model call interface can be but is not limited to an interface that allows the user to interact with the model on the remote server, including a Restful API interface based on the HTTP protocol and a WebSocket interface that provides a full-duplex communication mechanism.

[0069] Optionally, in this embodiment, the user initiates a model training request through a training account on a local terminal. The request specifies a reference model that the user wants to train, such as a visual language model (model A), and the type of task that the user wants to perform, such as an image caption training task.

[0070] Optionally, in this embodiment, when the local terminal needs to perform an image subtitle training task on model A, first select the interface corresponding to the training task performed by model A from multiple model call interfaces on the remote server, that is, the reference model call interface (such as interface A). Then send the target training data set (data set A) and target training instruction (instruction A) required to perform the training task to interface A, where data set A and instruction A are included in the target call request and sent to interface A. Subsequently, after receiving the target call request, interface A can perform the image subtitle training task on model A according to the data set A and instruction A carried in the target call request according to calling model A, and return the training result to the local terminal after model A completes the image subtitle training task.

[0071] Optionally, in this embodiment, the target training result can be the training result returned by interface A received by the local terminal, which can be, but is not limited to, the performance metrics obtained by Model A performing the image caption training task, including accuracy, loss value, precision, etc., or the trained model weights and the intermediate layer outputs of the model. The user can view the target training result on the local terminal and further optimize and debug the model based on the target training result.

[0072] As an optional solution, sending a target call request carrying the target training dataset and the target training instruction to the reference model call interface includes:

[0073] S21, obtaining the target interface address corresponding to the reference model call interface;

[0074] S22, adding the target training dataset and the target training instruction to the initial call request to obtain the target call request;

[0075] S23, sending the target call request to the reference model call interface according to the target interface address.

[0076] Optionally, in this embodiment, the manner in which the local terminal sends the target call request to interface A can be, but is not limited to, including: First, the local terminal obtains the address of interface A and adds dataset A and instruction A to the initial call request to obtain the target call request.

[0077] As an optional solution, when the reference model is a vision-language model, before adding the target training dataset and the target training instruction to the initial call request to obtain the target call request, the method further includes:

[0078] S31, obtaining an image caption dataset, a visual question answering dataset, and a visual foundation dataset, and determining at least one of the image caption dataset, the visual question answering dataset, and the visual foundation dataset as the target training dataset, where the image caption dataset includes multiple images and the descriptive text for each image, and the image caption dataset is used to train the performance of the vision-language model in the image description or caption generation task, the visual question answering dataset includes multiple images and the questions and corresponding answers proposed for the content of each image, and the visual question answering dataset is used to train the performance of the vision-language model in the visual question answering task, the visual foundation dataset includes multiple images and the descriptive text for the local area of each image, and the visual foundation dataset is used to train the performance of the vision-language model to recognize the content of the local area of the image;

[0079] S32, obtaining the reference model parameters corresponding to the reference model; generating a parameter adjustment instruction corresponding to the reference model parameters, where the parameter adjustment instruction is used to indicate adjusting the model parameters of the reference model to the reference model parameters; and determining the parameter adjustment instruction as the target training instruction.

[0080] Optionally, in this embodiment, when the model (reference model) that the user wants to execute the training task is a vision-language model, the target training dataset used by the user may include the following three types of datasets, and one or more datasets can be selected as the target training dataset:

[0081] Dataset 1, an image caption dataset, containing multiple images and descriptive text for each image. Its purpose is to train the performance of the vision-language model in image description or caption generation tasks. For example, given an image, the model needs to generate a piece of text to describe the image content or create a suitable title for the image.

[0082] Dataset 2, a visual question answering dataset, containing multiple images, questions raised for the content of each image, and corresponding answers. Its purpose is to train the performance of the vision-language model in visual question answering tasks. For example, given an image and a question (such as "What is the person in the picture doing?"), the model needs to provide the correct answer.

[0083] Dataset 3, a visual grounding dataset, containing multiple images and descriptive text for local regions of each image. Its purpose is to train the performance of the vision-language model to recognize the content of local regions of images. For example, given an image and a query pointing to a specific region in the image (such as "Where is the cat in the picture?"), the model needs to locate the corresponding region in the image and provide a description.

[0084] Optionally, in this embodiment, the reference model parameters can be, but are not limited to, model parameters that can control the model training effect during the model training process, including: hyperparameters (such as learning rate, batch size, optimizer type, and regularization coefficient, etc.) and model weight parameters, etc. The model parameters need to set initial values before the start of model training, and for the same model parameter in the model, different initial values may be set when training different datasets. The purpose is to enable the model to better adapt to different datasets and task requirements, thereby optimizing the performance of the model.

[0085] Optionally, in this embodiment, when the model (reference model) that the user wants to execute the training task is a vision-language model, the target training instruction used by the user can be obtained in the following way: Assume the reference model is Model A. First, obtain the model parameters of Model A, and at the same time parse the settings related to the initial values of the model parameters included in the model training request initiated by the user. Then generate a parameter adjustment instruction according to the model parameters and the initial values of the model parameters. Among them, the parameter adjustment instruction is an instruction that can adjust the model parameters of Model A to the initial values of the model parameters set by the user. Finally, determine the parameter adjustment instruction as the target training instruction (i.e., Instruction A).

[0086] Optionally, in this embodiment, the hyperparameter settings for the pre-training of the vision-language model can be as shown in Table 1:

[0087] Table 1

[0088]

[0089] The hyperparameter settings for the multi-task training of the vision-language model can be as shown in Table 2:

[0090] Table 2

[0091] Hyperparameter Multitask Learning rate <![CDATA[1e -5 > Total step length 6000 Batch size 1024 Adaptive Moment Estimation parameter ε <![CDATA[1e -8 > Adaptive Moment Estimation parameter β (0.9,0.95) Weight decay 0.1 Dropout rate 0.1 Input resolution <![CDATA[490 2 >

[0092] As an alternative solution, the reference model call interface calls the reference model to execute the model training task through the following steps:

[0093] S41, adjust the model parameters of the reference model to the reference model parameters according to the parameter adjustment instruction to obtain the reference model to be trained;

[0094] S42, perform multiple rounds of training on the reference model to be trained through the image caption dataset, visual question answering dataset, and visual grounding dataset until the model parameters of the reference model to be trained converge, and obtain the trained reference model;

[0095] S43, return the training result indicating the successful training of the model training task to the local terminal.

[0096] Optionally, in this embodiment, the manner in which interface A calls model A to execute the image caption training task can include but is not limited to: First, after interface A receives the target call request sent by the local terminal, it sets the model parameters of model A to the initial values indicated by the parameter adjustment instruction in instruction A (the target training instruction carried by the target call request), inputs dataset A (the target training dataset carried by the target call request) into model A for training, and the model parameters of model A are continuously iteratively updated during the training process. After all model parameters converge, that is, model A completes the training task, the model training result is returned to the local terminal.

[0097] As an alternative solution, before screening out the reference model call interface corresponding to the reference model from multiple model call interfaces on the remote server, the method further includes:

[0098] S51, the remote server obtains the terminal sending parameters and terminal receiving parameters of the local terminal, as well as the model call parameters of the reference model, where the terminal sending parameters are used to indicate the data sending method of the local terminal, the terminal receiving parameters are used to indicate the data receiving method allowed by the local terminal, and the model call parameters are the parameters required to call the reference model;

[0099] S52. The remote server generates a reference model call interface based on the terminal sending parameters, terminal receiving parameters, and model call parameters. The reference model call interface is used to receive data that meets the data sending method indicated by the terminal sending parameters, call the reference model according to the model call parameters, and send data to the local terminal according to the data receiving method indicated by the terminal receiving parameters.

[0100] Optionally, in this embodiment, the manner in which the remote server generates the reference model call interface may include, but is not limited to: determining the data exchange format between the local terminal and the remote server, including the terminal sending parameters, terminal receiving parameters, and model call parameters. The terminal sending parameters refer to the data format used by the local terminal when sending data to the remote server; the terminal receiving parameters refer to the format in which the local terminal can receive the data returned by the remote server; the model call parameters refer to the specific parameters required to call the reference model. Then, a reference model call interface is generated based on the above parameters. This interface can receive data that conforms to the terminal sending parameters to process the requests and responses of the local terminal, and after processing the model call, send data (such as the target training result) to the local terminal according to the data format specified by the terminal receiving parameters.

[0101] Optionally, in this embodiment, Figure 3 is a schematic diagram of the data exchange format in the communication between the local terminal and the remote server according to an embodiment of the present application. As Figure 3 shown, in the request header of the data format allowed to be sent from the local terminal to the remote server, Content-Type: application / json, and the terminal receiving parameter is Accept: application / json in the response header allowed to be received by the local terminal from the remote server.

[0102] Optionally, in this embodiment, the remote server sets one or more interfaces for each AI model deployed therein. Each interface can independently call the corresponding AI model, and different interfaces corresponding to the same AI model can call the model to perform different tasks. For example, model B corresponds to two interfaces, namely interface B1 and interface B1. Interface B1 can call model B to perform a training task, and interface B2 can call model B to perform an inference task. This can improve the efficiency of performing model tasks on the AI model.

[0103] In this embodiment, a method for executing a remote server model task is also provided, which is applied to a local terminal. Multiple models are deployed on the remote server. The remote server has pre-generated at least one model call interface for each of the models. The multiple models include the trained target model. Figure 4It is a flowchart of another method for executing a remote server model task according to an embodiment of the present application. As Figure 4 shown, the process includes the following steps:

[0104] As an optional solution, the method further includes:

[0105] S61, receiving a model inference request initiated by an inference account, where the model inference request is used to request the target model to perform a model inference task;

[0106] S62, in response to the model inference request, screening out the target model call interface corresponding to the target model from multiple model call interfaces on the remote server;

[0107] S63, sending a target inference request carrying model input data and a target inference instruction to the target model call interface, where the target inference request is used to request the reference model call interface to call the target model to perform a model inference task according to the model input data and the target inference instruction, and return the inference result of the model inference task to the local terminal;

[0108] S64, receiving the target inference result returned by the model call interface in response to the target inference request as the execution result of the model inference task.

[0109] Optionally, in this embodiment, the training account may but is not limited to an account pre-registered on the local terminal for training the model, and the inference account may but is not limited to an account pre-registered on the local terminal for calling the model to perform an inference task. The training account and the inference account may but is not limited to be the same account.

[0110] Optionally, in this embodiment, the user can initiate a model inference request on the local terminal through the inference account. This request specifies the target model the user wants to use, such as a vision-language model (Model A), and the type of task the user hopes to perform, such as an image captioning inference task.

[0111] Optionally, in this embodiment, when the local terminal needs to perform an image captioning inference task on Model A, first select the interface corresponding to Model A for performing the inference task, that is, the target model call interface (such as Interface A1) from multiple model call interfaces on the remote server. Then send the model input data (Data X) and the target inference instruction (Instruction X) required to execute this inference task to Interface A1, where Data X and Instruction X are included in the target inference request and sent to Interface A1. Subsequently, after receiving the target inference request, Interface A1 can execute the image captioning inference task on Model A according to Data X and Instruction X carried in the target inference request, and return the inference result to the local terminal after Model A completes the image captioning inference task.

[0112] Optionally, in this embodiment, the target inference result may be the inference result returned by interface A1 received by the local terminal, and may be, but is not limited to, the model output obtained by model A performing the image captioning inference task.

[0113] As an optional solution, the target model invocation interface invokes the target model to perform the model inference task through the following steps:

[0114] S71, input the model input data into the target model;

[0115] S72, control the target model to perform the inference operation indicated by the target inference instruction according to the input model input data;

[0116] S73, return the inference result output by the target model performing the inference operation to the local terminal.

[0117] Optionally, in this embodiment, the manner in which the target model invocation interface (interface A1) invokes the target model (model A) to perform the model inference task may include, but is not limited to: First, after interface A1 receives the target inference request sent by the local terminal, it inputs the data X that needs to perform the inference task into model A; then, control model A to perform the inference operation (such as image captioning inference) indicated by the target inference instruction on data X to obtain the model output result, that is, the inference result; finally, return the inference result to the local terminal.

[0118] Optionally, in this embodiment, assume that the user initiates a model inference request on the local terminal to request model A deployed on the remote server to perform the model inference task. Specifically, the model input data is data X (a photo containing a beach and palm trees) saved on the local terminal, and the inference operation indicated by the target inference instruction is image captioning inference (generating an image description for data X). After the local terminal initiates a target inference request containing data X and the target inference instruction, interface A1 invokes model A to perform the model inference task according to data X and the target inference instruction, obtains the model output result (a sunny beach and palm trees), and returns the model output result to the local terminal for the user to view.

[0119] Optionally, in this embodiment, to better understand the process of executing the above remote server model task, the following further describes the execution process of the above remote server model task in combination with optional embodiments, but is not used to limit the technical solutions of the embodiments of the present application.

[0120] In this embodiment, a method for executing a remote server model task is provided. Figure 5 It is a logical architecture diagram of the input of a vision-language model according to an embodiment of the present application, as Figure 5As shown, the vision - language model (Model A) can include the following components: a vision transformer encoder, a multi - layer perceptron adapter, a pre - trained large - language model, a vision expert module, and location embeddings. The vision transformer encoder can process image data and extract [CLS] features for understanding image content. The multi - layer perceptron adapter can transform image features into the same space as text features and simplify feature fusion by sharing location IDs. The pre - trained large - language model can utilize masked attention. The vision expert module is added to each layer of the model and works in conjunction with the QKV matrix and the multi - layer perceptron model to enhance the fusion of vision and language features. Location embeddings allow image features to share location IDs, reduce the attenuation of long - range dependencies, and improve the efficiency of image data processing.

[0121] The remote server sets a model call interface for Model A. Suppose the interface corresponding to Model A is Interface A. Figure 6 It is a schematic diagram of the Restful interface of a vision - language model according to an embodiment of the present application. As Figure 6 shown, Model A can be called through Interface A to execute model tasks. Even if the dataset for the model task to be executed is stored on the local terminal, the local terminal can still call Model A through data interaction with Interface A on the remote server, thereby completing the model task and obtaining the result output by Model A when executing the model task.

[0122] Through the description of the above - mentioned implementation manners, those skilled in the art can clearly understand that the method according to the above - mentioned embodiments can be implemented by means of software plus a necessary general - purpose hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of various embodiments of the present application.

[0123] In this embodiment, an execution device for remote - server model tasks is also provided. This device is used to implement the above - mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0124] Figure 7 It is a structural block diagram of an execution device for remote - server model tasks according to an embodiment of the present application. As Figure 7As shown, it is applied to a local terminal. The target training dataset of the reference model is stored in the local terminal, and multiple models are deployed on a remote server. The remote server has pre-generated at least one model call interface for each model. The multiple models include the reference model. The execution device of the remote server model task includes:

[0125] A first receiving module 702, configured to receive a model training request initiated by a training account, where the model training request is used to request to perform a model training task on the reference model;

[0126] A first response module 704, configured to respond to the model training request and screen out the reference model call interface corresponding to the reference model from multiple model call interfaces on the remote server;

[0127] A first sending module 706, configured to send a target call request carrying the target training dataset and the target training instruction to the reference model call interface, where the target call request is used to request the reference model call interface to call the reference model to perform the model training task according to the target training dataset and the target training instruction, and return the training result of the model training task to the local terminal;

[0128] A second receiving module 708, configured to receive the target training result returned by the model call interface in response to the target call request as the execution result of the model training task.

[0129] In an exemplary embodiment, the first sending module includes:

[0130] An obtaining unit, configured to obtain the target interface address corresponding to the reference model call interface;

[0131] An adding unit, configured to add the target training dataset and the target training instruction to the initial call request to obtain the target call request;

[0132] A sending unit, configured to send the target call request to the reference model call interface according to the target interface address.

[0133] In an exemplary embodiment, when the reference model is a vision-language model, the execution device of the remote server model task further includes:

[0134] A first acquisition module, configured to acquire an image caption dataset, a visual question answering dataset, and a visual foundation dataset before adding a target training dataset and a target training instruction to an initial call request to obtain a target call request, and determine at least one of the image caption dataset, the visual question answering dataset, and the visual foundation dataset as the target training dataset. The image caption dataset includes a plurality of images and descriptive texts for each image, and is used to train the performance of a vision-language model in image description or caption generation tasks. The visual question answering dataset includes a plurality of images, questions and corresponding answers raised for the content of each image, and is used to train the performance of a vision-language model in visual question answering tasks. The visual foundation dataset includes a plurality of images and descriptive texts for local regions of each image, and is used to train the performance of a vision-language model in recognizing the content of local regions of images;

[0135] A second acquisition module, configured to acquire reference model parameters corresponding to a reference model; generate a parameter adjustment instruction corresponding to the reference model parameters, where the parameter adjustment instruction is used to indicate adjusting the model parameters of the reference model to the reference model parameters; and determine the parameter adjustment instruction as the target training instruction.

[0136] In an exemplary embodiment, the reference model call interface calls the reference model to execute a model training task through the following steps:

[0137] Adjust the model parameters of the reference model to the reference model parameters according to the parameter adjustment instruction to obtain a reference model to be trained;

[0138] Perform multiple rounds of training on the reference model to be trained through the image caption dataset, the visual question answering dataset, and the visual foundation dataset until the model parameters of the reference model to be trained converge, to obtain a trained reference model;

[0139] Return a training result indicating the successful training of the model training task to the local terminal.

[0140] In an exemplary embodiment, the execution device for the remote server model task further includes:

[0141] A third acquisition module, configured to acquire terminal sending parameters and terminal receiving parameters of the local terminal, and model call parameters of the reference model before screening out a reference model call interface corresponding to the reference model from a plurality of model call interfaces on the remote server, where the terminal sending parameters are used to indicate the data sending method of the local terminal, the terminal receiving parameters are used to indicate the data receiving method allowed by the local terminal, and the model call parameters are the parameters required to call the reference model;

[0142] A generation and acquisition module, configured to generate a reference model call interface by a remote server based on terminal sending parameters, terminal receiving parameters, and model call parameters, where the reference model call interface is used to receive data that meets the data sending method indicated by the terminal sending parameters, call a reference model according to the model call parameters, and send the data to a local terminal according to the data receiving method indicated by the terminal receiving parameters.

[0143] Figure 8 is a structural block diagram of another execution device for a remote server model task according to an embodiment of the present application; as Figure 8 shown, applied to a local terminal, multiple models are deployed on the remote server, and the remote server has pre-generated at least one model call interface for each of the models, and the multiple models include a trained target model. Another execution device for a remote server model task includes:

[0144] A third receiving module 802, configured to receive a model inference request initiated by an inference account, where the model inference request is used to request the target model to execute a model inference task;

[0145] A second response module 804, configured to respond to the model inference request and screen out the target model call interface corresponding to the target model from multiple model call interfaces on the remote server;

[0146] A second sending module 806, configured to send a target inference request carrying model input data and a target inference instruction to the target model call interface, where the target inference request is used to request the reference model call interface to call the target model to execute a model inference task according to the model input data and the target inference instruction, and return the inference result of the model inference task to the local terminal;

[0147] A fourth receiving module 808, configured to receive the target inference result returned by the model call interface in response to the target inference request as the execution result of the model inference task.

[0148] In an exemplary embodiment, the target model call interface calls the target model to execute a model inference task through the following steps:

[0149] Input the model input data into the target model;

[0150] Control the target model to execute the inference operation indicated by the target inference instruction according to the input model input data;

[0151] Return the inference result output by the target model when executing the inference operation to the local terminal.

[0152] It should be noted that the above-mentioned modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: the above-mentioned modules are all located in the same processor; or, the above-mentioned modules are respectively located in different processors in any combination form.

[0153] The embodiments of the present application also provide a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the methods in the various embodiments of the present application; the computer program product also includes a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the steps of the methods in the various embodiments of the present application.

[0154] The embodiments of the present application also provide a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. Wherein, the computer program is set to execute the steps in any one of the above method embodiments when running.

[0155] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disc and other various media that can store computer programs.

[0156] The embodiments of the present application also provide an electronic device. Figure 9 It is a schematic diagram of an electronic device according to an embodiment of the present application, as Figure 9 shown. The electronic device includes a memory and a processor. A computer program is stored in the memory. The processor is set to run the computer program to execute the steps in any one of the above method embodiments.

[0157] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device. Wherein, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0158] For the specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary embodiments. This embodiment will not be elaborated here.

[0159] Obviously, those skilled in the art should understand that the various modules or steps of the present application described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present application is not limited to any specific combination of hardware and software.

[0160] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for executing a remote server model task, characterized in that: Applied to a local terminal, the local terminal stores a target training data set of a reference model, a plurality of models are deployed on a remote server, the remote server pre-generates at least one model calling interface for each of the models, the plurality of models include the reference model, and the method comprises: Receiving a model training request initiated by a training account, wherein the model training request is used to request execution of a model training task on the reference model; In response to the model training request, a reference model calling interface corresponding to the reference model is screened out from a plurality of model calling interfaces on the remote server; Sending a target call request carrying the target training data set and the target training instruction to the reference model call interface, wherein the target call request is used to request the reference model call interface to call the reference model to perform the model training task according to the target training data set and the target training instruction, and return the training result of the model training task to the local terminal; Receive the target training result returned by the model calling interface in response to the target calling request as the execution result of the model training task.

2. The method according to claim 1, characterized in that The sending a target call request carrying the target training data set and the target training instruction to the reference model call interface includes: Obtaining a target interface address corresponding to the reference model calling interface; Adding the target training data set and the target training instruction to an initial call request to obtain the target call request; The target call request is sent to the reference model call interface according to the target interface address.

3. The method according to claim 2, characterized in that In the case where the reference model is a visual language model, before adding the target training data set and the target training instruction to the initial call request to obtain the target call request, the method further includes: Acquire an image caption dataset, a visual question answering dataset, and a visual base dataset, and determine at least one of the image caption dataset, the visual question answering dataset, and the visual base dataset as the target training dataset, wherein the image caption dataset includes a plurality of images and a description text for each image, and the image caption dataset is used to train the performance of the visual language model on image description or caption generation tasks, the visual question answering dataset includes a plurality of images, and questions and corresponding answers raised for the content of each image, and the visual question answering dataset is used to train the performance of the visual language model on the visual question answering task, and the visual base dataset includes a plurality of images and a description text for a local area of ​​each image, and the visual base dataset is used to train the performance of the visual language model in identifying the content of a local area of ​​an image; Acquire reference model parameters corresponding to the reference model; generate parameter adjustment instructions corresponding to the reference model parameters, wherein the parameter adjustment instructions are used to instruct to adjust the model parameters of the reference model to the reference model parameters; determine the parameter adjustment instructions as the target training instructions.

4. The method according to claim 3, characterized in that The reference model calling interface calls the reference model to perform the model training task through the following steps: Adjusting the model parameters of the reference model to the reference model parameters according to the parameter adjustment instruction to obtain the reference model to be trained; Performing multiple rounds of training on the reference model to be trained by using the image caption dataset, the visual question answering dataset, and the visual basic dataset until the model parameters of the reference model to be trained converge, thereby obtaining the trained reference model; A training result indicating that the model training task is successfully trained is returned to the local terminal.

5. The method according to claim 1, characterized in that Before filtering out the reference model calling interface corresponding to the reference model from the multiple model calling interfaces on the remote server, the method further includes: The remote server acquires a terminal sending parameter and a terminal receiving parameter of the local terminal, and a model calling parameter of the reference model, wherein the terminal sending parameter is used to indicate a data sending mode of the local terminal, the terminal receiving parameter is used to indicate a data receiving mode allowed by the local terminal, and the model calling parameter is a parameter required to call the reference model; The remote server generates the reference model calling interface according to the terminal sending parameters, the terminal receiving parameters and the model calling parameters, wherein the reference model calling interface is used to receive data that satisfies the data sending method indicated by the terminal sending parameters, call the reference model according to the model calling parameters, and send data to the local terminal according to the data receiving method indicated by the terminal receiving parameters.

6. A method for executing a remote server model task, characterized in that: Applied to a local terminal, multiple models are deployed on a remote server, the remote server pre-generates at least one model calling interface for each model, the multiple models include a trained target model, and the method includes: Receiving a model inference request initiated by an inference account, wherein the model inference request is used to request the target model to perform a model inference task; In response to the model inference request, a target model calling interface corresponding to the target model is screened out from a plurality of model calling interfaces on the remote server; Sending a target reasoning request carrying model input data and target reasoning instructions to the target model calling interface, wherein the target reasoning request is used to request the target model calling interface to call the target model to perform the model reasoning task according to the model input data and the target reasoning instruction, and return the reasoning result of the model reasoning task to the local terminal; The target reasoning result returned by the model calling interface in response to the target reasoning request is received as the execution result of the model reasoning task.

7. The method according to claim 6, characterized in that The target model calling interface calls the target model to perform the model reasoning task through the following steps: inputting the model input data into the target model; Controlling the target model to execute the reasoning operation indicated by the target reasoning instruction according to the input model input data; The inference result output by the target model executing the inference operation is returned to the local terminal.

8. A remote server model task execution device, characterized in that: Applied to a local terminal, the local terminal stores a target training data set of a reference model, a plurality of models are deployed on a remote server, the remote server pre-generates at least one model calling interface for each of the models, the plurality of models include the reference model, and the execution device of the remote server model task includes: A first receiving module is used to receive a model training request initiated by a training account, wherein the model training request is used to request to perform a model training task on the reference model; A first response module, used to respond to the model training request and filter out a reference model calling interface corresponding to the reference model from a plurality of model calling interfaces on the remote server; A first sending module is used to send a target calling request carrying the target training data set and the target training instruction to the reference model calling interface, wherein the target calling request is used to request the reference model calling interface to call the reference model to perform the model training task according to the target training data set and the target training instruction, and return the training result of the model training task to the local terminal; The second receiving module is used to receive the target training result returned by the model calling interface in response to the target calling request as the execution result of the model training task.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in any one of claims 1 to 7 when executed by a processor.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • A method and device for sending information

    CN110288089A

  • Text pair classification model training method, classification method and equipment and storage medium

    CN112818658A

  • Method and device for realizing remote training

    CN114265690A

  • Model training method and device, computer readable storage medium and computer equipment

    CN115018043A

  • Model training method and device, computer equipment and computer readable storage medium

    CN115457572A