Task execution method and system based on large model, storage medium and electronic equipment

By receiving task execution requests and obtaining cloud server equipment to execute tasks as target computing nodes, the problems of low efficiency and poor flexibility of large-model tasks are solved, and efficient and flexible task execution is achieved.

CN120508381APending Publication Date: 2025-08-19PURPLE MOUNTAIN LAB
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510479703.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The task execution method based on large models in the prior art has the problem of low task execution efficiency and poor flexibility. Especially in a large number of business concurrency scenarios, the computing power resource scheduling is inflexible, resulting in long queue time and long response time, which affects the user experience.

Method used

By receiving task execution requests, obtaining target computing nodes and sending tasks to cloud server devices for execution, intelligent scheduling and automatic identification of remote computing nodes are realized, improving task execution efficiency and flexibility.

Benefits of technology

It realizes efficient execution and flexible scheduling based on large-model tasks, solves the problems of low task execution efficiency and poor flexibility, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508381A_ABST
    Figure CN120508381A_ABST
Patent Text Reader

Abstract

The invention discloses a task execution method and system based on a large model, a storage medium and electronic equipment. The method relates to the technical field of artificial intelligence and comprises the steps that a task execution request for a target task is received, and the target task is an inference task based on a large model or a fine tuning task based on the large model; a target computing power node for executing the target task is obtained according to the target task, and the target computing power node is server equipment deployed at the cloud; and sending the task execution request to the target computing power node, so that the target computing power node executes the target task based on the task execution request. The technical problems of low task execution efficiency and poor flexibility of a task execution method based on a large model in related technologies are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a large model-based task execution method, system, storage medium and electronic device. Background Art

[0002] Currently, big model technology is gradually penetrating various industries. The processing of hundreds of billions of data samples and model parameter training by big models requires the support of advanced computing power. Computing power has become the primary prerequisite for the development and application of big model technology. However, the computing power resources and the number of concurrent services supported by the task execution methods of big models in related technologies are limited. In scenarios with a large number of concurrent services, computing power resource scheduling flexibility is poor, and queuing and response times are long, resulting in low task execution efficiency. In some cases, it may also cause unstable operation, affecting the user experience.

[0003] Currently, no effective solution has been proposed to the problems of low task execution efficiency, poor flexibility and stability in the task execution methods based on large models in the above-mentioned related technologies. Summary of the Invention

[0004] The embodiments of the present invention provide a large-model-based task execution method, system, storage medium, and electronic device to at least solve the technical problems of low task execution efficiency and poor flexibility in the large-model-based task execution method in related technologies.

[0005] According to one aspect of an embodiment of the present invention, a task execution method based on a large model is provided, including: receiving a task execution request for a target task, wherein the target task is an inference task based on a large model, or a fine-tuning task based on the large model; obtaining a target computing power node for executing the target task according to the target task, wherein the target computing power node is a server device deployed in the cloud; sending the task execution request to the target computing power node, so that the target computing power node executes the target task based on the task execution request.

[0006] According to another aspect of an embodiment of the present invention, another task execution method based on a large model is provided, including: receiving a task execution request for a target task, wherein the target task is an inference task based on a large model, or a fine-tuning task based on the large model, the task execution request is sent by a data processing device when the preset execution conditions of the target task are not met, and the local computing power node is a server device deployed on the local side; executing the target task based on the task execution request.

[0007] According to another aspect of an embodiment of the present invention, another task execution method based on a large model is provided, including: receiving a service request for a target task, wherein the target task is an inference task based on a large model, and the service request is an inference service request for the inference task; or the target task is a fine-tuning task based on the large model, and the service request is a fine-tuning deployment request for the fine-tuning task; generating a task execution request based on the service request; sending the task execution request to a data processing device, for the data processing device to send the task execution request to a target computing power node if the preset execution conditions of the target task are not met, so that the target computing power node executes the target task.

[0008] According to another aspect of an embodiment of the present invention, a task execution system based on a large model is also provided, comprising: a target interaction device, a data processing device, and a target computing power node, wherein the target computing power node is a server device deployed in the cloud, wherein: the target interaction device is used to generate a task execution request based on the service request when receiving a service request for a target task, and send the task execution request to the data processing device, wherein the target task is an inference task based on a large model, and the service request is an inference service request for the inference task; or the target task is a fine-tuning task based on the large model, and the service request is a fine-tuning deployment request for the fine-tuning task; the data processing device is used to send the task execution request to the target computing power node; and the target computing power node is used to execute the target task based on the task execution request.

[0009] According to another aspect of an embodiment of the present invention, a non-volatile storage medium is provided, wherein the non-volatile storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing any one of the large model-based task execution methods.

[0010] According to another aspect of an embodiment of the present invention, an electronic device is also provided, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement any one of the large model-based task execution methods.

[0011] According to another aspect of an embodiment of the present invention, a computer program product is provided, including a computer program, which implements the steps of any one of the large model-based task execution methods when executed by a processor.

[0012] In an embodiment of the present invention, a task execution request for a target task is received, wherein the target task is an inference task based on a large model, or a fine-tuning task based on the large model; a target computing power node for executing the target task is obtained according to the target task, wherein the target computing power node is a server device deployed in the cloud; the task execution request is sent to the target computing power node, so that the target computing power node executes the target task based on the task execution request, thereby achieving the purpose of intelligent scheduling to automatically identify and select remote (target) computing power nodes to execute tasks, thereby achieving the technical effect of improving the execution efficiency and execution flexibility of tasks based on large models, and further solving the technical problems of low task execution efficiency and poor flexibility in the task execution method based on large models in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0014] Figure 1 is a flowchart of a task execution method based on a large model according to an embodiment of the present invention;

[0015] Figure 2 is a flowchart of another large model-based task execution method according to an embodiment of the present invention;

[0016] Figure 3 is a flowchart of another large model-based task execution method according to an embodiment of the present invention;

[0017] Figure 4 is a flowchart of an optional reasoning deployment phase of a reasoning task according to an embodiment of the present invention;

[0018] Figure 5 is a flowchart of an optional reasoning service phase of a reasoning task according to an embodiment of the present invention;

[0019] Figure 6 is a flowchart of an optional fine-tuning task execution according to an embodiment of the present invention;

[0020] Figure 7 2 is a schematic diagram of the structure of a task execution system based on a large model according to an embodiment of the present invention;

[0021] Figure 8 is a schematic diagram of a first system structure for performing a large model reasoning task according to an embodiment of the present invention;

[0022] Figure 9is a schematic diagram of a first system structure for performing a fine-tuning task of a large model according to an embodiment of the present invention;

[0023] Figure 10 is a schematic diagram of a second system structure for performing reasoning tasks of a large model according to an embodiment of the present invention;

[0024] Figure 11 is a schematic diagram of a second system structure for performing a fine-tuning task of a large model according to an embodiment of the present invention;

[0025] Figure 12 is a schematic diagram of a third system structure for performing reasoning tasks of a large model according to an embodiment of the present invention;

[0026] Figure 13 is a schematic diagram of a third system structure for performing a fine-tuning task of a large model according to an embodiment of the present invention;

[0027] Figure 14 is a schematic diagram of a fourth system structure for performing reasoning tasks of a large model according to an embodiment of the present invention;

[0028] Figure 15 is a schematic diagram of a fourth system structure for performing a fine-tuning task of a large model according to an embodiment of the present invention;

[0029] Figure 16 2 is a schematic diagram of a task execution device based on a large model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0030] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0031] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0032] First, to facilitate understanding of the embodiments of the present invention, some of the terms or nouns involved in the present invention are explained below:

[0033] The Large Model Training and Inference Machine is a high-performance computing device designed specifically for the training and inference of large-scale pre-trained models (such as natural language processing and computer vision models). It integrates the necessary hardware resources and software environment to provide a one-stop solution to support efficient training and inference services for large models.

[0034] A large model's model file refers to a file that stores the model's structure, weights, and parameters. In the fields of deep learning and machine learning, a model file is the concrete representation of a trained neural network model. It contains the structural definition of each layer, its connection structure, activation function, loss function, and most importantly, parameters such as the model weights and biases. These parameters are derived through learning from large amounts of data during the training process and determine how the model predicts output results from input data.

[0035] Image files refer to container images used to run large model tasks (such as fine-tuning and inference tasks) on specific hardware environments. Container images, particularly Docker images or similar container technology images, contain all the software environments, dependent libraries, fine-tuning code, and pre-configured parameters required to run large model tasks. Fine-tuning image files refer to container images used to run large model fine-tuning tasks on specific hardware environments.

[0036] At present, big model technology is gradually penetrating into all walks of life. Big models process hundreds of billions of data samples and train model parameters, which cannot be separated from the support of advanced computing power base. Computing power has become the primary prerequisite for the development and application of big model technology.

[0037] With the introduction of high-performance, low-cost open-source big models, AI big models have garnered unprecedented attention. Major chip and computing power providers have actively developed big model services. Some chip manufacturers have already completed the adaptation of large models of varying specifications, and several mainstream AI computing power vendors have launched big model services. Furthermore, big model training and inference appliances integrate advanced AI technologies with high-performance computing capabilities, providing a secure and stable hardware computing foundation for training and inference of large models. Users who purchase a training and inference appliance receive access to the built-in big model service. Due to its localized, specialized, and out-of-the-box capabilities, the appliance has become a critical infrastructure for building and deploying complex AI applications, gaining widespread adoption in industries such as finance, scientific research, and healthcare. To achieve a superior big model experience, some companies are opting for dedicated cloud computing or building their own computing clusters to ensure computing power availability during emergencies and other critical situations.

[0038] However, the task execution method based on large models in related technologies has the following problems:

[0039] 1) The computing power resources of the all-in-one machine and the number of concurrent services it can support are limited. In scenarios with a large number of concurrent services, the waiting time and response time are long, affecting the user experience.

[0040] 2) The all-in-one machine supports a small number of models locally and updates are not timely. Limited computing resources restrict the business development of large-parameter specification models. The local resources and performance of the all-in-one machine are difficult to meet the needs of the ever-increasing application scenarios.

[0041] 3) On the demand side for large-scale computing power, varying user roles and usage scenarios generate differentiated computing power requirements. Meanwhile, on the supply side, suppliers have varying capabilities and complex pricing rules. The lack of efficient distribution channels between supply and demand makes it difficult for computing power suppliers to identify potential users and provide targeted services, and for users to select cost-effective computing power resources that meet their business needs.

[0042] 4) The price of an all-in-one machine ranges from hundreds of thousands to millions. When enterprises choose to build their own computing clusters to ensure the supply of computing power during special periods, the cost is very high. After the emergency is over, the computing resources become idle and wasted.

[0043] 5) If the user has accessed specific third-party computing power in advance, the limited third-party computing power resources are not exclusive to the user. In the event of sudden high concurrency, the timeliness and stability of business response cannot be guaranteed.

[0044] 6) Whether it is a self-built and self-sold cloud-edge collaborative model or a model that accesses specific third-party computing resources, it is necessary to transmit non-privacy data sets and large model files between the cloud and the edge. These data sets and model files are huge in size, and the security, reliability and speed of network transmission become business bottlenecks.

[0045] In response to the above problems, an embodiment of the present invention provides an embodiment of a method for executing tasks based on a large model. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0046] Figure 1 is a flowchart of a task execution method based on a large model according to an embodiment of the present invention. Figure 1 As shown, the method includes the following steps:

[0047] Step S102: receiving a task execution request for a target task, wherein the target task is an inference task based on a large model, or a fine-tuning task based on a large model;

[0048] Optionally, the execution subject of steps S102 to S106 may be a data processing device, on which a scheduling system and / or a load balancing system is deployed. For example, when the target task is a fine-tuning task based on a large model, a scheduling system may be deployed on the data processing device to execute the method steps of steps S102 to S106. When the target task is an inference task based on a large model, a load balancing system may be deployed on the data processing device to execute the method steps of steps S102 to S106, and the load balancing system may also interact with the scheduling system to complete the execution of the method steps, wherein the scheduling system may be deployed on the data processing device, or on other processor devices or server devices other than the data combing device.

[0049] Optionally, the big model can be, but is not limited to, a big language model focusing on natural language processing, a multimodal big model that simultaneously processes and fuses multiple data modalities (such as text, images, videos, audio, etc.), a visual big model focusing on image and video data, a generative big model focusing on generating new data (covering multiple forms such as text, images, audio, etc.), etc.

[0050] Optional, large-model-based reasoning tasks involve leveraging a large model (either pre-trained or fine-tuned) to achieve specific tasks through direct invocation or with minimal input, without modifying model parameters. Large-model-based fine-tuning tasks involve using a pre-trained large model to adjust model parameters through supervised learning to improve performance for specific tasks or domains.

[0051] In an optional embodiment, when the target task is an inference task, the task execution request is an inference task execution request, wherein the inference task execution request is used to request the execution of the inference task to obtain an inference service result, and the inference task execution request carries a task execution request with inference context information and a knowledge retrieval result, wherein the inference context information is based on the inference service request for the target task initiated by the terminal device and is obtained by querying the state library; the knowledge retrieval result is based on the inference service request and is obtained by querying the knowledge library; the inference service request carries request information input through the terminal device, wherein the request information can be but is not limited to text, images or instructions input through the terminal device.

[0052] Optionally, when the target task is an inference task, the corresponding received task execution request is an inference task execution request. The inference context information and knowledge retrieval results carried in the inference task execution request are based on the inference service request initiated by the terminal device and are obtained by querying the state library and knowledge library. In this way, relevant information can be dynamically obtained based on different inference service requests, thereby improving the flexibility and adaptability of the inference task. Different inference service requests may involve different contexts and knowledge requirements, and this method can better meet these requirements.

[0053] The optional state repository is primarily used to store and manage state information during the inference service process. This includes, but is not limited to, inference context, providing real-time feedback to the scheduling system and administrators to assist with resource management and troubleshooting, and result caching. The knowledge base contains external information to assist in inference, such as domain knowledge, specialized terminology, fact databases, and corpora. During inference, large models can access the knowledge base to obtain additional information, providing more detailed and professional responses when answering user requests or executing tasks. During the knowledge retrieval phase, relevant knowledge items or information fragments can be found based on the user's original input prompt or query, and provided to the model as inference input. Furthermore, the knowledge base should be able to receive the latest data or information updates to ensure that the model references the latest and most comprehensive knowledge materials during inference. Furthermore, to ensure data security and privacy, access permissions can be set for the knowledge base to ensure that only authorized models or users can access specific knowledge content.

[0054] Optionally, when the target task is an inference task, the corresponding inference task execution request can be generated and sent by the inference session platform. The inference session platform is used to interact with the data processing device (such as a load balancing system) and the terminal device based on the inference service request sent by the terminal device during the inference service phase of the inference task to complete the execution of the inference task.

[0055] Optionally, during the inference service phase, the terminal device user initiates a high-concurrency inference service request to the inference session platform; the inference session platform queries the inference context from the state library; the inference session platform initiates a knowledge retrieval request to the user knowledge base based on the user's original input information prompt (i.e., the request information entered by the user, such as text, image, instruction, etc.); the state library returns the inference context information to the inference session platform; the knowledge base returns the knowledge retrieval results to the inference session platform; the inference session platform initiates an inference task execution request carrying the context information and knowledge retrieval results to the load balancing system.

[0056] In an optional embodiment, when the target task is a fine-tuning task, the task execution request is a fine-tuning task execution request, wherein the fine-tuning task execution request is used to request execution of the fine-tuning task, obtain the fine-tuning execution result, and the fine-tuned model file, and the fine-tuning execution result is at least used to indicate whether the fine-tuning task is executed successfully.

[0057] Optionally, in the case where the target task is a fine-tuning task, the corresponding received task execution request is a fine-tuning task execution request. The computing power node of the target task can execute the fine-tuning task for the large model based on the received fine-tuning task execution request and the relevant information pre-deployed on the computing power node (such as the source model file of the large model, the fine-tuning image file, and the data set required to execute the fine-tuning task, etc.), and obtain a fine-tuning execution result containing the fine-tuned model file of the large model. In the above manner, corresponding model files can be generated according to different fine-tuning task requirements, thereby improving the flexibility and adaptability of the fine-tuning task. Different fine-tuning tasks may require different degrees of adjustment to the model, and this method can better meet these adjustment requirements.

[0058] Optionally, when the target task is a fine-tuning task, the corresponding fine-tuning task execution request may be sent by the target interactive device (such as a fine-tuning deployment enabling platform). The fine-tuning deployment enabling platform is used to interact with the data processing device (such as a scheduling system) during the fine-tuning deployment phase and the fine-tuning service phase of the fine-tuning task to complete the execution of the fine-tuning task.

[0059] Step S104: obtaining a target computing node for executing the target task according to the target task, wherein the target computing node is a server device deployed in the cloud;

[0060] Optionally, the target computing node can be selected from multiple remote computing nodes, which are server devices deployed in the cloud. Flexibly calling remote computing nodes based on task requirements can improve the efficiency and flexibility of task execution based on large models.

[0061] In an optional embodiment, a target computing power node for executing a target task is obtained according to the target task, including: obtaining a candidate computing power node determined in advance based on a target scheduling scheme, wherein the target scheduling scheme is obtained based on computing power resource information and task requirement information matched with the target task; and obtaining a target computing power node based on the candidate computing power node.

[0062] Optionally, in the task deployment phase of the target task (such as the reasoning task deployment phase or the fine-tuning task deployment phase), a corresponding target scheduling scheme can be generated based on the computing power resource information and task requirement information matched by the target task. The target scheduling scheme may include relevant information of one or more candidate computing power nodes selected from multiple remote computing power nodes for executing the target task. Based on the target scheduling scheme, the candidate computing power nodes for executing the target task can be determined. In the task execution phase of the target task (such as the reasoning task execution phase or the fine-tuning task execution phase), the target computing power node for executing the target task can be determined based on the candidate computing power nodes. In the above process of determining the target computing power node, the computing power resource information and task requirement information matched by the target task are comprehensively considered, and the factors considered are more comprehensive. The target computing power nodes selected on this basis are also more in line with the needs of task execution, and are more conducive to the efficient and stable execution of the task.

[0063] Optionally, the computing resource information includes, but is not limited to, computing resource information such as computing power specifications (CPU, GPU, memory, etc.). Based on the target user's usage habits, computing resource information may also include any of the following items or combinations: computing resource location, network specifications (bandwidth, latency, packet loss rate, etc.), storage specifications, various resource prices, and other computing resource information. Task requirement information may include, but is not limited to, the model of the large model, network latency, cost, and the task deployment mode for the inference task. Similarly, those skilled in the art can comprehensively consider and select any item or combination of this task requirement information.

[0064] In an optional embodiment, obtaining candidate computing power nodes determined in advance based on a target scheduling scheme includes: receiving a task deployment request for a target task, wherein the task deployment request is an inference deployment request based on an inference task, or a fine-tuning deployment request based on a fine-tuning task; determining the target scheduling scheme based on the computing power resource information and task requirement information carried in the task deployment request; and determining the candidate computing power nodes based on the target scheduling scheme.

[0065] Optionally, during the task deployment phase (such as the inference task deployment phase or the fine-tuning task deployment phase), the target interactive device will send a task deployment request to the data processing device. The data processing device can perform computing power scheduling based on the relevant information carried in the task deployment request (such as computing power resource information required to execute the target task, task requirement information, etc.), and generate a corresponding target scheduling plan for screening candidate computing power nodes. In the above process of determining candidate computing power nodes, the computing power resource information and task requirement information matching the target task are comprehensively considered, and the factors considered are more comprehensive. The candidate computing power nodes screened on this basis are also more in line with the needs of task execution, which helps to ensure efficient and stable execution of tasks.

[0066] Optionally, when the task deployment stage is the inference task deployment stage and the corresponding task deployment request is the inference task deployment request, a scheduling system can be deployed in the data processing device, and a inference deployment enabling platform can be deployed in the target interactive device. The inference deployment enabling platform is used to interact with the data processing device and the target device (i.e., the administrator user device) for the inference deployment request of the inference task during the inference deployment stage of the inference task. Among them, when the task deployment stage is the inference task deployment stage, the inference deployment enabling platform can serve as a back-end device, and the target device can be a front-end device corresponding to the inference deployment enabling platform. Specifically, the target device sends an inference task deployment request to the inference deployment enabling platform, and forwards the inference task deployment request to the scheduling system through the inference deployment enabling platform; the scheduling system performs computing power scheduling based on the computing power resource information and task requirement information carried in the inference deployment request, and generates a target scheduling plan as a basis for screening candidate computing power nodes.

[0067] Optionally, when the task deployment stage is the fine-tuning task deployment stage and the corresponding task deployment request is a fine-tuning task deployment request, a scheduling system may be deployed in the data processing device, and a fine-tuning deployment enabling platform may be deployed in the target interactive device. Specifically, the target device (i.e., the administrator user device) sends a fine-tuning deployment request to the fine-tuning deployment enabling platform, and the fine-tuning deployment enabling platform forwards the fine-tuning deployment request to the scheduling system; the scheduling system performs computing power scheduling based on the computing power resource information and task requirement information carried in the fine-tuning deployment request, and generates a target scheduling plan as a basis for screening candidate computing power nodes. Specifically, when the task deployment stage is the fine-tuning task deployment stage, the fine-tuning deployment enabling platform can serve as a front-end device, and the target device can be a front-end device corresponding to the fine-tuning deployment enabling platform.

[0068] In an optional embodiment, when the task requirement information includes a task deployment mode, a target scheduling scheme is determined based on the computing power resource information carried in the task deployment request and the task requirement information, including: determining an initial scheduling scheme based on the computing power resource information carried in the task deployment request, wherein the initial scheduling scheme includes the initially determined information of the computing power nodes to be scheduled for the target task; obtaining the task deployment mode of the target task; when the task deployment mode is a first mode, determining a target scheduling scheme based on the initial scheduling scheme, wherein the first mode is used to indicate that there is no confirmation process for the scheduling scheme; or when the task deployment mode is a second mode, sending a scheme confirmation instruction to the target device, wherein the second mode is used to indicate that there is a confirmation process for the scheduling scheme, and the scheme confirmation instruction is used to indicate whether the scheduling scheme is confirmed; and determining the target scheduling scheme based on the scheme confirmation result returned by the target device.

[0069] Optionally, after determining the initial scheduling plan based on the computing resource information carried in the task deployment request and the task requirement information, it is necessary to further optimize the initial scheduling plan according to the task deployment mode of the target task to obtain the target scheduling plan. The task deployment mode includes a first mode and a second mode, which are used to indicate whether there is a scheduling plan confirmation link, that is, whether the determined initial scheduling plan is reconfirmed through manual intervention. In the first mode, there is no need for manual intervention (such as administrator user device intervention) to perform the scheduling plan confirmation process. This mode can be understood as a "worry-free mode". In the "worry-free mode", the target scheduling plan is directly determined based on the initial scheduling plan; in the second mode, manual intervention and the initial scheduling plan confirmation process are required. This mode can be understood as a "rest assured mode". In the "rest assured mode", the target interactive device sends the initial scheduling plan to the administrator user device, and the administrator user confirms the scheduling plan through the target interactive device (such as the inference deployment enabling platform, the fine-tuning deployment enabling platform, etc.), and determines the target scheduling plan based on the plan confirmation result returned by the target device. By setting up "worry-free mode" and "rest assured mode", user participation can be increased, and the target scheduling plan can be flexibly determined according to user needs.

[0070] Optionally, when there is only one initial scheduling plan, if the task deployment mode is "worry-free mode", the initial scheduling plan will be directly used as the target scheduling plan; if the task deployment mode is "safe mode", if the plan confirmation result returned by the target device indicates that the initial scheduling plan is passed, the initial scheduling plan will be used as the target scheduling plan.

[0071] Optionally, when there are multiple initial scheduling plans, if the task deployment mode is "worry-free mode", at least one target scheduling plan is automatically selected from the multiple initial scheduling plans; if the task deployment mode is "safe mode", the plan confirmation result carries the scheduling plan selected from the multiple initial scheduling plans; the selected scheduling plan is determined as the target scheduling plan.

[0072] In an optional embodiment, the method also includes: when the target task is an inference task and there are multiple candidate computing nodes, binding multiple candidate computing nodes in a serverless placeholder manner, and synchronizing the model files and image files required to execute the target task to the multiple candidate computing nodes.

[0073] Optionally, during the inference task deployment phase, multiple candidate computing nodes can be identified. After obtaining multiple candidate computing nodes, the scheduling system can bind the resource information in the target scheduling plan (i.e., multiple candidate computing nodes) in a serverless placeholder manner, open the corresponding wide area network transmission tunnel, and synchronize the model files and image files required to execute the target task to multiple candidate computing nodes.

[0074] In an optional embodiment, after binding multiple candidate computing nodes in a serverless placeholder manner, the method further includes: deploying an initial inference instance (also known as a veteran instance) in one of the multiple candidate computing nodes.

[0075] Optionally, the scheduling system can also initiate deployment requests for veteran instances to computing nodes. For a serverless computing resource cluster with available network resources, only one veteran instance (the initial instance for inference) needs to be deployed. This reduces resource waste when there is no business traffic access and reduces cold start latency for initial user access. Alternatively, if local computing nodes are deployed, there's no need to deploy veteran instances on multiple candidate computing nodes.

[0076] It should be noted that in this embodiment, during the inference deployment phase, when there are multiple available computing nodes (i.e., multiple candidate computing nodes), the scheduling system only deploys one inference instance (i.e., deploys the initial inference instance) on the serverless placeholder's available computing nodes across the network. This reduces resource waste when there is no business traffic access and reduces cold start latency for initial user access. During the inference deployment phase, other serverless placeholder resources only synchronize the model files and image files corresponding to the large oracle model, i.e., the "network-follows-computing" mode.

[0077] Optionally, before synchronizing the model files and image files required for executing the target task (such as the reasoning task), it is possible to detect whether the model files and image files required for executing the target task already exist in the candidate computing power node in the target scheduling scheme. If the model files and image files required for executing the target task do not exist in the candidate computing power node, the model files and image files required for executing the target task need to be synchronized to the candidate computing power node. Otherwise, there is no need to synchronize the model files and image files required for executing the target task to the candidate computing power node. Specifically, when the model files or image files required for executing the target task do not exist in the candidate computing power node, the scheduling system can request the reasoning deployment enabling platform to synchronize (copy) the model files or image files required for executing the target task to the computing power node in the scheduling scheme. Among them, the model files and image files required for executing the target task need to be de-domained under the premise that the administrator user confirms that there is no privacy risk, and the image files need to be adapted to the computing power resources in advance. The reasoning deployment enabling platform requests the model management system to synchronize the model files or image files to the candidate computing power nodes that are missing the model files or image files. The model management system synchronizes the model files or image files to the candidate computing power nodes that are missing the model files or image files. The model management system returns the synchronization results of the model file and the image file to the inference deployment enabling platform. The inference deployment enabling platform returns the synchronization results of the model file and the image file to the scheduling system.

[0078] Optionally, after the deployment of the veteran instance is completed, all candidate computing nodes can also return the veteran instance deployment results to the scheduling system; the scheduling system synchronizes the inference deployment results to the load balancing system, including computing node routing information, as well as monitoring information such as resource status and inference engine deployment status; the scheduling system returns the inference deployment results to the inference deployment enabling platform; the inference deployment enabling platform presents the inference deployment results to the administrator user.

[0079] In an optional embodiment, before receiving a task execution request for a target task, the method further includes: in the case where the target task is a fine-tuning task, initiating a file preparation request to the target interactive device, so that the target interactive device performs a file synchronization operation based on the file preparation request to obtain a file preparation result, wherein the file preparation result is used to send to the target computing power node, the file synchronization operation is used to synchronize the source model file of the large model, the fine-tuning image file, and the data set required to execute the fine-tuning task to the target computing power node, and the file preparation result is used to indicate the execution result of the file synchronization operation.

[0080] Optionally, during the fine-tuning task deployment phase, after the target scheduling plan is confirmed, the data processing equipment (such as the scheduling system) binds the resource information in the scheduling plan (i.e., the candidate computing power nodes) and opens the corresponding wide area network transmission tunnel; the scheduling system initiates a file preparation request to the fine-tuning deployment enabling platform, requesting the fine-tuning deployment enabling platform to synchronize (copy) the fine-tuning source model file, fine-tuning mirror file, and the user data set used for fine-tuning to the candidate computing power nodes in the target scheduling plan. File out-of-domain needs to be carried out under the premise that the administrator user confirms that there is no privacy risk, and the mirror file needs to be adapted to the computing power resources in advance. The above method realizes file synchronization through the "network follows computing" method, that is, after the computing power node required to execute the fine-tuning task is determined through computing power scheduling, the fine-tuning source model file, fine-tuning mirror file, and the user data set used for fine-tuning are synchronized (copied) to the computing power node.

[0081] Optionally, the scheduling system initiates a file preparation request to the fine-tuning deployment enabling platform, requesting the enabling platform to synchronize (copy) the fine-tuning source model file, fine-tuning image file, and user data set used for fine-tuning to the computing power nodes in the scheduling plan. File out-of-domain needs to be carried out under the premise that the administrator user confirms that there is no privacy risk, and the image file needs to be adapted to the computing power resources in advance. Specifically, the fine-tuning deployment enabling platform requests the model management system to synchronize the model file or fine-tuning image file to the node that is missing the source model file or fine-tuning image file of the large model; the fine-tuning deployment enabling platform requests the user data center to synchronize the user local data set used for this fine-tuning to the remote computing power node (i.e., the target computing power node); the model management system synchronizes the source model file or fine-tuning image file to the target computing power node that is missing the source model file or fine-tuning image file; the user data center synchronizes the user local data set used for this fine-tuning to the target computing power node.

[0082] In an optional embodiment, a target computing power node is obtained based on the candidate computing power node, including: when the target task is an inference task and there are multiple candidate computing power nodes, the target computing power node is obtained in the following manner: a load balancing method is used to determine the target computing power node from multiple candidate computing power nodes; or when the target task is a fine-tuning task, the target computing power node is obtained in the following manner: the candidate computing power node is used as the target computing power node.

[0083] Optionally, during the execution of an inference task, a load balancing method can be used to flexibly determine the remote computing node (i.e., target computing node) required to execute the target task from multiple candidate computing nodes. This allows for flexible call-up of remote computing resources to improve the efficiency and flexibility of large-model inference tasks. During the execution of a fine-tuning task, the candidate computing node selected through computing power scheduling is used as the target computing node to execute the inference task, thereby achieving efficient allocation of computing resources required to execute the fine-tuning task.

[0084] Optionally, during the inference service phase, the load balancing system can provide global routing information including remote computing nodes for business traffic (inference service requests) from terminal devices based on the configured load balancing strategy, determine the target computing node based on the global routing information, and request the scheduling system to open a wide area network transmission tunnel; the scheduling system opens the wide area network transmission tunnel and returns the wide area network transmission tunnel opening result to the load balancing system.

[0085] In an optional embodiment, obtaining a target computing node for executing a target task includes: using a load balancing method to determine the target computing node from multiple candidate computing nodes, and determining task execution requests corresponding to the target computing node and the local computing node respectively.

[0086] Optionally, when the target task is an inference task, the execution of the inference task can also be completed through the collaboration of local computing nodes and remote computing nodes. In this way, on the basis of fully utilizing local computing resources, remote computing nodes can be flexibly called to perform business execution, so as to effectively improve the efficiency and stability of task execution. Specifically, in the inference service stage, the load balancing system can provide global routing information including local computing nodes and remote computing nodes for business traffic (inference service requests) from terminal devices according to the configured load balancing strategy, determine the target computing node based on the global routing information, and request the scheduling system to open a wide area network transmission tunnel; the scheduling system opens the wide area network transmission tunnel and returns the wide area network transmission tunnel opening result to the load balancing system. Optionally, the load balancing method can be used to determine the target computing node from multiple candidate computing nodes based on, but not limited to, the amount of computing resources required to execute the target task, the location distance between the candidate computing node and the local computing node, and the health status of the candidate computing node.

[0087] Optionally, the local computing power node is a server device deployed on the local side, which can be, but is not limited to, a large model training and pushing all-in-one machine, a server with AI computing capabilities (including a graphics processing unit (GPU) card), and a bare metal server.

[0088] In an optional embodiment, obtaining a target computing power node for executing the target task according to the target task includes: when the target computing power resource amount exceeds the computing power resource ownership of the local computing power node, obtaining the target computing power node for executing the target task according to the target task, wherein the target computing power resource ownership is the computing power resource amount required to execute the target task, and the local computing power node is a server device deployed on the local side; or when the computing power resource occupancy rate of the local computing power node is greater than a preset threshold, obtaining the target computing power node for executing the target task according to the target task; or when the waiting time required for executing the target task at the local computing power node is greater than a predetermined time, obtaining the target computing power node for executing the target task according to the target task; or when there is no local computing power node for executing the target task (such as the local computing power node does not support executing the target task), obtaining the target computing power node for executing the target task according to the target task.

[0089] Optionally, in the presence of a local computing power node to execute the target task, when the access concurrency of the terminal user's target service (such as an inference service or a fine-tuning service, etc.) exceeds the upper limit of the local computing power resource carrying capacity, or the computing power resource occupancy rate of the local computing power node exceeds the upper limit, or the local computing power resources are insufficient to support the specifications of the target service model selected by the user, or the waiting time required to execute the target task at the local computing power node is too long, the target service will face problems such as response delay and limited service scope, affecting the user experience. By calling remote computing power nodes to execute tasks, the scale and capability range of AI computing power resources can be expanded by integrating cloud computing power, and the most cost-effective cloud computing power resources can be matched for the execution of large model tasks. In the absence of a local computing power node to execute the target task, the remote computing power node can be flexibly determined based on the computing power resource information and task requirement information required to execute the target task, so as to match the most cost-effective cloud computing power resources to execute the target task.

[0090] Step S106: Send the task execution request to the target computing power node, so that the target computing power node executes the target task based on the task execution request.

[0091] Optionally, a data processing device (such as a scheduling system or a load balancing system) sends a task execution request to a selected remote computing power node (i.e., a target computing power node). The target computing power node can execute the target task according to the task execution request, and return the task execution result to the corresponding device, thereby completing the flexible scheduling of computing power resources and the efficient execution of large model tasks.

[0092] Optionally, when the target task is an inference task, the target computing power node can execute the inference task based on the received inference task execution request, and return the obtained inference service result to the data processing device (such as a load balancing system), and the data processing device can return the inference service result to the terminal device via the target interactive device (such as an inference session platform).

[0093] Optionally, when the target task is a fine-tuning task, the target computing power node can execute the fine-tuning task based on the received fine-tuning task execution request, and return the obtained fine-tuning execution result to the data processing device (such as the scheduling system), and the fine-tuning execution result is at least used to indicate whether the fine-tuning task is executed successfully; the data processing device further sends the fine-tuning execution result to the target interactive device (such as the fine-tuning deployment enabling platform), and presents the fine-tuning execution result to the administrator user through the target interactive device.

[0094] In an optional embodiment, a task execution request is sent to a target computing power node, so that the target computing power node executes the target task based on the task execution request, including: sending a corresponding task execution request to the target computing power node and the local computing power node, so that the target computing power node and the local computing power node respectively execute the target task based on the corresponding task execution request.

[0095] Optionally, when the target task is an inference task, when it is determined through the load balancing method that the inference task is executed simultaneously through the target computing power node and the local computing power node, the inference task execution requests allocated to the target computing power node and the local computing power node are sent to the corresponding computing power node. The target computing power node and the local computing power node execute the inference task based on the task execution requests they receive, and return the corresponding inference deployment results to the corresponding device. The above method integrates cloud computing power, local computing power nodes (such as idle all-in-one machines) and other network-wide computing power resources, expands the scale and capability range of AI computing power resources, breaks the business model of pure all-in-one machines, self-built and self-sold cloud computing power, and connecting to specific third-party computing power, so as to improve the execution efficiency and flexibility of inference business.

[0096] In an optional embodiment, a task execution request is sent to a target computing power node, so that the target computing power node executes the target task based on the task execution request, including: when the target task is an inference task and the corresponding task execution request is an inference task execution request, after distributing the task execution request to the target computing power node through a load balancing method, an inference task instance corresponding to the inference task is deployed on the target computing power node, so that the target computing power node executes the inference task.

[0097] Optionally, since only one computing node is deployed with a veteran instance during the inference deployment phase, during the inference service phase, the load balancing system distributes the terminal user's business traffic (such as task execution requests) to the target computing node before the scheduling system completes the actual deployment of the inference task instance on the target computing node, that is, the inference service is completed in a "computing follows the network" manner, thereby improving the execution efficiency and flexibility of the inference task while ensuring the reduction of computing resource usage.

[0098] In an optional embodiment, a task execution request is sent to a target computing power node, and the target computing power node executes the target task based on the task execution request, including: when the target task is a fine-tuning task and the corresponding task execution request is a fine-tuning task execution request, a fine-tuning instance of the fine-tuning task is deployed to the target computing power node, and the target computing power node executes the fine-tuning task, obtains a fine-tuning execution result and a fine-tuned model file, and returns the fine-tuning execution result to the target device, and returns the fine-tuned model file to the model management device, wherein the fine-tuning execution result is at least used to indicate whether the fine-tuning task is successfully executed.

[0099] Optionally, during the fine-tuning service phase, the fine-tuning instance of the fine-tuning task is first deployed to the target computing node. The target computing node then executes the fine-tuning task and returns the fine-tuning execution results obtained from executing the fine-tuning task to the target device. This allows the "network-follows-computing" model (first scheduling computing resources and deploying instances, then transmitting business traffic (such as fine-tuning execution results and fine-tuned model files)) to minimize the use of computing network resources while ensuring efficient execution of fine-tuning tasks.

[0100] In an optional embodiment, when the target computing power node is selected from multiple candidate computing power nodes and one of the multiple candidate computing power nodes deploys an inference initial instance, the method further includes: detecting whether the deployment of the inference task instance on the target computing power node is completed; when the deployment of the inference task instance on the target computing power node is not completed, distributing the task execution request to the computing power node deployed with the inference initial instance, so that the computing power node deployed with the inference initial instance executes the inference task; when the deployment of the inference task instance on the target computing power node is completed, the target task is jointly executed by the computing power node deployed with the inference initial instance and the target computing power node.

[0101] Optionally, if a veteran instance has been deployed during the inference deployment phase, when the process of deploying the inference task instance on the target computing power node has not yet been completed, and the inference task execution request has been distributed to the cloud, the veteran instance will first generate the inference service result. After the inference task instance is deployed on the target computing node, the newly deployed inference task instance and the veteran instance will jointly provide inference services to generate remote inference service results. This ensures that even during the instance deployment process, timely execution of tasks can be achieved through the veteran instance, avoiding delays in inference task execution due to instance deployment and improving the efficiency of inference task execution. If the veteran instance is not deployed during the inference deployment phase, when the inference task instance of the target computing power node has not yet been deployed and the inference task execution request has been distributed to the cloud, wait until the inference task instance of the target computing power node is deployed before generating the remote inference service result.

[0102] As an optional embodiment, a task execution request is sent to a target computing power node, so that the target computing power node executes the target task based on the task execution request, including: determining a computing-network collaboration mode for the target task, wherein the computing-network collaboration mode is used to indicate a mode for collaboratively scheduling computing power resources and network resources of the computing power node; and sending a task execution request to the target computing power node, so that the target computing power node executes the target task based on the task execution request in accordance with the computing-network collaboration mode.

[0103] Optionally, when the task type (or task request type) of the target task is different, there are certain differences in the corresponding selected computing network collaboration mode. For example, when the target task is an inference task and the corresponding task execution request is an inference task execution request, the inference task is executed by using the computing network dynamic mode, that is, first distributing the business traffic, and then deploying the inference instance on the corresponding computing power node to which the business traffic is distributed; when the target task is a fine-tuning task and the corresponding task execution request is a fine-tuning task execution request, the network follows the computing dynamic mode, that is, first scheduling the computing power node, and deploying the fine-tuning instance of the fine-tuning task to the corresponding computing power node to execute the fine-tuning task, and then returning the fine-tuning execution result and the fine-tuned model file. In this way, the corresponding computing network collaboration mode is selected in a targeted manner according to different types of large model tasks to execute the target task, thereby improving the execution efficiency and execution flexibility of large model tasks.

[0104] In an optional embodiment, when the target task is an inference task, after sending the task execution request to the target computing power node for the target computing power node to execute the target task based on the task execution request, the method also includes: obtaining the computing power resource occupancy rate of the target computing power node; when the computing power resource occupancy rate is lower than a preset lower limit value, controlling the target computing power node to release part of the computing power resources used to execute the target task; or when the computing power resource occupancy rate is higher than a preset upper limit value, determining a first computing power node from multiple candidate computing power nodes, and deploying an inference task instance corresponding to the inference task on the first computing power node, wherein the first computing power node is used to execute the target task together with the target computing power node.

[0105] Optionally, the computing resource utilization rate may be the computing resource utilization rate of the inference task instance in the target computing node. If the load balancing system detects a decrease in the computing resource utilization rate of the inference task instance in the target computing node, and if the computing resource utilization rate drops below the resource threshold, the load balancing system requests elastic scaling of computing resources from the scheduling system. This involves releasing the instance resources whose resource utilization rate is below the resource threshold and distributing the traffic carried by them to other instances for execution. If the load balancing system detects an increase in computing resource utilization rate, and if the utilization rate rises above the resource threshold, the load balancing system requests elastic scaling of computing resources from the scheduling system. This involves selecting a new computing node (i.e., the first computing node) within the serverless resource range and deploying a new inference instance on this node to execute the target task alongside the target computing node. Flexible scaling of computing resources based on computing resource utilization ensures efficient execution of inference tasks while effectively improving computing resource utilization.

[0106] In an optional embodiment, obtaining the computing resource occupancy rate of the target computing node includes: monitoring the concurrency of task execution requests received by the target computing node; and determining the computing resource occupancy rate based on the concurrency.

[0107] Optionally, when the concurrency of inference service requests sent by end users decreases, the load balancing system monitors and perceives a decrease in the computing resource utilization rate of the inference task instances in the target computing node, and can elastically scale down the computing resources. When the concurrency of inference service requests sent by end users increases, the load balancing system monitors and perceives an increase in the computing resource utilization rate of the inference task instances in the target computing node, and can elastically scale up the computing resources. Therefore, the computing resource utilization rate can be flexibly determined based on the concurrency of task execution requests received by the known target computing node, and changes in computing resource utilization can be timely perceived, allowing for timely elastic scaling of computing resources.

[0108] Optionally, when it is detected that the concurrency of task execution requests received by the target computing power node is reduced, part of the computing power resources occupied by the target computing power node for executing the target task can be released in a timely manner, thereby ensuring that the large model inference task is completed efficiently and stably while minimizing the occupation of remote computing power resources. Specifically, when it is detected that the concurrency of inference service requests from terminal device users is reduced and the load balancing system perceives that there is no inference service traffic at the target computing power node, the load balancing system requests the scheduling system to reduce the computing power resources; the computing power scheduling system requests the target computing power node to reduce the capacity; and the target computing power node releases part or all of the computing power resources occupied by the current inference task based on the received request.

[0109] As an optional embodiment, when the target task is an inference task, after executing the target task based on the task execution request, the computing resources occupied by executing the target task are released, and the user data set, model file, and fine-tuning image file uploaded to the remote computing node (i.e., the target computing node) for this execution of the target task are cleared. This allows the relevant computing resources to be released in a timely manner after the task is completed, minimizing the computing resource usage.

[0110] Through the above steps S102 to S106, the purpose of intelligent scheduling to automatically identify and select remote (target) computing nodes to perform tasks can be achieved, thereby realizing the technical effect of improving the efficiency, flexibility and stability of task execution based on large models, and then solving the technical problems of low task execution efficiency and poor flexibility in task execution methods based on large models in related technologies.

[0111] According to an embodiment of the present invention, another method embodiment of task execution based on a large model is provided. Figure 2 is a flowchart of another task execution method based on a large model according to an embodiment of the present invention. Figure 2 As shown, the method includes the following steps:

[0112] Step S202: Receive a task execution request for a target task, where the target task is an inference task based on a large model or a fine-tuning task based on a large model. The task execution request is sent by a data processing device, and the local computing power node is a server device deployed locally.

[0113] Step S204: executing the target task based on the task execution request.

[0114] Optionally, the execution subject of steps S202 to S204 can be the target computing node. The target computing node can execute the target task according to the received task execution request and return the corresponding task execution result to the relevant device. Through the above steps S202 to S204, the purpose of intelligent scheduling to automatically identify and select remote (target) computing nodes to execute tasks can be achieved, thereby achieving the technical effect of improving the task execution efficiency and execution flexibility based on the large model, and thus solving the technical problems of low task execution efficiency and poor flexibility in the task execution method based on the large model in the related art.

[0115] Optionally, when the target task is an inference task, the target computing power node can execute the inference task based on the received inference task execution request, and return the obtained inference service result to the data processing device (such as the load balancing system and scheduling system), and the data processing device can return the inference service result to the terminal device via the target interactive device (such as the inference session platform).

[0116] Optionally, when the target task is a fine-tuning task, the target computing power node can execute the fine-tuning task based on the received fine-tuning task execution request, and return the obtained fine-tuning model file to the model management device, and return the obtained fine-tuning execution result to the data processing device (such as the scheduling system). The fine-tuning execution result is at least used to indicate whether the fine-tuning task is executed successfully; the data processing device further sends the fine-tuning execution result to the target interactive device (such as the fine-tuning deployment enabling platform), and presents the fine-tuning execution result to the administrator user through the target interactive device.

[0117] In an optional embodiment, executing a target task based on a task execution request includes: when the target task is an inference task and the task execution request is an inference task execution request, executing the inference task based on the inference task execution request to obtain an inference service result, wherein the inference task execution request carries a task execution request with inference context information and a knowledge retrieval result, the inference context information is based on the inference service request for the target task initiated by the terminal device and is obtained by querying a state library; the knowledge retrieval result is based on the inference service request and is obtained by querying a knowledge library; the inference service request carries request information input through the terminal device, wherein the request information includes at least one of the following: text, instructions, images; and the inference service result is sent to a data processing device so that the data processing device sends the inference service result to the terminal device via the target interactive device.

[0118] Optionally, when the target task is an inference task, the task execution request received by the target computing power node is an inference task execution request. The inference context information and knowledge retrieval results carried in the inference task execution request are based on the inference service request initiated by the terminal device and are obtained by querying the state library and knowledge library. This method can dynamically obtain relevant information based on different inference service requests, thereby improving the flexibility and adaptability of inference tasks. Different inference service requests may involve different context and knowledge requirements, and this method can better meet these requirements.

[0119] Optionally, when the target task is an inference task, the corresponding inference task execution request can be generated by the inference session platform and sent to the data processing device (such as a load balancing system), and then sent to the target computing power node through the data processing device. The inference session platform is used to interact with the data processing device and the terminal device based on the inference service request sent by the terminal device during the inference service phase of the inference task to complete the execution of the inference task.

[0120] Optionally, during the inference service phase, the user of the terminal device initiates a high-concurrency inference service request to the target interactive device (such as the inference session platform); the inference session platform queries the state library for the inference context; the inference session platform initiates a knowledge retrieval request to the user knowledge library based on the user's original input prompt (i.e., the request information entered by the user); the state library returns the inference context information to the inference session platform; the knowledge library returns the knowledge retrieval results to the inference session platform; the inference session platform initiates an inference task execution request carrying the context information and knowledge retrieval results to the data processing device (such as a load balancing system); and the inference task execution request is sent to the target computing power node via the data processing device. Optionally, after the inference task is executed and the inference service result is obtained, the user's original input prompt (i.e., the request information entered by the user) and the inference context information can be refreshed to the state library accordingly.

[0121] In an optional embodiment, executing a target task based on a task execution request includes: when the target task is a fine-tuning task and the task execution request is a fine-tuning task execution request, executing the fine-tuning task based on the task execution request, obtaining a fine-tuning execution result and a fine-tuning model file, wherein the fine-tuning execution result is at least used to indicate whether the fine-tuning task is executed successfully; sending the fine-tuning execution result to a data processing device, so that the data processing device sends the fine-tuning execution result to the target device via the target interactive device, and sending the fine-tuning model file to the model management system device.

[0122] Optionally, in the case where the target task is a fine-tuning task, the task execution request received by the target computing power node is a fine-tuning task execution request. The target computing power node can execute the fine-tuning task for the large model based on the fine-tuning task execution request and the relevant information pre-deployed on the computing power node (such as the source model file of the large model, the fine-tuning image file, and the data set required to execute the fine-tuning task, etc.), and obtain the fine-tuned model file and the fine-tuning execution result. In the above manner, corresponding model files can be generated according to different fine-tuning task requirements, thereby improving the flexibility and adaptability of the fine-tuning task. Different fine-tuning tasks may require different degrees of adjustment to the model, and this method can better meet these adjustment requirements.

[0123] Optionally, when the target task is a fine-tuning task, the corresponding fine-tuning task execution request can be generated and sent by the target interactive device (such as a fine-tuning deployment enabling platform). The fine-tuning deployment enabling platform is used to interact with the data processing device (such as a scheduling system) during the fine-tuning deployment phase and the fine-tuning service phase of the fine-tuning task to complete the execution of the fine-tuning task.

[0124] According to an embodiment of the present invention, another method embodiment of task execution based on a large model is provided. Figure 3 is a flowchart of another method for executing a task based on a large model according to an embodiment of the present invention. Figure 3 As shown, the method includes the following steps:

[0125] Step S302: Receive a service request for a target task, wherein the target task is a large-model-based reasoning task and the service request is a reasoning service request for the reasoning task; or the target task is a large-model-based fine-tuning task and the service request is a fine-tuning deployment request for the fine-tuning task;

[0126] Step S304: Generate a task execution request based on the service request;

[0127] Step S306: Send the task execution request to the data processing device, so that the data processing device sends the task execution request to the target computing power node, and executes the target task through the target computing power node.

[0128] The execution subject of the above steps S302 to S306 can be a target interactive device, which serves as a communication bridge between the user end (such as a terminal user or an administrator user), the data processing device (such as a scheduling device, a load balancing device) and other devices. On the one hand, it can receive the service request sent by the user end and generate a corresponding execution request based on the service request, wherein the service request carries the original input information prompt (i.e., the request information input by the user); on the other hand, it can also send the generated task execution request to the data processing device to realize the scheduling of computing resources and the execution of large model tasks. Through the participation of the target interactive device, the purpose of intelligent scheduling to automatically identify and select remote (target) computing nodes to execute tasks can be achieved, thereby achieving the technical effect of improving the efficiency and flexibility of task execution based on large models, and then solving the technical problems of low task execution efficiency and poor flexibility in the task execution method based on large models in related technologies.

[0129] In an optional embodiment, a task execution request is generated based on a service request, including: when the target task is an inference task, the service request is an inference service request for the inference task, and the task execution request is an inference task execution request, the inference service request is sent to the state library and the knowledge library respectively; receiving the inference context information returned by the state library based on the inference service request, and the knowledge detection result returned by the knowledge library; and generating an inference task execution request based on the inference context information and the knowledge detection result.

[0130] Optionally, after receiving the reasoning service request sent by the terminal device, the target interactive device (such as the reasoning session platform) will send the reasoning service request to the state library and the knowledge library respectively. The state library can query the reasoning context information corresponding to the reasoning task according to the reasoning service request, and return the queried reasoning context information to the target interactive device; the knowledge library can query the knowledge detection result corresponding to the reasoning service according to the reasoning service request, and return the queried knowledge detection result to the target interactive device; the target interactive device can generate a reasoning task execution request based on the reasoning context information and the knowledge detection result. In the above manner, relevant information can be dynamically obtained according to different reasoning service requests, and reasoning task execution requests can be generated in a targeted manner, thereby improving the flexibility and adaptability of reasoning tasks. Different reasoning service requests may involve different contexts and knowledge requirements, and this method can better meet these requirements.

[0131] In an optional embodiment, a task execution request is generated based on a service request, including: when the target task is a fine-tuning task, the service request is a fine-tuning deployment request for the fine-tuning task, and the task execution request is a fine-tuning task execution request, a file preparation request sent by a data processing device is received, wherein the file preparation request is used to request execution of a file synchronization operation, and the file synchronization operation is used to synchronize the source model file of the large model, the fine-tuning image file, and the data set required to execute the fine-tuning task to the target computing power node; based on the execution result of the file synchronization operation, a file preparation result is obtained, and the file preparation result is sent to the data processing device, so that the data processing device generates a fine-tuning task execution request based on the file preparation result.

[0132] Optionally, after receiving the fine-tuning deployment request sent by the target device (i.e., the administrator user) and completing the scheduling of computing resources (i.e., the selection of the target computing node), the scheduling system initiates a file preparation request to the fine-tuning deployment enabling platform, requesting the enabling platform to synchronize (copy) the fine-tuning source model file, fine-tuning image file, and user data set used for fine-tuning to the computing node in the scheduling plan. File out-of-domain needs to be carried out under the premise that the administrator user confirms that there is no privacy risk, and the image file needs to be adapted to the computing resources in advance. Specifically, the fine-tuning deployment enabling platform requests the model management system to synchronize the model file or fine-tuning image file to the node that is missing the source model file or fine-tuning image file of the large model; the fine-tuning deployment enabling platform requests the user data center to synchronize the user local data set used for this fine-tuning to the remote computing node (i.e., the target computing node); the model management system synchronizes the source model file or fine-tuning image file to the target computing node that is missing the source model file or fine-tuning image file; the user data center synchronizes the user local data set used for this fine-tuning to the target computing node; the model management system returns the source model file or fine-tuning image file to the fine-tuning deployment enabling platform. The user data center returns the synchronization result of the data set to the fine-tuning deployment enabling platform; when the file synchronization result indicates that the file synchronization is completed, a fine-tuning task execution request is generated and sent to the data processing device (such as a scheduling system), and the fine-tuning task execution request is sent to the target computing power node via the data processing device, so that the target computing power node executes the fine-tuning task for the large model based on the received fine-tuning task execution request and the relevant information pre-deployed on the computing power node (such as the source model file of the large model, the fine-tuning image file, and the data set required to execute the fine-tuning task, etc.), and obtains the fine-tuned model file and the fine-tuning execution result.

[0133] In an optional embodiment, after sending a task execution request to a data processing device, for the data processing device to send a task execution request to a target computing power node when the local computing power node does not meet the preset execution conditions of the target task, for the target computing power node to execute the target task, the method also includes: when the target task is an inference task and the task execution request is an inference task execution request, receiving an inference service result obtained by the target computing power node executing the inference task, and sending the inference service result to the terminal device.

[0134] Optionally, after the target computing power node completes the inference task based on the inference task request, it will return the obtained inference service result to the data processing system (such as a load balancing system); the inference service result will be sent to the target interactive device (such as an inference session platform) through the data processing system; the target interactive device serves as a medium for interaction with the terminal device, and can synchronize the obtained inference service result to the terminal device, thereby completing the execution of the inference task and the synchronization of the results.

[0135] In an optional embodiment, after sending a task execution request to a data processing device, for the data processing device to send the task execution request to a target computing power node when the local computing power node does not meet the preset execution conditions of the target task, for the target computing power node to execute the target task, the method further includes: when the target task is a fine-tuning task and the task execution request is a fine-tuning task execution request, receiving a fine-tuning execution result obtained by the target computing power node executing the fine-tuning task; and sending the fine-tuning execution result to the target device, wherein the fine-tuning execution result is at least used to indicate whether the fine-tuning task is executed successfully.

[0136] Optionally, after the target computing power node executes the fine-tuning task based on the fine-tuning task execution request and obtains the fine-tuning execution result including the fine-tuning model file, the fine-tuning model file is transmitted back from the remote computing power node to the user model management system; the resources related to this fine-tuning are released from the remote computing power node, and the user data set, model file and fine-tuning image file uploaded to the remote computing power node for this task are cleaned up; the target computing power node returns the fine-tuning execution result to the scheduling system; the scheduling system returns the fine-tuning execution result to the target interactive device (such as the fine-tuning deployment enabling platform), and the fine-tuning deployment enabling platform presents the fine-tuning execution result to the administrator user, thereby completing the fine-tuning of the large model.

[0137] Based on the above embodiments and optional embodiments, the present invention proposes an optional implementation method of a task execution method based on a large model, which is used to execute reasoning tasks based on a large model. The reasoning business of cloud-edge collaboration (i.e., collaboration between cloud servers and local servers) includes two stages, namely, the reasoning task collaborative scheduling and deployment stage (referred to as the reasoning deployment stage) and the load balancing reasoning service stage (the reasoning service stage). In the reasoning deployment stage, the administrator user creates the reasoning deployment task, and the computing network scheduling system performs comprehensive scheduling based on the computing network resource specifications matched with the business, as well as the expected cost value, computing power resource specification expected value, network delay expected value, computing power resource location, number of queries per second (QPS) / number of queries per minute (QPM), operator information and other task demand information entered by the user on the enabling platform, and forms a corresponding scheduling plan. According to the scheduling plan, multiple available computing power nodes (i.e., multiple candidate computing power nodes) for executing the reasoning task are determined, and the selected multiple available computing power nodes are bound to the scheduling system in a serverless placeholder manner, and the corresponding resource information is registered with the load balancing system. During the inference service phase, the end user initiates an inference session, and the load balancing system diverts the end user's business traffic to local computing power and remote computing power based on load balancing strategies such as the minimum path.

[0138] During the inference deployment phase, when there are multiple available computing nodes, the scheduling system can actually deploy only one initial inference instance on the serverless placeholder's available computing nodes in the entire network (that is, deploy the veteran instance on only one computing node) to reduce resource waste when there is no business traffic access, while also reducing the cold start latency of the user's first access. In particular, if there are local computing nodes that support the execution of inference tasks, the deployment of veteran instances is also optional. During the inference deployment phase, other serverless placeholder resources only complete the synchronization of the model files and image files corresponding to the large prediction model. The scheduling system will not complete the actual deployment of the inference engine image on the node until the load balancing system distributes the end user's business traffic to the node during the inference service phase, which is the "calculation follows the network" mode.

[0139] When the user's local computing resources are sufficient to support the access concurrency and service quality requirements of the inference service terminal users, priority will be given to inference deployment on low-cost, secure and reliable local computing nodes, and inference services will be provided to the terminal users. When the terminal user's inference service access concurrency exceeds the upper limit of the local computing resources, or the local computing resources are insufficient to support the specifications of the inference service model selected by the user, the inference service will face problems such as response delays and limited service scope, affecting the user experience. At this time, the administrator user can obtain business needs such as cloud acceleration and service scope expansion by viewing monitoring data, direct feedback from terminal users, and querying the large model computing network demand map to determine whether it is necessary to scale computing resources, and initiate cloud-edge collaborative inference deployment tasks. Among them:

[0140] Figure 4 FIG. 1 is a flow chart of an optional reasoning deployment phase of a reasoning task according to an embodiment of the present invention. Figure 4 As shown in the figure, the reasoning deployment process based on the reasoning task specifically includes:

[0141] S11: The administrator user initiates a cloud-edge collaborative inference deployment request on the inference deployment enabling platform, which carries task requirement information such as the inference model, cost, computing resource specifications, network latency, computing resource location, etc. related to the current task, as well as requirement information such as the task deployment mode.

[0142] S12: The inference enabling platform sends a synchronous inference deployment request to the scheduling system.

[0143] In step S13, the scheduling system coordinates the scheduling of computing power, network, and storage resources for this inference deployment task based on the large-scale computing network demand map and the task requirement information contained in the inference deployment request, and generates a corresponding initial scheduling plan. Based on the task deployment mode configured by the administrator user in step S11, the scheduling system can automatically schedule and bind the most cost-effective computing power resources for this task in "easy mode", or send the resource scheduling plan information to the administrator user for confirmation in "safe mode". If the administrator user selects "easy mode", the process jumps to step S18; if the administrator user selects "safe mode", the process continues to step S14.

[0144] S14: The scheduling system feeds back the initial scheduling plan to the inference deployment enabling platform.

[0145] S15: The inference deployment enabling platform presents the initial scheduling plan to the administrator user.

[0146] S16: The administrator user selects and confirms the initial scheduling plan through the inference deployment enabling platform to obtain the target scheduling plan.

[0147] S17: The inference deployment enabling platform returns the user confirmation result to the scheduling system.

[0148] In step S18, the scheduling system binds the resource information in the target scheduling solution using a serverless placeholder. If the model file and image file used by the current task already exist at the computing node in the target scheduling solution, the system proceeds to step S115. If the model file or image file used by the current task is missing at the computing node in the target scheduling solution, the system proceeds to step S19.

[0149] S19, the dispatching system opens the wide area network transmission tunnel.

[0150] At step S110, the scheduling system requests the inference deployment enabling platform to synchronize (copy) the model files and image files required for executing the inference task to the candidate computing nodes in the target scheduling solution. The model files and image files required for executing the inference task must be exported to the domain after the administrator confirms that there is no privacy risk. The image files must be pre-configured with computing resources.

[0151] S111 , the inference deployment enabling platform requests the model management system to synchronize the model file or image file with the node that lacks the model file or image file required to execute the inference task.

[0152] S112 , the model management system synchronizes the model file or the mirror file with the node that lacks the model file or the mirror file required for executing the reasoning task.

[0153] In step S113, the model management system returns the synchronization results of the model files and image files required to execute the inference task to the inference deployment enabling platform. If the veteran instance is not deployed remotely, the process proceeds to step S118; if the veteran instance is deployed remotely, the process continues to step S115.

[0154] S114, the inference deployment enabling platform returns the synchronization result of the model file and the image file required to execute the inference task to the scheduling system.

[0155] In S115, the scheduling system initiates a veteran instance deployment request to the computing power node. The serverless computing resource cluster with available computing power in the entire network only needs to deploy one veteran instance to reduce resource waste when there is no business traffic access and reduce the cold start delay of the user's first access.

[0156] S116, deploy veteran instances on candidate computing nodes.

[0157] S117, the candidate computing power node returns the veteran instance deployment result to the scheduling system.

[0158] S118, the scheduling system synchronizes the inference deployment results to the load balancing system, including computing power node routing information, as well as monitoring information such as resource status and inference engine deployment status.

[0159] S119: The scheduling system returns the inference deployment result to the inference deployment enabling platform.

[0160] S120: The inference deployment enabling platform presents the inference deployment result to the administrator user.

[0161] Figure 5 FIG. 1 is a flowchart of an optional reasoning service phase of a reasoning task according to an embodiment of the present invention. Figure 5 As shown in the figure, the reasoning service process based on the reasoning task specifically includes:

[0162] S21: The terminal user initiates a high-concurrency reasoning service request to the reasoning session platform.

[0163] S22 includes the following two parallel steps:

[0164] (S22-1) The reasoning session platform queries the state library for the reasoning context and proceeds to step (S23-1).

[0165] (S22-2) The reasoning conversation platform initiates a knowledge retrieval request to the user knowledge base based on the user's original prompt and continues with step (S23-2).

[0166] S23 includes the following two parallel steps:

[0167] (S23-1) The state library returns the reasoning context information to the reasoning session platform, and proceeds to step S24.

[0168] (S23-2) The knowledge base returns the knowledge retrieval results to the reasoning conversation platform, and then proceeds to step S24.

[0169] S24, the reasoning session platform initiates a reasoning task execution request carrying context information and knowledge retrieval results to the load balancing system.

[0170] S25: The load balancing system provides global routing information including local computing nodes and remote computing nodes for business traffic based on the configured load balancing strategy.

[0171] S26, the load balancing system synchronizes the routing information (load balancing result) balanced in step S25 to the scheduling system.

[0172] S27, includes the following two parallel steps:

[0173] (S27-1) The scheduling system sends an inference deployment task request to the selected remote computing power node (i.e., the target computing power node), requesting that the inference instance be deployed at the node where the business traffic is routed in the load balancing result, and continues with step (S28-1).

[0174] (S27-2) The scheduling system opens a wide area network transmission tunnel and proceeds to step (S28-2).

[0175] S28, consists of the following two parallel steps:

[0176] (S28-1) Deploy the inference instance on the selected remote computing node, and after step (S29-2) occurs, jump to step (S210-2).

[0177] (S28-2) The scheduling system returns the wide area network transmission tunnel opening result to the load balancing system, and continues with steps (S29-1) and (S29-2).

[0178] S29, consists of the following two parallel steps:

[0179] (S29-1) The load balancing system distributes the reasoning task execution request carrying the context information and knowledge retrieval results to the local computing power node and continues with step (S210-1).

[0180] (S29-2) The load balancing system distributes the reasoning task execution request carrying the context information and knowledge retrieval results to the selected remote computing power node (ie, the target computing power node), and continues with step (S210-2).

[0181] S210 includes the following two parallel steps:

[0182] (S210-1) The local computing power node generates an inference service result and continues with step (S211-1).

[0183] (S210-2) The selected remote computing power node (i.e., the target computing power node) generates the reasoning service result and proceeds to step (S211-2). If the veteran instance has been deployed in the reasoning deployment stage, when the reasoning instance deployed in step (S28-1) has not yet been deployed, and the reasoning task execution request of step (S29-2) has been distributed to the cloud, the veteran instance first generates the reasoning service result, and after the reasoning instance deployed in step (S28-1) is deployed, it provides reasoning services together with the veteran instance to generate remote reasoning results. If the veteran instance has not been deployed in the reasoning deployment stage, when the reasoning instance deployed in step (S28-1) has not yet been deployed, and the reasoning request of step (S29-2) has been distributed to the cloud, wait until the reasoning instance deployed in step (S28-1) is deployed, and then generate the remote reasoning result.

[0184] S211 includes the following two parallel steps:

[0185] (S211-1) The local computing power node returns the inference service result to the load balancing system and proceeds to step S212.

[0186] (S211-2) The remote computing node returns the inference service result to the load balancing system and continues with step S212.

[0187] S212: The load balancing system returns the inference service result to the inference session platform.

[0188] S213, the reasoning conversation platform presents the reasoning service result to the terminal user.

[0189] S214: The reasoning conversation platform updates the context state information in the state library.

[0190] S215 includes the following two parallel steps:

[0191] (S215-1) When the number of concurrent terminal user inference service requests decreases, the load balancing system monitors and senses that the computing resource utilization rate of the remote instance has decreased. When the utilization rate decreases below the resource threshold, step (S216-1) is continued.

[0192] (S215-2) When the number of concurrent terminal user inference service requests increases, the load balancing system monitors and senses the increase in the computing resource occupancy rate of the remote instance. When the occupancy rate increases to above the resource upper limit, step (S216-2) is continued.

[0193] S216 includes the following two parallel steps:

[0194] (S216-1) The load balancing system requests the scheduling system to elastically scale down computing resources, i.e., to release instance resources whose resource utilization is lower than the resource threshold, and distribute the business traffic carried by them to other instances for execution, and continue with step (S217-1).

[0195] (S216-2) The load balancing system requests the scheduling system to elastically expand computing resources, that is, to deploy a new inference instance within the resource range occupied by the serverless system, and then continues with step (S217-2).

[0196] S217 includes the following two parallel steps:

[0197] (S217-1) The computing power scheduling system requests the remote computing power node to elastically shrink, and continues with step (S218-1).

[0198] (S217-2) The computing power scheduling system requests elastic expansion of the remote computing power node and continues with step (S218-2).

[0199] S218 includes the following two parallel steps:

[0200] (S218-1) The remote computing power node performs elastic scaling of computing power resources.

[0201] (S218-2) The remote computing nodes perform elastic expansion of computing resources.

[0202] Based on the above embodiments and optional embodiments, the present invention proposes another optional implementation method of a task execution method based on a large model, which is used to execute fine-tuning tasks based on a large model, and the fine-tuning tasks only involve administrator users. When the user's local computing resources are sufficient to support the administrator user's fine-tuning deployment service quality requirements (mainly latency requirements), fine-tuning deployment is prioritized on low-cost, safe and reliable localized computing nodes. When the concurrency of fine-tuning tasks exceeds the upper limit of the local computing resources, or the local computing resources are insufficient to support the fine-tuning task requirements, the fine-tuning deployment tasks will face problems such as queuing and inability to fine-tune, affecting the user experience. At this time, the administrator user can initiate a cloud-edge collaborative fine-tuning deployment task, deploy the fine-tuning task on a remote computing node, and then execute the fine-tuning task on the node to shorten the task queuing time and expand the local business scope, that is, the "network-following computing" mode. Specifically, Figure 6 is a flow chart of an optional fine-tuning task execution according to an embodiment of the present invention, such as Figure 6 As shown in the figure, the specific process of fine-tuning task execution is as follows:

[0203] S31: The administrator user initiates a fine-tuning task deployment request on the fine-tuning deployment enabling platform.

[0204] S32, the fine-tuning deployment enabling platform requests the computing network scheduling system to fine-tune the remote deployment.

[0205] In step S33, the scheduling system coordinates the scheduling of computing power, network, and storage resources for this deployment task based on the large-scale computing network demand map and the user requirements contained in the request, and generates a corresponding initial scheduling plan. Based on the task deployment mode configured by the administrator user in step S31, the scheduling system can automatically schedule and bind the most cost-effective computing power resources for this task in "Easy Mode", or send the resource plan information to the administrator user for confirmation in "Rest assured mode". If the administrator user selects "Easy Mode", the process jumps to step S38; if the administrator user selects "Rest assured mode", the process continues to step S34.

[0206] S34, the scheduling system feeds back the initial scheduling plan to the fine-tuning deployment enabling platform.

[0207] S35, the fine-tuning deployment enabling platform presents the initial scheduling plan to the administrator user.

[0208] S36: The administrator user selects and confirms the initial scheduling plan by fine-tuning the deployment enabling platform to obtain the target scheduling plan.

[0209] S37, the fine-tuning deployment enabling platform returns user confirmation information to the scheduling system.

[0210] S38, the scheduling system binds the resource information in the target scheduling solution.

[0211] S39, the dispatching system opens a wide area network transmission tunnel.

[0212] S310: The scheduling system initiates a file preparation request to the fine-tuning deployment enabling platform, requesting that the enabling platform synchronize (copy) the fine-tuning source model file, fine-tuning image file, and user dataset used for fine-tuning to the computing nodes in the target scheduling solution. Files must be transferred out of the domain after the administrator confirms that there is no privacy risk. Image files must be pre-configured for computing resources.

[0213] S311 includes the following two parallel steps:

[0214] (S311-1) The fine-tuning deployment enabling platform requests the model management system to synchronize the model file or fine-tuning image file with the remote computing node (i.e., the target computing node) that is missing the model file or fine-tuning image file in the target scheduling plan, and then proceed to step (S312-1). If the remote computing node already has the model file and image file for this fine-tuning, this step is skipped.

[0215] (S311-2) The fine-tuning deployment enabling platform requests the user data center to synchronize the user local data set used for this fine-tuning to the remote computing power node, and continues with step (S312-2).

[0216] S312 includes the following two parallel steps:

[0217] (S312-1) The model management system synchronizes the model file or the fine-tuning image file to the remote computing node where the model file or the fine-tuning image file is missing, and continues with step (S313-1).

[0218] (S312-2) The user data center synchronizes the user local data set used for this fine-tuning to the remote computing power node and continues with step (S313-2).

[0219] S313 includes the following two parallel steps:

[0220] (S313-1) The model management system returns the synchronization result of the model file or the fine-tuning image file to the fine-tuning deployment enabling platform, and then proceeds to step S314.

[0221] (S313-2) The user data center returns the data set synchronization result to the fine-tuning deployment enabling platform, and proceeds to step S314.

[0222] S314, the fine-tuning deployment enabling platform returns the file preparation result to the scheduling system.

[0223] S315, the scheduling system initiates a fine-tuning task execution request to the remote computing power node.

[0224] S316: The remote computing node performs fine-tuning task deployment and generates a fine-tuned model file.

[0225] S317: The fine-tuned model file is transmitted back from the remote computing node to the user model management system.

[0226] In step S318, the resources related to this fine-tuning are released from the remote computing node, and the user dataset, model file, and fine-tuning image file uploaded to the remote computing node for this task are cleaned up.

[0227] S319, the remote computing node returns the fine-tuning deployment result to the scheduling system.

[0228] S320: The scheduling system returns the fine-tuning deployment result to the fine-tuning deployment enabling platform.

[0229] S321, the fine-tuning deployment enabling platform presents the fine-tuning deployment result to the administrator user.

[0230] It should be noted that in this embodiment, by integrating cloud computing power, idle all-in-one machines and other network-wide computing power resources, the scale and capability range of AI computing power resources are expanded, breaking the business models of pure all-in-one machines, self-built and self-sold cloud computing power, and connecting to specific third-party computing power. Build and evaluate the network demand map for large-scale model computing power, and according to the business needs of large-scale models, conduct cloud-edge computing and network collaborative scheduling for all-network remote, heterogeneous, and heterogeneous resources, match the cloud computing power resources with the best cost-effectiveness, and provide all-in-one machine users with large-scale model training and promotion services with traffic distribution modes such as network-following computing and global load balancing computing-following network. This embodiment is based on the deterministic connection capability of the wide area network, breaks through the bottlenecks of network transmission security, reliability, and speed, and realizes the deterministic flow of large model files and data sets at a rate of up to 100GiB. This embodiment is conducive to achieving the connection between the supply and demand sides of the large-model computing network, and steplessly directing user services to the remote computing power supply side for acceleration, expansion and business development (cloud-edge collaboration mode), and activating idle computing power resources including remote training and push all-in-one machines (edge sharing mode), enabling demanders to obtain the best computing power, and enabling suppliers to open up distribution channels and provide precise services.

[0231] According to an embodiment of the present invention, a system embodiment of task execution based on a large model is provided. Figure 7 is a structural diagram of a task execution system based on a large model according to an embodiment of the present invention. Figure 7 As shown, the above-mentioned task execution system based on the large model includes: a target interaction device 70, a data processing device 72, and a target computing node 74, wherein the target computing node 74 is a server device deployed in the cloud, wherein:

[0232] The target interaction device 70 is configured to, upon receiving a service request for a target task, generate a task execution request based on the service request and send the task execution request to the data processing device, wherein the target task is an inference task based on a large model and the service request is an inference service request for the inference task; or the target task is a fine-tuning task based on a large model and the service request is a fine-tuning deployment request for the fine-tuning task;

[0233] The data processing device 72 is used to send the task execution request to the target computing power node;

[0234] The target computing power node 74 is used to execute the target task based on the task execution request.

[0235] Optionally, the target computing power node can be a large model training and inference all-in-one machine. The target interactive device can receive a service request sent by a user end (such as a terminal device user or an administrator user), and generate a task execution request based on the service request, and send the task execution request to a data processing device (such as a scheduling system or a load balancing system); the data processing device sends the task execution request to the target computing power node; the target computing power node executes the target task based on the task execution request, and sends the obtained task execution result (such as the inference service result or the fine-tuning execution result, etc.) to the data processing device; the data processing device returns the task execution result to the user end via the target interactive device.

[0236] It should be noted that, when the business type of the target business is different, there are differences in the corresponding user end, task execution request, target interactive device, data processing device and the obtained task execution result. For example, when the target task is an inference task, the corresponding user end is a terminal device, the task execution request is an inference task execution request, the target interactive device is deployed with an inference deployment enabling platform and an inference session platform, the data processing device is deployed with a scheduling system and a load balancing system, and the obtained task execution result is an inference service result. When the target business is a fine-tuning business, the corresponding user end is an administrator user, the task execution request is a fine-tuning task execution request, the target interactive device is deployed with a fine-tuning deployment enabling platform, the data processing device is deployed with a scheduling system, and the obtained task execution result is a fine-tuning execution result. The method of selecting the target computing power node, the method of generating the task execution request, etc. are the same as those in the aforementioned embodiment and will not be repeated here.

[0237] By collaboratively executing target tasks based on large models through target interaction devices, data processing devices, and target computing nodes, the purpose of intelligent scheduling can be achieved by automatically identifying and selecting remote (target) computing nodes to execute tasks, thereby achieving the technical effect of improving the efficiency, flexibility, and stability of task execution based on large models, and thus solving the technical problems of low task execution efficiency and poor flexibility in task execution methods based on large models in related technologies.

[0238] In an optional embodiment, when the task type of the target task is different, the target interactive device is set up differently. For example, when the target task is an inference task, an inference deployment enabling platform and an inference conversation platform are deployed on the target interactive device, wherein the inference deployment enabling platform is used to interact with the data processing device and the target device for the inference deployment request of the inference task during the inference deployment phase of the inference task, wherein the inference deployment enabling platform can be locally deployed; or it can be deployed locally and deployed in the cloud at the same time; the inference conversation platform is used to interact with the data processing device and the terminal device for the inference service request of the inference task during the inference service phase of the inference task; for another example, when the target task is a fine-tuning task, the target interactive device is deployed with a fine-tuning deployment enabling platform, wherein the fine-tuning deployment enabling platform is used to interact with the data processing device and the target device for the fine-tuning task. Among them, the operations performed by the inference deployment enabling platform, the inference conversation platform, and the fine-tuning deployment enabling platform during the task execution process and the interactions with other devices in the system are the same as in other embodiments and will not be repeated here.

[0239] In an optional embodiment, when the target task is a reasoning task, the system further includes a state library and a knowledge library. The role of the state library and the knowledge library in executing the reasoning task, as well as the interaction between the state library and the knowledge library and other devices in the system, are the same as those in the previous embodiment and will not be repeated here.

[0240] In an optional embodiment, the task execution system of the large model can have multiple hardware deployment forms. For example, the system can be set to also include local computing power nodes and CPU servers, wherein the local computing power nodes are server devices deployed on the local side, and the local computing power nodes are deployed with model files and image files required to execute the target tasks, and the target interactive device is deployed on the CPU server. Figure 8 is a schematic diagram of a first system structure for performing a large model reasoning task according to an embodiment of the present invention. Figure 8 As shown in the figure, in this system, the inference deployment enablement platform (including local deployment and cloud-based value-added deployment) and the inference session platform are deployed on CPU servers. The large-model training and inference all-in-one machine serves as a local computing node, containing model files and image files (as well as inference models and inference engine images) for executing large-model inference tasks, and deploying corresponding inference instances. The knowledge base, state library, and model management system are deployed independently. Figure 9 is a schematic diagram of a first system structure for performing a fine-tuning task of a large model according to an embodiment of the present invention, such as Figure 9As shown in the figure, in this system, the fine-tuning deployment enablement platform (including local deployment and cloud-based value-added deployment) is deployed on the CPU server. The large model training and induction machine serves as a local computing node and contains the source model file and fine-tuning image file of the large model before fine-tuning. The user data center and model management system are deployed independently.

[0241] In an optional embodiment, the system may also be configured to include a CPU server, wherein the target interactive device is deployed on the CPU server. Figure 10 is a schematic diagram of a second system structure for performing reasoning tasks of a large model according to an embodiment of the present invention. Figure 10 As shown in the figure, in this system, the inference deployment enablement platform (cloud-based value-added deployment) and the inference session platform are deployed on CPU servers. Users lack local computing power and can only use remote computing power for model deployment and inference services. The knowledge base, state library, and model management system are deployed independently. Figure 11 is a schematic diagram of a second system structure for performing a fine-tuning task of a large model according to an embodiment of the present invention, such as Figure 11 As shown in the figure, in this system, the fine-tuning deployment enabling platform (cloud-based value-added deployment) is deployed on the CPU server. Users lack local computing power and can only use remote computing power for fine-tuning. The user data center and model management system are deployed independently.

[0242] In an optional embodiment, the system can also be configured to include a local computing node, wherein the local computing node is deployed with the model file and the image file required to execute the target task, and the target interactive device is deployed on the local computing node, for example, Figure 12 is a schematic diagram of a third system structure for performing reasoning tasks of a large model according to an embodiment of the present invention, such as Figure 12 As shown in the figure, in this system, the inference deployment enablement platform (including local deployment and cloud-based value-added deployment), inference session platform, inference model, and inference engine image are all deployed on the large model training and inference all-in-one machine. The knowledge base, state library, and model management system are deployed independently. Figure 13 is a schematic diagram of a third system structure for performing a fine-tuning task of a large model according to an embodiment of the present invention, such as Figure 13 As shown in the figure, in this system, the fine-tuning deployment enabling platform (including local deployment and cloud-based value-added deployment), pre-fine-tuning source model files, and fine-tuning image files are all deployed on the large model training and push all-in-one machine. The user data center and model management system are deployed independently.

[0243] In an optional embodiment, the system can also be configured to include a local computing node and a CPU server, where the model files and image files required to execute the target task are deployed on the local computing node, and the target interactive device is deployed on the local computing node and the CPU server. Figure 14 is a schematic diagram of a fourth system structure for performing reasoning tasks of a large model according to an embodiment of the present invention, such as Figure 14 As shown in the figure, in this system, the inference deployment enablement platform (local deployment), inference session platform, inference model, and inference engine image are all deployed on a large model training and inference machine. The inference deployment enablement platform (cloud-based value-added deployment) is deployed on a separate CPU server. When administrators need to perform remote inference deployment, they use the independent cloud-based value-added deployment interface to issue tasks. The knowledge base, state library, and model management system are deployed independently. Figure 15 is a schematic diagram of a fourth system structure for performing a fine-tuning task of a large model according to an embodiment of the present invention, such as Figure 15 As shown in the figure, in this system, the fine-tuning deployment enablement platform (local deployment), pre-fine-tuning source model files, and fine-tuning image files are all deployed on the large model training and push all-in-one machine. The fine-tuning deployment enablement platform (cloud-based value-added deployment) is deployed on a separate CPU server. When administrators need to perform remote fine-tuning deployment, they use the independent cloud-based value-added deployment interface to issue tasks. The user data center and model management system are deployed independently.

[0244] It should be noted that the specific structure of the large model-based task execution system shown in this application is only a schematic illustration. In specific applications, the large model-based task execution system in this application may have more or less structures than those recorded in this embodiment.

[0245] It should be noted that any optional or preferred large model-based task execution method in the above method embodiments can be executed or implemented in the large model-based task execution provided in this embodiment.

[0246] In addition, it should be noted that the optional or preferred implementation of this embodiment can be found in the relevant description in the method embodiment, which will not be repeated here.

[0247] In this embodiment, a task execution device based on a large model is also provided. The device is used to implement the above-mentioned embodiments and preferred implementation methods. The details that have been described will not be repeated here. As used below, the terms "module" and "device" can be a combination of software and / or hardware that implements the predetermined functions. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0248] According to an embodiment of the present invention, there is also provided an embodiment of a device for implementing the above-mentioned large model-based task execution method. Figure 16 is a structural diagram of a task execution device based on a large model according to an embodiment of the present invention. Figure 16As shown, the above-mentioned task execution device based on the large model includes: a receiving module 1600, an acquisition module 1602, and a sending module 1604, wherein:

[0249] A receiving module 1600 is configured to receive a task execution request for a target task, wherein the target task is an inference task based on a large model or a fine-tuning task based on a large model;

[0250] An acquisition module 1602, connected to the receiving module 1600, is configured to acquire a target computing node for executing a target task according to the target task, wherein the target computing node is a server device deployed in the cloud;

[0251] The sending module 1604 is connected to the obtaining module 1602 and is used to send the task execution request to the target computing power node, so that the target computing power node executes the target task based on the task execution request.

[0252] It should be noted that the above modules can be implemented by software or hardware. For example, for the latter, it can be implemented in the following ways: the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0253] It should be noted that the receiving module 1600, obtaining module 1602, and sending module 1604 correspond to steps S102 to S106 in the embodiment. The examples and application scenarios implemented by the modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above embodiment. It should be noted that the modules, as part of the device, can be run on a computer terminal.

[0254] It should be noted that the optional or preferred implementation of this embodiment can be found in the relevant description in the embodiment, which will not be repeated here.

[0255] The above-mentioned task execution device based on the large model can also include a processor and a memory. The above-mentioned receiving module 1600, acquisition module 1602, sending module 1604, etc. are all stored in the memory as program modules, and the processor executes the above-mentioned program modules stored in the memory to realize the corresponding functions.

[0256] The processor includes a core, which retrieves corresponding program modules from memory. There can be one or more cores. Memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory includes at least one memory chip.

[0257] According to an embodiment of the present application, an embodiment of a non-volatile storage medium is also provided. Optionally, in this embodiment, the non-volatile storage medium includes a stored program, wherein when the program is executed, the device containing the non-volatile storage medium is controlled to execute any of the above-mentioned large model-based task execution methods.

[0258] Optionally, in this embodiment, the non-volatile storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group, and the non-volatile storage medium includes a stored program.

[0259] According to an embodiment of the present application, an embodiment of a processor is further provided. Optionally, in this embodiment, the processor is used to run a program, wherein when the program is run, any of the above-mentioned large model-based task execution methods is executed.

[0260] According to an embodiment of the present application, an embodiment of a computer program product is also provided, which, when executed on a data processing device, is suitable for executing a program that initializes any of the steps of the above-mentioned large model-based task execution method.

[0261] An embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a program stored in the memory and runnable on the processor. When the processor executes the program, the program implements any one of the above-mentioned steps of the large model-based task execution method.

[0262] The above sequence of the embodiments of the present invention is for description only and does not represent the superiority or inferiority of the embodiments.

[0263] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0264] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the above modules can be a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, modules or indirect coupling or communication connection of modules, which can be electrical or other forms.

[0265] The modules described above as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.

[0266] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or software functional modules.

[0267] If the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable non-volatile storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a non-volatile storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned non-volatile storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), block storage (Block Storage), mobile hard disk, magnetic disk or optical disk and other media that can store program code.

[0268] The above are only preferred embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A task execution method based on a large model, characterized in that: include: Receiving a task execution request for a target task, wherein the target task is an inference task based on a large model, or a fine-tuning task based on the large model; Obtaining a target computing node for executing the target task according to the target task, wherein the target computing node is a server device deployed in the cloud; The task execution request is sent to the target computing power node, so that the target computing power node executes the target task based on the task execution request.

2. The method according to claim 1, characterized in that The step of obtaining a target computing power node for executing the target task according to the target task includes: Obtaining candidate computing nodes pre-determined based on a target scheduling scheme, wherein the target scheduling scheme is obtained based on computing resource information and task requirement information matched with the target task; Based on the candidate computing power nodes, the target computing power node is obtained.

3. The method according to claim 2, characterized in that The obtaining of candidate computing nodes pre-determined based on the target scheduling scheme includes: receiving a task deployment request for the target task, wherein the task deployment request is an inference deployment request based on the inference task, or a fine-tuning deployment request based on the fine-tuning task; Determining the target scheduling solution based on the computing resource information and the task requirement information carried in the task deployment request; Based on the target scheduling scheme, the candidate computing power nodes are determined.

4. The method according to claim 3, characterized in that In a case where the task requirement information includes a task deployment mode, determining the target scheduling scheme according to the computing resource information carried in the task deployment request and the task requirement information includes: Determine an initial scheduling plan based on the computing resource information carried in the task deployment request, wherein the initial scheduling plan includes initially determined information on computing nodes to be scheduled for the target task; Obtaining the task deployment mode of the target task; determining the target scheduling scheme based on the initial scheduling scheme when the task deployment mode is a first mode, wherein the first mode is used to indicate that there is no confirmation process for the scheduling scheme; or When the task deployment mode is the second mode, sending a solution confirmation instruction to the target device, wherein the second mode is used to indicate that there is a confirmation process of the scheduling solution, and the solution confirmation instruction is used to indicate whether the scheduling solution is passed; The target scheduling solution is determined based on the solution confirmation result returned by the target device.

5. The method according to claim 2, characterized in that The method further comprises: When the target task is the inference task and there are multiple candidate computing nodes, multiple candidate computing nodes are bound in a serverless placeholder manner, and the model files and image files required to execute the target task are synchronized to the multiple candidate computing nodes; or In the case where the target task is the fine-tuning task, a file preparation request is initiated to the target interactive device, so that the target interactive device performs a file synchronization operation based on the file preparation request to obtain a file preparation result, wherein the file preparation result is used to be sent to the target computing power node, and the file synchronization operation is used to synchronize the source model file of the large model, the fine-tuning image file, and the data set required to execute the fine-tuning task to the target computing power node, and the file preparation result is used to indicate the execution result of the file synchronization operation.

6. The method according to claim 5, characterized in that After binding multiple candidate computing nodes in a serverless placeholder manner, the method further includes: Deploy an initial inference instance on one of the multiple candidate computing nodes.

7. The method according to claim 2, characterized in that The obtaining the target computing power node based on the candidate computing power node includes: When the target task is the inference task and there are multiple candidate computing nodes, the target computing node is obtained by: determining the target computing node from the multiple candidate computing nodes using a load balancing method; or In the case where the target task is the fine-tuning task, the target computing power node is obtained in the following manner: the candidate computing power node is used as the target computing power node.

8. The method according to claim 1, characterized in that In the case where the target task is the reasoning task, the task execution request is a reasoning task execution request, wherein the reasoning task execution request is used to request execution of the reasoning task to obtain a reasoning service result, and the reasoning task execution request carries a task execution request with reasoning context information and a knowledge retrieval result, the reasoning context information is based on the reasoning service request for the target task initiated by the terminal device, and is obtained by querying the state library; the knowledge retrieval result is based on the reasoning service request, and is obtained by querying the knowledge library; the reasoning service request carries request information input through the terminal device; or In the case where the target task is the fine-tuning task, the task execution request is a fine-tuning task execution request, wherein the fine-tuning task execution request is used to request execution of the fine-tuning task, obtain a fine-tuning execution result, and a fine-tuned model file, and the fine-tuning execution result is at least used to indicate whether the fine-tuning task is executed successfully.

9. The method according to claim 1, characterized in that The step of obtaining a target computing power node for executing the target task according to the target task includes: In the case where the target computing power resource amount exceeds the computing power resource amount of the local computing power node, obtaining the target computing power node for executing the target task according to the target task, wherein the target computing power resource amount is the computing power resource amount required to execute the target task, and the local computing power node is a server device deployed on the local side; or When the computing resource occupancy rate of the local computing node is greater than a preset threshold, obtaining the target computing node for executing the target task according to the target task; or When the waiting time required for executing the target task on the local computing power node is longer than a predetermined time, obtaining the target computing power node for executing the target task according to the target task; or In the case that the local computing power node that executes the target task does not exist, the target computing power node that executes the target task is obtained according to the target task.

10. The method according to claim 1, characterized in that The obtaining of the target computing node for executing the target task includes: determining the target computing node from a plurality of candidate computing nodes using a load balancing method, and determining task execution requests corresponding to the target computing node and a local computing node, respectively, wherein the local computing node is a server device deployed on the local side; The sending of the task execution request to the target computing power node, so that the target computing power node executes the target task based on the task execution request, includes: sending the corresponding task execution request to the target computing power node and the local computing power node, so that the target computing power node and the local computing power node respectively execute the target task based on the corresponding task execution request.

11. The method according to claim 1, wherein The step of sending the task execution request to the target computing power node, so that the target computing power node executes the target task based on the task execution request, includes: In a case where the target task is the inference task and the corresponding task execution request is an inference task execution request, after distributing the task execution request to the target computing power node through a load balancing method, an inference task instance corresponding to the inference task is deployed on the target computing power node for the target computing power node to execute the inference task; or When the target task is the fine-tuning task and the corresponding task execution request is a fine-tuning task execution request, the fine-tuning instance of the fine-tuning task is deployed to the target computing power node, and the target computing power node is used to execute the fine-tuning task, and the fine-tuning execution result and the fine-tuning model file are obtained, and the fine-tuning execution result is returned to the target device, and the fine-tuning model file is returned to the model management device, wherein the fine-tuning execution result is at least used to indicate whether the fine-tuning task is executed successfully.

12. The method according to claim 11, characterized in that In a case where the target computing node is selected from a plurality of candidate computing nodes, and one of the plurality of candidate computing nodes deploys an initial inference instance, the method further includes: Detecting whether the deployment of the inference task instance on the target computing power node is completed; In the case where the deployment of the inference task instance on the target computing power node is not completed, the task execution request is distributed to the computing power node where the initial inference instance is deployed, so that the computing power node where the initial inference instance is deployed executes the inference task; When the deployment of the inference task instance on the target computing power node is completed, the target task is jointly executed by the computing power node on which the initial inference instance is deployed and the target computing power node.

13. The method according to any one of claims 1 to 12, characterized in that In a case where the target task is the inference task, after sending the task execution request to the target computing power node for the target computing power node to execute the target task based on the task execution request, the method further includes: Obtaining the computing power resource occupancy rate of the target computing power node; When the computing power resource occupancy rate is lower than a preset lower limit, controlling the target computing power node to release part of the computing power resources used to execute the target task; or When the computing power resource occupancy rate is higher than a preset upper limit value, a first computing power node is determined from multiple candidate computing power nodes, and an inference task instance corresponding to the inference task is deployed on the first computing power node, wherein the first computing power node is used to execute the target task together with the target computing power node.

14. The method according to any one of claims 13, characterized in that The obtaining of the computing resource occupancy rate of the target computing node includes: Monitoring the concurrency of the task execution requests received by the target computing power node; The computing power resource occupancy rate is determined according to the concurrency amount.

15. A task execution method based on a large model, characterized in that: include: receiving a task execution request for a target task, wherein the target task is an inference task based on a large model or a fine-tuning task based on the large model, and the task execution request is sent by a data processing device; The target task is executed based on the task execution request.

16. The method according to claim 15, characterized in that The executing the target task based on the task execution request includes: In the case where the target task is the reasoning task and the task execution request is an reasoning task execution request, the reasoning task is executed based on the reasoning task execution request to obtain an reasoning service result, wherein the reasoning task execution request carries a task execution request with reasoning context information and a knowledge retrieval result, the reasoning context information is based on the reasoning service request for the target task initiated by the terminal device and is obtained by querying the state library; the knowledge retrieval result is based on the reasoning service request and is obtained by querying the knowledge library; the reasoning service request carries request information input through the terminal device; the reasoning service result is sent to the data processing device, so that the data processing device sends the reasoning service result to the terminal device via the target interactive device.

17. The method according to claim 15, characterized in that The executing the target task based on the task execution request includes: In the case where the target task is the fine-tuning task and the task execution request is a fine-tuning task execution request, the fine-tuning task is executed based on the task execution request to obtain a fine-tuning execution result and a fine-tuning model file, wherein the fine-tuning execution result is at least used to indicate whether the fine-tuning task is successfully executed; the fine-tuning execution result is sent to the data processing device, so that the data processing device sends the fine-tuning execution result to the target device via the target interaction device, and sends the fine-tuning model file to the model management device.

18. A task execution method based on a large model, characterized in that: include: Receive a service request for a target task, wherein the target task is an inference task based on a large model, and the service request is an inference service request for the inference task; or the target task is a fine-tuning task based on the large model, and the service request is a fine-tuning deployment request for the fine-tuning task; generating a task execution request based on the service request; The task execution request is sent to a data processing device, so that the data processing device sends the task execution request to a target computing power node, and the target task is executed by the target computing power node.

19. The method according to claim 18, characterized in that Generating a task execution request based on the service request includes: When the target task is the reasoning task, the service request is an reasoning service request for the reasoning task, and the task execution request is an reasoning task execution request, the reasoning service request is sent to the state library and the knowledge library respectively; the reasoning context information returned by the state library based on the reasoning service request and the knowledge detection result returned by the knowledge library are received; and the reasoning task execution request is generated based on the reasoning context information and the knowledge detection result.

20. The method according to claim 18, wherein Generating a task execution request based on the service request includes: When the target task is the fine-tuning task, the service request is a fine-tuning deployment request for the fine-tuning task, and the task execution request is a fine-tuning task execution request, a file preparation request sent by the data processing device is received, wherein the file preparation request is used to request the execution of a file synchronization operation, and the file synchronization operation is used to synchronize the source model file of the large model, the fine-tuning image file, and the data set required to execute the fine-tuning task to the target computing power node; based on the execution result of the file synchronization operation, a file preparation result is obtained, and the file preparation result is sent to the data processing device, so that the data processing device generates the fine-tuning task execution request based on the file preparation result.

21. The method according to claim 18, wherein After sending the task execution request to the data processing device, for the data processing device to send the task execution request to the target computing node if the local computing power node does not meet the preset execution conditions of the target task, for the target computing node to execute the target task, the method further includes: In a case where the target task is the inference task and the task execution request is an inference task execution request, receiving an inference service result obtained by the target computing power node executing the inference task, and sending the inference service result to the terminal device; or In the case where the target task is the fine-tuning task and the task execution request is a fine-tuning task execution request, a fine-tuning execution result obtained by the target computing power node executing the fine-tuning task is received; the fine-tuning execution result is sent to the target device, wherein the fine-tuning execution result is at least used to indicate whether the fine-tuning task is executed successfully.

22. A task execution system based on a large model, characterized in that: It includes a target interactive device, a data processing device, and a target computing node, wherein the target computing node is a server device deployed in the cloud, wherein: The target interaction device is configured to, upon receiving a service request for a target task, generate a task execution request based on the service request, and send the task execution request to the data processing device, wherein the target task is an inference task based on a large model, and the service request is an inference service request for the inference task; or the target task is a fine-tuning task based on the large model, and the service request is a fine-tuning deployment request for the fine-tuning task; The data processing device is used to send the task execution request to the target computing power node; The target computing power node is used to execute the target task based on the task execution request.

23. The system according to claim 22, wherein: The system further includes a local computing node and a CPU server, wherein the local computing node is a server device deployed on the local side, the model file and the image file required to execute the target task are deployed on the local computing node, and the target interaction device is deployed on the CPU server; or The system further includes a CPU server, wherein the target interactive device is deployed on the CPU server; or The system further includes a local computing node, wherein the model file and the image file required for executing the target task are deployed on the local computing node, and the target interactive device is deployed on the local computing node; or The system also includes a local computing node and a CPU server. The model file and the image file required to execute the target task are deployed on the local computing node, and the target interactive device is deployed on the local computing node and the CPU server.

24. The system according to claim 22, wherein: When the target task is the reasoning task, the system further includes a state library and a knowledge library.

25. The system according to claim 22, wherein: In the case where the target task is the reasoning task, an inference deployment enabling platform and an inference conversation platform are deployed on the target interactive device, wherein the inference deployment enabling platform is used to interact with the data processing device and the target device in response to the inference deployment request of the inference task during the inference deployment phase of the inference task; and the inference conversation platform is used to interact with the data processing device and the terminal device in response to the inference service request of the inference task during the inference service phase of the inference task; or In the case that the target task is the fine-tuning task, the target interactive device is deployed with a fine-tuning deployment enabling platform, wherein the fine-tuning deployment enabling platform is used to interact with the data processing device and the target device for the fine-tuning task.

26. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a plurality of instructions, which are suitable for being loaded by a processor and executing the large model task execution method described in any one of claims 1 to 21.

27. An electronic device, characterized in that: It includes one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the task execution method of the large model described in any one of claims 1 to 21.

28. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the task execution method of the large model according to any one of claims 1 to 21 are implemented.

Citation Information

Cited By

  • Data transmission method and device, electronic equipment and storage medium

    CN120935164A