Model deployment method and device based on keywords, equipment and storage medium

Through the keyword-based model deployment method, large-model services are automatically generated and configured, which solves the problems of cumbersome deployment of large-model services and difficult resource allocation in the existing technology, and simplified deployment and automated resource management are realized.

CN120104141AActive Publication Date: 2025-06-06BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510205292.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-06
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

In the prior art, the deployment process of large-model services is cumbersome and resource allocation is difficult, so users need to manually deploy multiple model components and resource allocation.

Method used

Provide a keyword-based model deployment method, which automatically generates the target processing model and configures resources by obtaining model structure keywords and model resource keywords, including dynamically adjusting the resource occupation of model components.

Benefits of technology

It simplifies the deployment process of large-model services, reduces the difficulty of resource allocation, and realizes automated model deployment and resource management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104141A_ABST
    Figure CN120104141A_ABST
Patent Text Reader

Abstract

The invention discloses a keyword-based model deployment method, apparatus and device, and a storage medium. The method comprises the steps of obtaining a model structure keyword and a model resource keyword; generating a target processing model based on the model structure keyword; the target processing model comprises a plurality of model components; and carrying out resource configuration on the plurality of model components in the target processing model based on the model resource keywords. By adopting the method provided by the invention, the deployment process of the large model service can be simplified, and the resource configuration difficulty in the large model deployment process can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a keyword-based model deployment method, device, equipment and storage medium. Background Art

[0002] A large model refers to a machine learning model with large-scale parameters and complex computing structures. In order for users to easily integrate the powerful capabilities of large models into applications, developers need to deploy the model into a large model service in the form of an interface. Therefore, large model service deployment has become an indispensable part of the entire production cycle of large models. However, in the prior art, users usually manually deploy multiple model components into a large model service, and the large model service deployment process is cumbersome. Summary of the invention

[0003] The embodiments of the present application provide a keyword-based model deployment method, apparatus, device and storage medium, which can simplify the deployment process of large model services and reduce the difficulty of resource configuration during the large model deployment process.

[0004] The technical solution adopted by the present invention to solve the problem is as follows: In a first aspect, the present application provides a keyword-based model deployment method, comprising: Get model structure keywords and model resource keywords; Based on the model structure keywords, a target processing model is generated; the target processing model includes multiple model components; Based on the model resource keywords, resources are configured for multiple model components in the target processing model.

[0005] In some embodiments of the present application, after resource configuration is performed on multiple model components in the target processing model based on the model resource keyword, the following steps are included: When the target processing model is running, the running status information of multiple model components is obtained; The resources occupied by multiple model components are dynamically adjusted based on the running status information.

[0006] In some embodiments of the present application, a target processing model is generated based on the model structure keyword, including: Verify multiple model keywords in the model structure keyword and the dependency relationship between multiple model keywords; If the multiple model keywords and the dependency relationships between the multiple model keywords are verified, a target processing model is generated based on the multiple model keywords and the dependency relationships between the multiple model keywords; If the verification of the plurality of model keywords and the dependency relationship between the plurality of model keywords fails, outputting first verification result information; The adjusted model structure keyword input by the user based on the first verification result information is obtained, and a target processing model is generated based on the adjusted model structure keyword.

[0007] In some embodiments of the present application, multiple model keywords in the model structure keyword and the dependency relationship between the multiple model keywords are verified, including: Determine a corresponding model structure based on multiple model keywords in the model structure keyword and dependencies between the multiple model keywords; If the model structure includes a preset model structure, it is determined that the dependency check between multiple model keywords and multiple model keywords fails; If the model structure does not include a preset model structure, it is determined that a plurality of model keywords and dependency relationships between the plurality of model keywords are verified.

[0008] In some embodiments of the present application, resource configuration is performed on multiple model components in the target processing model based on the model resource keywords, including: Verify model resource keywords; If the model resource keyword verification passes, determine the required resource type and required resource amount corresponding to each model component from the model resource keyword, and allocate computing resources in the computing cluster to multiple model components in the target processing model based on the required resource type and required resource amount; If the model resource keyword verification fails, output the second verification result information; The adjusted model resource keyword input by the user based on the second verification result information is obtained, and resource configuration is performed on multiple model components in the target processing model based on the adjusted model resource keyword.

[0009] In some implementation schemes of the present application, the model resource keyword includes the required resource type and required resource amount corresponding to each model component, and the model resource keyword is verified, including: Match the required resource type corresponding to each model component with the resource type in the computing cluster; If the required resource type does not match the resource type in the computing cluster, it is determined that the model resource keyword check fails; If the required resource type matches the resource type in the computing cluster, determine the total required resource amount of each required resource type based on the required resource type and required resource amount corresponding to each model component; If the total amount of required resources of any required resource type is greater than the total amount of resources of the required resource type in the computing cluster, it is determined that the model resource keyword verification fails.

[0010] In some embodiments of the present application, the resources occupied by multiple model components are dynamically adjusted based on the running status information, including: Based on the running status information, determining whether the multiple model components include a first model component; the first model component is a model component that is running and the corresponding task is in a queued state among the multiple model components; If the plurality of model components include the first model component, determining whether the plurality of model components include the second model component; the second model component is a model component among the plurality of model components that has not been run and occupies resources; If the multiple model components include a second model component, the resources of the second model component are dynamically adjusted based on the amount of resources required by the tasks in the queue state in the first model component.

[0011] In some embodiments of the present application, dynamically adjusting the resources of the second model component based on the amount of resources required by the tasks in the queued state in the first model component includes: If the amount of resources required by the tasks in the queued state in the first model component is greater than or equal to the total amount of resources occupied by the second model component, the resources occupied by the second model component are allocated to the first model component; If the amount of resources required by the tasks in the queued state in the first model component is less than the total amount of resources occupied by the second model component, determining the priority of the second model component based on the dependency relationship of the second model component; Determine a third model component from the second model component based on the priority and the amount of resources required by the tasks in the queued state in the first model component; the third model component is a model component with a higher priority in the second model component, and the total amount of resources occupied by the third model component is greater than or equal to the amount of resources required by the tasks in the queued state in the first model component; The resources occupied by the third model component are allocated to the first model component.

[0012] In a second aspect, the present application provides a keyword-based model deployment device, comprising: Information acquisition module, used to obtain model structure keywords and model resource keywords; A model generation module is used to generate a target processing model based on a model structure keyword; the target processing model includes multiple model components; The resource configuration module is used to perform resource configuration on multiple model components in the target processing model based on model resource keywords.

[0013] In a third aspect, the present application further provides a computer device, the computer device comprising: one or more processors; Memory; and One or more applications, wherein the one or more applications are stored in a memory and configured to be executed by a processor to implement any keyword-based model deployment method of the first aspect.

[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and the computer program is loaded by a processor to execute the steps in the keyword-based model deployment method of any one of the first aspects.

[0015] The beneficial effects of the present invention are as follows: by obtaining model structure keywords and generating a target processing model based on the model structure keywords, a computer device can automatically build a large model service based on the model structure keywords, without the need for users to manually build the large model service, thereby simplifying the deployment process of the large model service; by obtaining model resource keywords, resource configuration is performed for multiple model components in the target processing model based on the model resource keywords, and a computer device can automatically configure resources for the model components based on the model resource keywords, thereby reducing the difficulty of resource configuration during the large model deployment process. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 is a scenario diagram of a keyword-based model deployment system provided by an embodiment of the present invention; Figure 2 is a flow chart of an embodiment of a keyword-based model deployment method provided by an embodiment of the present invention; Figure 3 is a schematic diagram of the structure of an embodiment of a target processing model provided by an embodiment of the present invention; Figure 4 is a flowchart of a specific embodiment of generating a target processing model provided by an embodiment of the present invention; Figure 5 is a schematic diagram of a structure of a model structure provided by an embodiment of the present invention; Figure 6 is a schematic structural diagram of another embodiment of the model structure provided by an embodiment of the present invention; Figure 7 is a schematic diagram of another embodiment of the model structure provided by an embodiment of the present invention; Figure 8is a schematic structural diagram of another embodiment of the model structure provided by an embodiment of the present invention; Fig. 9 is a schematic diagram of an embodiment of a ring model structure provided by an embodiment of the present invention; Fig.10 It is a schematic diagram of an embodiment of a model structure of node isolation provided by an embodiment of the present invention; Fig.11 is a schematic diagram of another embodiment of a model structure of isolated nodes provided by an embodiment of the present invention; Fig.12 It is a flowchart of a specific embodiment of resource configuration of multiple model components in a target processing model provided by an embodiment of the present invention; Fig.13 is a flow chart of another embodiment of a keyword-based model deployment method provided by an embodiment of the present invention; Fig.14 It is a schematic diagram of the operation result when the model provided by the embodiment of the present invention is run without dynamic adjustment of resources; Fig.15 It is a schematic diagram of the operation result of dynamically adjusting resources when the model provided by the embodiment of the present invention is running; Fig.16 It is a principle block diagram of a keyword-based model deployment device provided by an embodiment of the present invention; Fig.17 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0019] In the description of this application, it should be understood that the terms "first", "second", "third", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, features defined as "first", "second", "third", etc. may explicitly or implicitly include one or more features.

[0020] In this application, the word "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described in this application as "exemplary" is not necessarily to be construed as being preferred or advantageous over other embodiments. The following description is given to enable any technician in the field to implement and use the present application. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present application can be implemented without using these specific details. In other instances, well-known structures and processes will not be elaborated in detail to avoid obscuring the description of the present application with unnecessary details. Therefore, the present application is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in the present application.

[0021] It should be noted that since the method of the embodiment of the present application is executed in a computer device, the processing objects of each computer device exist in the form of data or information. For example, time is actually time information. It can be understood that if size, quantity, position, etc. are mentioned in subsequent embodiments, they are all corresponding data for processing by the computer device. The details will not be repeated here.

[0022] The embodiments of the present application provide a keyword-based model deployment method, apparatus, device, and storage medium, which are described in detail below.

[0023] See also Figure 1 , Figure 1 A scenario diagram of a keyword-based model deployment system provided in an embodiment of the present application, wherein the keyword-based model deployment system may include a computer device 100, wherein the computer device 100 is integrated with a keyword-based model deployment device, such as Figure 1 Computer equipment in.

[0024] In the embodiment of the present application, the computer device 100 is mainly used to obtain model structure keywords and model resource keywords; based on the model structure keywords, a target processing model is generated; the target processing model includes multiple model components; based on the model resource keywords, resources are configured for multiple model components in the target processing model. The computer device can automatically build a large model service based on the model structure keywords and automatically configure resources for the model components based on the model resource keywords, thereby simplifying the deployment process of the large model service and reducing the difficulty of resource configuration during the large model deployment process.

[0025] In the embodiment of the present application, the computer device 100 may be an independent server, or a server network or server cluster composed of servers. For example, the computer device 100 described in the embodiment of the present application includes but is not limited to a computer, a network host, a single network server, a plurality of network server sets or a cloud server composed of a plurality of servers. The cloud server is composed of a large number of computers or network servers based on cloud computing.

[0026] It is understandable that the computer device 100 used in the embodiments of the present application may be a device including both receiving and transmitting hardware, that is, a device having receiving and transmitting hardware capable of performing two-way communication on a two-way communication link. Such a device may include: a cellular or other communication device having a single-line display or a multi-line display or a cellular or other communication device without a multi-line display. The computer device 100 may specifically be a desktop terminal or a mobile terminal, and the computer device 100 may specifically be one of a television, a mobile phone, a tablet computer, a laptop computer, and the like.

[0027] Those skilled in the art will understand that Figure 1 The application environment shown in the figure is only one application scenario of the present application solution and does not constitute a limitation on the application scenario of the present application solution. Other application environments may also include Figure 1 More or less computer equipment as shown in Figure 1 Only one computer device is shown in the figure. It can be understood that the keyword-based model deployment system can also include one or more other services, which are not specifically limited here.

[0028] In addition, if Figure 1 As shown, the keyword-based model deployment system may also include a memory 200 for storing data, such as keyword information, such as model structure keywords, model resource keywords, etc., such as verification result information, such as first verification result information, second verification result information, etc.

[0029] It should be noted that Figure 1 The scenario diagram of the keyword-based model deployment system shown is merely an example. The keyword-based model deployment system and scenario described in the embodiment of the present application are intended to more clearly illustrate the technical solution of the embodiment of the present application, and do not constitute a limitation on the technical solution provided in the embodiment of the present application. A person of ordinary skill in the art will appreciate that with the evolution of the keyword-based model deployment system and the emergence of new business scenarios, the technical solution provided in the embodiment of the present application is equally applicable to similar technical problems.

[0030] First, a keyword-based model deployment method is provided in an embodiment of the present application. The executor of the keyword-based model deployment method is a keyword-based model deployment device. The keyword-based model deployment device is applied to a computer device. The keyword-based model deployment method includes: obtaining model structure keywords and model resource keywords; based on the model structure keywords, generating a target processing model; the target processing model includes multiple model components; based on the model resource keywords, performing resource configuration on multiple model components in the target processing model.

[0031] like Figure 2 As shown, it is a flowchart of an embodiment of a keyword-based model deployment method in an embodiment of the present application. The keyword-based model deployment method may include steps S201 to S203, which are as follows: S201. Obtain model structure keywords and model resource keywords.

[0032] In an embodiment of the present application, the model structure keyword is a keyword related to the model structure of the large model. The model structure keyword includes multiple model keywords and dependency relationships between multiple model keywords. For example, when the model structure keyword is "input: model a, model a: model b, model c, model b: model c, model d, model c: model d, model d: model e, model e: output", the model structure keyword includes 7 model keywords, namely, input, model a, model b, model c, model d, model e and output. Model a depends on input, model b and model c depend on model a, model c and model d depend on model b, model d depends on model c, model e depends on model d, and output depends on model e. That is to say, model a is used to receive input information, model b and model c respectively receive the output of model a, model c and model d respectively receive the output of model b, model d receives the output of model c, model e receives the output of model d, and model e is used to output the final information.

[0033] Furthermore, the model resource keyword is resource configuration information associated with the model components corresponding to the multiple model keywords, and the model resource keyword includes the required resource type and required resource amount of the model components corresponding to the multiple model keywords. The required resource type may include any one or more of a graphics processing unit (GPU), a central processing unit (CPU), a random access memory (RAM), and a custom computing resource. For example, the model components corresponding to the multiple model keywords include model a, model b, model c, model d, and model e, and the model resource keyword may be "model a: GPU: 2, model b: GPU: 1, model c: GPU: 0.5, model d: GPU: 0.25, model e: GPU: 0.25", that is, based on the model resource keyword, it can be determined that the required resource type corresponding to model a, model b, model c, model d, and model e is a graphics processing unit (GPU), and the required resource amounts corresponding to model a, model b, model c, model d, and model e are 2, 1, 0.5, 0.25, and 0.25, respectively.

[0034] Optionally, the computer device can obtain the model structure keywords and model resource keywords in a variety of ways. For example, the computer device can receive the model structure keywords and model resource keywords input by the user through input devices such as a keyboard, mouse, and touch screen. The computer device can also obtain the model structure keywords and model resource keywords from other devices through the network, Bluetooth, etc., which is not limited in this embodiment.

[0035] S202. Generate a target processing model based on the model structure keywords; the target processing model includes multiple model components.

[0036] In an embodiment of the present application, the target processing model is a large model service deployed based on a model structure keyword. The target processing model includes multiple model components. This embodiment generates a target processing model based on the model structure keyword. The user only needs to manually input the model structure keyword, and the computer device can automatically deploy the large model service based on the model structure keyword, thereby simplifying the deployment process of the large model service.

[0037] In some embodiments, based on the model structure keyword, the step of generating the target processing model specifically includes: generating the target processing model based on multiple model keywords in the model structure keyword and the dependency relationship between the multiple model keywords. For example, when the model structure keyword is "input: model a, model a: model b, model c, model b: model c, model d, model c: model d, model d: model e, model e: output", based on the multiple model keywords in the model structure keyword and the dependency relationship between the multiple model keywords, the target processing model can be generated. Figure 3 The target processing model shown includes five model components: model a, model b, model c, model d and model e.

[0038] In other embodiments, reference Figure 4 As shown, in the above step S202, generating a target processing model based on the model structure keyword may include steps S301 to S304, which are specifically as follows: S301. Verify multiple model keywords in the model structure keyword and the dependency relationship between the multiple model keywords.

[0039] In some embodiments, the step of verifying multiple model keywords in the model structure keywords and the dependency relationships between the multiple model keywords specifically includes: determining the corresponding model structure based on the multiple model keywords in the model structure keywords and the dependency relationships between the multiple model keywords; if the model structure includes a preset model structure, determining that the verification of the multiple model keywords and the dependency relationships between the multiple model keywords has failed; if the model structure does not include a preset model structure, determining that the verification of the multiple model keywords and the dependency relationships between the multiple model keywords has passed.

[0040] In the embodiment of the present application, the model structure is a model structure determined based on multiple model keywords in the model structure keyword and the dependency relationship between the multiple model keywords. For example, when the model structure keyword is "model c: model d, model a: model b, model c, model b: model c, input: model a, model e: output", based on the model structure keyword "model c: model d" can be determined Figure 5 The model structure shown in the figure can be determined based on the model structure keywords "model a: model b, model c, model b: model c" Figure 6 The model structure shown in the figure can be determined based on the "input: model a" in the model structure keyword. Figure 7 The model structure shown can be determined based on the "model e: output" in the model structure keyword Figure 8 The model structure shown.

[0041] Furthermore, the preset model structure is a preset invalid model structure, and the preset model structure can measure whether the model structure is valid. Optionally, the preset model structure can include a ring model structure and a node isolated model structure. For example, referring to Figures 9 to 11 As shown, Fig. 9 FIG. 1 is a schematic diagram of an embodiment of a loop model structure, in which the input and output of multiple model components form a loop. Fig.10 and Fig.11 It is a schematic diagram of a specific embodiment of a node-isolated model structure, in which the model component has only output but no input or the model component has only input but no output.

[0042] S302: If the plurality of model keywords and the dependency relationships between the plurality of model keywords are verified, a target processing model is generated based on the plurality of model keywords and the dependency relationships between the plurality of model keywords.

[0043] In some embodiments, if the dependency relationship between multiple model keywords and the multiple model keywords is verified, the target processing model can be directly generated based on the multiple model keywords and the dependency relationship between the multiple model keywords. For example, when the model structure keyword is "input: model a, model a: model b, model c, model b: model c, model d, model c: model d, model d: model e, model e: output", the model structure determined based on the multiple model keywords in the model structure keyword and the dependency relationship between the multiple model keywords does not contain a preset model structure, that is, the multiple model keywords in the model structure keyword and the dependency relationship between the multiple model keywords are verified, then the target processing model can be directly generated based on the multiple model keywords in the model structure keyword and the dependency relationship between the multiple model keywords.

[0044] S303: If the verification of the plurality of model keywords and the dependency relationships between the plurality of model keywords fails, output first verification result information.

[0045] In an embodiment of the present application, the first verification result information is used to characterize that a plurality of model keywords and the dependency verification between the plurality of model keywords have failed. Furthermore, in order to facilitate the user to adjust the model structure keywords based on the first verification result information, the first verification result information may include the reason for the failure of the verification. For example, the first verification result information is "the model structure corresponding to 'model a: model b, model b: model c, model c: model a' forms a loop and the verification fails."

[0046] In an embodiment of the present application, the computer device can output the first verification result information in a variety of ways. For example, the computer device can output the first verification result information through a display screen configured by itself, the computer device can also output the first verification result information through a microphone configured by itself, and the computer device can also transmit the first verification result information to other devices via a network, Bluetooth, etc., and output the first verification result information through other devices.

[0047] S304: Obtain the adjusted model structure keyword input by the user based on the first verification result information, and generate a target processing model based on the adjusted model structure keyword.

[0048] Optionally, after the computer outputs the first verification result information, the user can adjust the model structure keywords based on the first verification result information, the computer device obtains the adjusted model structure keywords input by the user based on the first verification result information, and generates a target processing model based on the adjusted model structure keywords, thereby improving the accuracy of the generated target processing model.

[0049] In some embodiments, the step of generating a target processing model based on the adjusted model structure keyword specifically includes: determining the adjusted model structure keyword as a new model structure keyword, and continuing to execute the step of verifying multiple model keywords in the model structure keyword and the dependency relationship between the multiple model keywords until the multiple model keywords and the dependency relationship between the multiple model keywords are verified, and generating a target processing model based on the multiple model keywords and the dependency relationship between the multiple model keywords.

[0050] S203: Based on the model resource keywords, perform resource configuration on multiple model components in the target processing model.

[0051] In an embodiment of the present application, after the target processing model is generated, resource configuration is performed on multiple model components in the target processing model based on the model resource keywords. In this way, the computer device can automatically perform resource configuration on multiple model components in the target processing model based on the model resource keywords, thereby solving the problem that developers need to modify the code when manually configuring model resources, resulting in high difficulty in resource configuration.

[0052] In some embodiments, based on the model resource keyword, the step of configuring resources for multiple model components in the target processing model specifically includes: determining the required resource type and required resource amount corresponding to each model component from the model resource keyword, and allocating computing resources in the computing cluster to multiple model components in the target processing model based on the required resource type and required resource amount. For example, when the model resource keyword is "model a: GPU: 2, model b: GPU: 1, model c: GPU: 0.5, model d: GPU: 0.25, model e: GPU: 0.25", 2 GPUs in the computing cluster are allocated to model a, 1 GPU in the computing cluster is allocated to model b, 0.5 GPUs in the computing cluster are allocated to model c, 0.25 GPUs in the computing cluster are allocated to model d, and 0.25 GPUs in the computing cluster are allocated to model e.

[0053] In other embodiments, referring to Fig.12 As shown, in the above step S203, resource configuration is performed on multiple model components in the target processing model based on the model resource keyword, which may include steps S401 to S404, which are specifically as follows: S401: Verify model resource keywords.

[0054] In some embodiments, the model resource keyword includes the required resource type and required resource amount corresponding to each model component. The above-mentioned step of verifying the model resource keyword specifically includes: matching the required resource type corresponding to each model component with the resource type in the computing cluster; if the required resource type does not match the resource type in the computing cluster, determining that the model resource keyword verification has failed; if the required resource type matches the resource type in the computing cluster, determining the total required resource amount of each required resource type based on the required resource type and required resource amount corresponding to each model component; if the total required resource amount of any required resource type is greater than the total resource amount of the required resource type in the computing cluster, determining that the model resource keyword verification has failed.

[0055] In an embodiment of the present application, matching the required resource type with the resource type in the computing cluster means that the required resource type is the resource type in the computing cluster. For example, if the required resource type is GPU, and the resource types in the computing cluster include GPU and CPU, then the required resource type matches the resource type in the computing cluster. Conversely, if the required resource type is GPU, and the resource types in the computing cluster include CPU and ARM, then the required resource type does not match the resource type in the computing cluster.

[0056] Furthermore, when determining the total amount of required resources for each required resource type based on the required resource type and required resource amount corresponding to each model component, the required resource amounts of the same required resource type in the required resource types corresponding to multiple model components can be added together to obtain the total amount of required resources for each required resource type.

[0057] Optionally, if the required resource type does not match the resource type in the computing cluster, it is determined that the model resource keyword verification has failed; conversely, if the required resource type matches the resource type in the computing cluster, the total required resource amount of each required resource type is determined based on the required resource type and required resource amount corresponding to each model component, and the verification result of the model resource keyword is determined based on the total required resource amount of each required resource type.

[0058] In some embodiments, based on the total amount of required resources of each required resource type, the step of determining the verification result of the model resource keyword specifically includes: comparing the total amount of required resources of each required resource type with the total amount of resources of the required resource type in the computing cluster, if the total amount of required resources of any required resource type is greater than the total amount of resources of the required resource type in the computing cluster, determining that the model resource keyword verification has failed; conversely, if the total amount of required resources of each required resource type is less than the total amount of resources of the required resource type in the computing cluster, determining that the model resource keyword verification has passed. For example, if the total amount of required resources of the GPU is 5, and the total amount of resources of the GPU in the computing cluster is 4, then the total amount of required resources of the GPU is greater than the total amount of resources of the GPU in the computing cluster, and the model resource keyword verification has failed.

[0059] S402: If the model resource keyword verification passes, determine the required resource type and required resource amount corresponding to each model component from the model resource keyword, and allocate computing resources in the computing cluster to multiple model components in the target processing model based on the required resource type and required resource amount.

[0060] In an embodiment of the present application, if the model resource keyword verification passes, the required resource type and required resource amount corresponding to each model component are determined directly from the model resource keyword, and the computing resources in the computing cluster are allocated to multiple model components in the target processing model based on the required resource type and required resource amount.

[0061] S403: If the model resource keyword verification fails, output second verification result information.

[0062] In an embodiment of the present application, the second verification result information is used to characterize that the model resource keyword verification has failed. Furthermore, in order to facilitate the user to adjust the model resource keyword based on the second verification result information, the second verification result information may include the reason why the model resource keyword verification has not passed. For example, the second verification result information is "GPU resources exceed the total resources and the verification failed". The second verification result information may also be "GPU resource demand is 5, the total GPU resources are 4, the GPU resource demand exceeds the total resources, and the verification failed". The second verification result information may also be "GPU resource type does not match and the verification failed".

[0063] In an embodiment of the present application, the computer device can output the second verification result information in a variety of ways. For example, the computer device can output the second verification result information through a display screen configured by itself, the computer device can also output the second verification result information through a microphone configured by itself, and the computer device can also transmit the second verification result information to other devices via a network, Bluetooth, etc., and output the second verification result information through other devices.

[0064] S404: Obtain the adjusted model resource keyword input by the user based on the second verification result information, and perform resource configuration on multiple model components in the target processing model based on the adjusted model resource keyword.

[0065] Optionally, after the computer device outputs the second verification result information, the user can adjust the model resource keywords based on the second verification result information. The computer device obtains the adjusted model resource keywords input by the user based on the second verification result information, and performs resource configuration for multiple model components in the target processing model based on the adjusted model resource keywords, thereby improving the accuracy of resource configuration.

[0066] In some embodiments, the step of configuring resources for multiple model components in the target processing model based on the adjusted model resource keyword specifically includes: determining the adjusted model resource keyword as a new model resource keyword, and continuing to execute the step of verifying the model resource keyword until the model resource keyword verification passes, determining the required resource type and required resource amount corresponding to each model component from the model resource keyword, and allocating computing resources in the computing cluster to the multiple model components in the target processing model based on the required resource type and required resource amount.

[0067] In some embodiments, reference Fig.13 As shown, after performing resource configuration on multiple model components in the target processing model based on the model resource keyword in the above step S203, steps S204 to S205 may be included, which are specifically as follows: S204. When the target processing model is running, the running status information of multiple model components is obtained.

[0068] In an embodiment of the present application, the running status information of multiple model components is used to characterize whether the multiple model components are running. The running status information of multiple model components may include running, and / or already running, and / or not running.

[0069] S205. Dynamically adjust the resources occupied by multiple model components based on the running status information.

[0070] In some embodiments, the step of dynamically adjusting the resources occupied by multiple model components based on the running status information specifically includes: determining a fourth model component from the multiple model components based on the running status information, the fourth model component being an already running model component among the multiple model components; releasing the resources occupied by the fourth model component, thereby preventing the already running model components in the target processing model from continuing to occupy computing resources, thereby improving resource utilization.

[0071] In other embodiments, the step of dynamically adjusting the resources occupied by multiple model components based on the running status information specifically includes: determining whether the multiple model components include the first model component based on the running status information; if the multiple model components include the first model component, determining whether the multiple model components include the second model component; if the multiple model components include the second model component, dynamically adjusting the resources of the second model component based on the amount of resources required for the tasks in the queue in the first model component.

[0072] In the embodiment of the present application, the first model component is a model component that is currently running and the corresponding task is in a queue state among the multiple model components, for example, Figure 3 For example, if task 1 is being executed in model a, and tasks 2 and 3 are in a queue waiting for model a to execute, then model a is the first model component. The second model component is the model component that is not running and occupies resources among multiple model components. For example, continue to refer to Figure 3 As shown, if model d and model e have not been run, model d and model e are determined to be the second model components.

[0073] Optionally, if the first model component is included in multiple model components, it means that the resources configured for the first model component are insufficient. The resources of the second model component are dynamically adjusted based on the amount of resources required for the tasks in the queue state in the first model component. The idle resources occupied by the second model component can be allocated to the first model component for task processing, thereby avoiding resource waste caused by idle resources and shortening the running time of multi-task scenarios.

[0074] In some embodiments, the above-mentioned step of dynamically adjusting the resources of the second model component based on the amount of resources required by the tasks in the queued state in the first model component specifically includes: if the amount of resources required by the tasks in the queued state in the first model component is greater than or equal to the total amount of resources occupied by the second model component, the resources occupied by the second model component are allocated to the first model component; if the amount of resources required by the tasks in the queued state in the first model component is less than the total amount of resources occupied by the second model component, the priority of the second model component is determined based on the dependency relationship of the second model component; the third model component is determined from the second model component based on the priority and the amount of resources required by the tasks in the queued state in the first model component; and the resources occupied by the third model component are allocated to the first model component.

[0075] In an embodiment of the present application, when determining the priority of a second model component based on the dependency relationship of the second model component, the priority of the dependent party in the two model components with a dependency relationship can be set higher than the priority of the dependent party, or the priority of the dependent party in the two model components with a dependency relationship can be set higher than the priority of the dependent party, and this embodiment does not limit this. For example: when model b depends on model a, the priority of model a can be set higher than the priority of model b, or the priority of model a can be set lower than the priority of model b.

[0076] Furthermore, the third model component is a model component with a higher priority in the second model component, and the total amount of resources occupied by the third model component is greater than or equal to the amount of resources required by the tasks in the queued state in the first model component. Optionally, when the priority of the dependent party is higher than the priority of the dependent party in two model components with a dependency relationship, the third model component is a model component that is ranked higher in priority from low to high; when the priority of the dependent party is higher than the priority of the dependent party in two model components with a dependency relationship, the third model component is a model component that is ranked higher in priority from high to low. For example, Figure 3 Take the target model in as an example, if model a is the first model component, the amount of resources required for the tasks in the queued state in model a is 1 GPU, if model d and model e are both model components that are not running and occupy resources, model d and model e both occupy 1 GPU and model e depends on model d, and according to the dependency relationship, it is determined that the priority of model e is higher than that of model d, then the 1 GPU resource occupied by model e is allocated to model a to execute the tasks in the queued state.

[0077] For example, refer to Fig.14 and Fig.15As shown, GPU1, GPU2 and GPU3 are allocated to model a, model b and model c respectively. Model b depends on model a, and model c depends on model b. If the resources allocated to model a, model b and model c are not dynamically adjusted, GPU2 allocated to model b will have a certain amount of idle time during the running stage of model a, and GPU3 allocated to model c will have a certain amount of idle time during the running stages of model a and model b, which will lead to idle resource waste. Fig.15 As shown, if the resources GPU2 and GPU3 allocated to model b and model c are allocated to GPU1 when model a is running, and GPU1, GPU2 and GPU3 are allocated to model b after model a is finished running, and GPU1, GPU2 and GPU3 are allocated to model c after model b is finished running, all resources allocated to the target processing model can be fully utilized during the entire target processing model running phase, which can improve resource utilization and shorten the running time of multi-task scenarios.

[0078] In order to better implement the keyword-based model deployment method in the embodiment of the present application, on the basis of the keyword-based model deployment method, the embodiment of the present application also provides a keyword-based model deployment device, such as Fig.16 As shown, the keyword-based model deployment device 600 includes: Information acquisition module 610, used to acquire model structure keywords and model resource keywords; A model generation module 620 is used to generate a target processing model based on the model structure keyword; the target processing model includes multiple model components; The resource configuration module 630 is used to perform resource configuration on multiple model components in the target processing model based on the model resource keywords.

[0079] In an embodiment of the present application, by obtaining a model structure keyword and generating a target processing model based on the model structure keyword, a computer device can automatically build a large model service based on the model structure keyword, without the need for a user to manually build the large model service, thereby simplifying the deployment process of the large model service; by obtaining a model resource keyword and performing resource configuration for multiple model components in the target processing model based on the model resource keyword, a computer device can automatically configure resources for the model components based on the model resource keyword, thereby reducing the difficulty of resource configuration during the large model deployment process.

[0080] In some embodiments of the present application, the keyword-based model deployment device 600 further includes: A status acquisition module is used to obtain the running status information of multiple model components when the target processing model is running; The resource adjustment module is used to dynamically adjust the resources occupied by multiple model components based on the running status information.

[0081] In some embodiments of the present application, the model generation module 620 is specifically used to: Verify multiple model keywords in the model structure keyword and the dependency relationship between multiple model keywords; If the multiple model keywords and the dependency relationships between the multiple model keywords are verified, a target processing model is generated based on the multiple model keywords and the dependency relationships between the multiple model keywords; If the verification of the plurality of model keywords and the dependency relationship between the plurality of model keywords fails, outputting first verification result information; The adjusted model structure keyword input by the user based on the first verification result information is obtained, and a target processing model is generated based on the adjusted model structure keyword.

[0082] In some embodiments of the present application, the model generation module 620 is further used to: Determine a corresponding model structure based on multiple model keywords in the model structure keyword and dependencies between the multiple model keywords; If the model structure includes a preset model structure, it is determined that the dependency check between multiple model keywords and multiple model keywords fails; If the model structure does not include a preset model structure, it is determined that a plurality of model keywords and dependency relationships between the plurality of model keywords are verified.

[0083] In some embodiments of the present application, the resource configuration module 630 is specifically used to: Verify model resource keywords; If the model resource keyword verification passes, determine the required resource type and required resource amount corresponding to each model component from the model resource keyword, and allocate computing resources in the computing cluster to multiple model components in the target processing model based on the required resource type and required resource amount; If the model resource keyword verification fails, output the second verification result information; The adjusted model resource keyword input by the user based on the second verification result information is obtained, and resource configuration is performed on multiple model components in the target processing model based on the adjusted model resource keyword.

[0084] In some embodiments of the present application, the model resource keyword includes the required resource type and required resource amount corresponding to each model component, and the resource configuration module 630 is specifically used to: Match the required resource type corresponding to each model component with the resource type in the computing cluster; If the required resource type does not match the resource type in the computing cluster, it is determined that the model resource keyword check fails; If the required resource type matches the resource type in the computing cluster, determine the total required resource amount of each required resource type based on the required resource type and required resource amount corresponding to each model component; If the total amount of required resources of any required resource type is greater than the total amount of resources of the required resource type in the computing cluster, it is determined that the model resource keyword verification fails.

[0085] In some embodiments of the present application, the resource adjustment module is specifically used to: Based on the running status information, determining whether the multiple model components include a first model component; the first model component is a model component that is running and the corresponding task is in a queued state among the multiple model components; If the multiple model components include the first model component, determine whether the multiple model components include the second model component; the second model component is a model component that has not yet run and occupies resources among the multiple model components; If the multiple model components include a second model component, the resources of the second model component are dynamically adjusted based on the amount of resources required by the tasks in the queue state in the first model component.

[0086] In some embodiments of the present application, the resource adjustment module is further used to: If the amount of resources required by the tasks in the queued state in the first model component is greater than or equal to the total amount of resources occupied by the second model component, the resources occupied by the second model component are allocated to the first model component; If the amount of resources required by the tasks in the queued state in the first model component is less than the total amount of resources occupied by the second model component, determining the priority of the second model component based on the dependency relationship of the second model component; Determine a third model component from the second model component based on the priority and the amount of resources required by the tasks in the queued state in the first model component; the third model component is a model component with a higher priority in the second model component, and the total amount of resources occupied by the third model component is greater than or equal to the amount of resources required by the tasks in the queued state in the first model component; The resources occupied by the third model component are allocated to the first model component.

[0087] The present application also provides a computer device, the computer device comprising: one or more processors; Memory; and One or more applications, wherein the one or more applications are stored in a memory and configured to be executed by a processor in the steps of the keyword-based model deployment method in any of the above-mentioned keyword-based model deployment method embodiments.

[0088] The present application also provides a computer device, such as Fig.17 As shown, it shows a schematic diagram of the structure of the computer device involved in the embodiment of the present application, specifically: The computer device may include one or more processing core processors 801, one or more computer-readable storage media memories 802, a power supply 803, an input unit 804 and other components. Those skilled in the art will appreciate that Fig.17 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. Among them: The processor 801 is the control center of the computer device. It uses various interfaces and lines to connect various parts of the entire computer device. By running or executing software programs and / or modules stored in the memory 802 and calling data stored in the memory 802, it executes various functions of the computer device and processes data, thereby monitoring the computer device as a whole. Optionally, the processor 801 may include one or more processing cores; optionally, the processor 801 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 801.

[0089] The memory 802 can be used to store software programs and modules. The processor 801 executes various functional applications and data processing by running the software programs and modules stored in the memory 802. The memory 802 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 802 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 802 may also include a memory controller to provide the processor 801 with access to the memory 802.

[0090] The computer device also includes a power supply 803 for supplying power to various components. Optionally, the power supply 803 can be logically connected to the processor 801 through a power management system, so as to manage charging, discharging, and power consumption through the power management system. The power supply 803 can also include any components such as one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, and power status indicators.

[0091] The computer device may further include an input unit 804, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0092] Although not shown, the computer device may further include a display unit, etc., which will not be described in detail herein. Specifically in this embodiment, the processor 801 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 802 according to the following instructions, and the processor 801 will run the application programs stored in the memory 802, thereby realizing various functions, as follows: Get model structure keywords and model resource keywords; Based on the model structure keywords, a target processing model is generated; the target processing model includes multiple model components; Based on the model resource keywords, resources are configured for multiple model components in the target processing model.

[0093] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0094] To this end, the embodiment of the present application provides a computer-readable storage medium, which may include: a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any keyword-based model deployment method provided in the embodiment of the present application. For example, the computer program can be loaded by a processor to execute the following steps: Get model structure keywords and model resource keywords; Based on the model structure keywords, a target processing model is generated; the target processing model includes multiple model components; Based on the model resource keywords, resources are configured for multiple model components in the target processing model.

[0095] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the detailed description of other embodiments above, and will not be repeated here.

[0096] In specific implementation, the above units or structures can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units or structures can refer to the previous method embodiments, which will not be repeated here.

[0097] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0098] The above is a detailed introduction to a keyword-based model deployment method, device, equipment and storage medium provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A keyword-based model deployment method, characterized in that: include: Get model structure keywords and model resource keywords; Based on the model structure keywords, generating a target processing model; The target processing model includes a plurality of model components; Based on the model resource keywords, resource configuration is performed on the plurality of model components in the target processing model.

2. The keyword-based model deployment method according to claim 1, characterized in that: After performing resource configuration on the plurality of model components in the target processing model based on the model resource keyword, the method includes: When the target processing model is running, obtaining the running status information of the plurality of model components; The resources occupied by the plurality of model components are dynamically adjusted based on the running status information.

3. The keyword-based model deployment method according to claim 1, characterized in that: The generating a target processing model based on the model structure keyword comprises: Verifying a plurality of model keywords in the model structure keyword and dependency relationships between the plurality of model keywords; If the plurality of model keywords and the dependency relationships between the plurality of model keywords are verified, generating a target processing model based on the plurality of model keywords and the dependency relationships between the plurality of model keywords; If the verification of the plurality of model keywords and the dependency relationship between the plurality of model keywords fails, outputting first verification result information; The adjusted model structure keywords input by the user based on the first verification result information are obtained, and a target processing model is generated based on the adjusted model structure keywords.

4. The keyword-based model deployment method according to claim 3, characterized in that: The verifying of the plurality of model keywords in the model structure keyword and the dependency relationship between the plurality of model keywords includes: Determine a corresponding model structure based on a plurality of model keywords in the model structure keyword and dependencies between the plurality of model keywords; If the model structure includes a preset model structure, determining that the dependency check between the plurality of model keywords and the plurality of model keywords fails; If the model structure does not include the preset model structure, it is determined that the plurality of model keywords and the dependency relationship between the plurality of model keywords are verified.

5. The keyword-based model deployment method according to claim 1, characterized in that: The configuring resources for the plurality of model components in the target processing model based on the model resource keyword includes: Verifying the model resource keywords; If the model resource keyword verification passes, determining the required resource type and required resource amount corresponding to each of the model components from the model resource keyword, and allocating computing resources in the computing cluster to the multiple model components in the target processing model based on the required resource type and the required resource amount; If the model resource keyword verification fails, outputting second verification result information; The adjusted model resource keyword input by the user based on the second verification result information is obtained, and resource configuration is performed on the plurality of model components in the target processing model based on the adjusted model resource keyword.

6. The keyword-based model deployment method according to claim 5, characterized in that: The model resource keywords include the required resource type and required resource amount corresponding to each of the model components; The verifying of the model resource keyword includes: Matching the required resource type corresponding to each of the model components with the resource type in the computing cluster; If the required resource type does not match the resource type in the computing cluster, determining that the model resource keyword check fails; If the required resource type matches the resource type in the computing cluster, determining the total required resource amount of each required resource type based on the required resource type and required resource amount corresponding to each model component; If the total amount of required resources of any one of the required resource types is greater than the total amount of resources of the required resource type in the computing cluster, it is determined that the model resource keyword check has failed.

7. The keyword-based model deployment method according to claim 2, characterized in that: The dynamically adjusting the resources occupied by the plurality of model components based on the running status information includes: Based on the running status information, determining whether the plurality of model components include a first model component; the first model component is a model component that is running and whose corresponding tasks are in a queued state among the plurality of model components; If the plurality of model components include the first model component, determining whether the plurality of model components include a second model component; the second model component is a model component among the plurality of model components that has not been run and occupies resources; If the second model component is included in the plurality of model components, resources of the second model component are dynamically adjusted based on the amount of resources required by the tasks in the queued state in the first model component.

8. The keyword-based model deployment method according to claim 7, characterized in that: The dynamically adjusting the resources of the second model component based on the amount of resources required by the tasks in the queued state in the first model component includes: If the amount of resources required by the tasks in the queued state in the first model component is greater than or equal to the total amount of resources occupied by the second model component, allocating the resources occupied by the second model component to the first model component; If the amount of resources required by the tasks in the queued state in the first model component is less than the total amount of resources occupied by the second model component, determining the priority of the second model component based on the dependency relationship of the second model component; Determine a third model component from the second model component based on the priority and the amount of resources required by the tasks in the queued state in the first model component; the third model component is a model component with a higher priority in the second model component, and the total amount of resources occupied by the third model component is greater than or equal to the amount of resources required by the tasks in the queued state in the first model component; Allocate resources occupied by the third model component to the first model component.

9. A keyword-based model deployment device, characterized in that: include: Information acquisition module, used to obtain model structure keywords and model resource keywords; A model generation module, used to generate a target processing model based on the model structure keywords; The target processing model includes a plurality of model components; A resource configuration module is used to perform resource configuration on the plurality of model components in the target processing model based on the model resource keywords.

10. A computer device, characterized in that: The computer device comprises: one or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the keyword-based model deployment method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in the keyword-based model deployment method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Deployment constraint automatic detection method for Web application

    CN101957794A

  • Application configuration deploying method and device

    CN107766060A

  • Resource adjustment method and device

    CN110597614A

  • Resource management method and system, electronic equipment and storage medium

    CN112269659A

  • Cloud resource dynamic allocation method oriented to complex information system

    CN112995341A