Data processing method and device, storage medium and electronic equipment

By selecting the target GPU based on GPU resource priority and availability in a serverless architecture, resources are provided for generative AI applications, solving the problem of GPU resource limitations and achieving stable application operation and efficient resource allocation.

CN120849084APending Publication Date: 2025-10-28ALIBABA CLOUD COMPUTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410526130.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-28
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In a serverless architecture, when multiple generative AI applications request the same GPU model, resource limitations may cause some applications to degrade performance or even fail to run properly.

Method used

By identifying target graphics cards from multiple graphics cards based on their resource priority and availability, resources are provided for generative AI applications. This includes priority determination, resource prediction, and candidate graphics card selection to ensure the stable operation of the applications.

Benefits of technology

Effective management and allocation of graphics card resources ensures the continuity and performance of generative artificial intelligence applications, improving application stability and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849084A_ABST
    Figure CN120849084A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, a storage medium and electronic equipment. The method relates to the technical field of data processing, and comprises the steps that a target data processing request is acquired, and the target data processing request at least carries data information to be processed; determining a target generation type artificial intelligence application of the to-be-processed target data processing request; determining a target graphics card from a plurality of graphics cards corresponding to the target generative artificial intelligence application according to the graphics card resource priority and the graphics card resource availability; and running the target generation type artificial intelligence application in the target display card to process the to-be-processed data information through the target generation type artificial intelligence application to obtain a processing result. According to the method and the device, the technical problem that the generative artificial intelligence application cannot run due to resource limitation when a plurality of generative artificial intelligence applications request the same type of display card resources for data processing in the related technology is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more specifically, to a data processing method and apparatus, a storage medium and an electronic device. Background Technology

[0002] With the rapid development of cloud computing technology, serverless architecture, as a new computing model, is gradually becoming the preferred solution for enterprises and developers to deploy applications. Serverless architecture enables developers to focus on code and application logic by automatically managing underlying computing resources, without having to worry about server operation and management. This model has significant advantages in improving development efficiency, reducing costs, and enhancing application resilience.

[0003] Meanwhile, the rapid advancements in Artificial Intelligence (AI) and Artificial Intelligence Generated Content (AIGC) technologies have led to a surge in demand for computing resources, particularly Graphics Processing Units (GPUs). GPUs, due to their efficiency in parallel processing of large-scale data, have become indispensable resources for AI and AIGC applications. However, the high cost and limited supply of GPU resources have, in many cases, become bottlenecks restricting the development of these applications.

[0004] While server-less architectures offer elastic management of computing resources, traditional resource scheduling mechanisms often fail to effectively address GPU resource shortages. When multiple applications simultaneously request GPU resources of a specific type, resource limitations may cause some applications to experience performance degradation or malfunction.

[0005] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0006] This application provides a data processing method and apparatus, storage medium and electronic device to at least solve the technical problem in the related art where multiple generative artificial intelligence applications cannot run due to resource limitations when requesting the same type of graphics card resources for data processing.

[0007] According to one aspect of the embodiments of this application, a data processing method is provided, comprising: acquiring a target data processing request, wherein the target data processing request carries at least data information to be processed; determining a target generative artificial intelligence application to process the target data processing request; determining a target graphics card from a plurality of graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority and graphics card resource availability, wherein the graphics card among the plurality of graphics cards is used to provide resources for the operation of the target generative artificial intelligence application; and running the target generative artificial intelligence application on the target graphics card to process the data information to be processed through the target generative artificial intelligence application to obtain a processing result.

[0008] Further, determining the target graphics card from multiple graphics cards corresponding to the target generative AI application based on graphics card resource priority and graphics card resource availability includes: determining multiple first graphics cards with a first graphics card resource priority from multiple graphics cards corresponding to the target generative AI application based on the graphics card resource priority; detecting whether there is a first graphics card among the multiple first graphics cards with a remaining resource amount greater than the resource amount required by the target generative AI application; if there is no first graphics card among the multiple first graphics cards with a remaining resource amount greater than the resource amount required by the target generative AI application, then determining multiple second graphics cards with a second graphics card resource priority from multiple graphics cards corresponding to the target generative AI application based on the graphics card resource priority, wherein the second graphics card resource priority is lower than the first graphics card resource priority; if there is a candidate graphics card among the multiple second graphics cards with a remaining resource amount greater than the resource amount required by the target generative AI application, then determining the target graphics card based on the candidate graphics card.

[0009] Furthermore, before detecting whether there is a first graphics card among the plurality of first graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, the method further includes: determining the amount of unused resources of the first graphics cards among the plurality of first graphics cards based on the data processing requests currently being processed by the first graphics cards among the plurality of first graphics cards; predicting the amount of data to be processed by the first graphics cards among the plurality of first graphics cards at the next moment based on the data processing requests currently being processed by the first graphics cards among the plurality of first graphics cards, to obtain a predicted resource amount; and determining the remaining resource amount of the first graphics cards among the plurality of first graphics cards based on the unused resource amount and the predicted resource amount.

[0010] Furthermore, if there are no candidate graphics cards among the plurality of second graphics cards with remaining resources greater than the resources required by the target generative artificial intelligence application, the method further includes: repeatedly executing the step of determining a third graphics card with a third graphics card resource priority from among the plurality of graphics cards corresponding to the target generative artificial intelligence application based on the graphics card resource priority, until there are candidate graphics cards among the plurality of graphics cards with the current graphics card resource priority with remaining resources greater than the resources required by the target generative artificial intelligence application.

[0011] Further, determining the target graphics card based on the candidate graphics cards includes: if there are multiple candidate graphics cards, obtaining the amount of resources required by the target generative AI application and the initial resource utilization rate of the candidate graphics cards; estimating the target resource utilization rate of the candidate graphics cards based on the amount of resources required by the target generative AI application and the initial resource utilization rate; and determining the target graphics card from the candidate graphics cards based on the target resource utilization rate.

[0012] Further, determining the target graphics card from the candidate graphics cards based on the target resource utilization rate includes: determining the server where the candidate graphics card is located; detecting whether the target server has cached the code package corresponding to the target generative artificial intelligence application in the server where the candidate graphics card is located; if the target server does not have cached the code package corresponding to the target generative artificial intelligence application, then determining the target graphics card from the candidate graphics cards based on the target resource utilization rate.

[0013] Further, a target generative AI application is run on the target graphics card to process the data information to be processed, and the processing result is obtained by: acquiring parameter information corresponding to the target graphics card; adjusting the target configuration information in the configuration file corresponding to the target generative AI application based on the parameter information corresponding to the target graphics card to obtain an adjusted configuration file, wherein the target configuration information is configuration information related to the parameter information corresponding to the target graphics card; running the target generative AI application and the adjusted configuration file on the target graphics card to process the data information to be processed through the target generative AI application to obtain the processing result.

[0014] According to another aspect of the embodiments of this application, a data processing method is also provided, comprising: acquiring a target data processing request sent by a client, wherein the target data processing request carries at least data information to be processed; determining a target generative artificial intelligence application to process the target data processing request in a cloud server, determining a target graphics card from a plurality of graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority and graphics card resource availability, wherein the graphics card among the plurality of graphics cards is used to provide resources for the operation of the target generative artificial intelligence application; returning the identification information corresponding to the target graphics card to the client, wherein the target generative artificial intelligence application is run on the target graphics card to process the data information to be processed through the target generative artificial intelligence application to obtain a processing result.

[0015] According to another aspect of the embodiments of this application, a data processing apparatus is also provided, comprising: an acquisition unit, configured to acquire a target data processing request, wherein the target data processing request carries at least data information to be processed; a first determination unit, configured to determine a target generative artificial intelligence application for processing the target data processing request; a second determination unit, configured to determine a target graphics card from a plurality of graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority and graphics card resource availability, wherein the graphics card among the plurality of graphics cards is used to provide resources for the operation of the target generative artificial intelligence application; and a processing unit, configured to run the target generative artificial intelligence application on the target graphics card to process the data information to be processed through the target generative artificial intelligence application and obtain a processing result.

[0016] Further, the second determining unit includes: a first determining subunit, configured to determine a plurality of first graphics cards with a first graphics card resource priority from a plurality of graphics cards corresponding to the target generative AI application based on the graphics card resource priority; a detecting subunit, configured to detect whether there is a first graphics card among the plurality of first graphics cards with a remaining resource amount greater than the resource amount required by the target generative AI application; a second determining subunit, configured to determine a plurality of second graphics cards with a second graphics card resource priority from a plurality of graphics cards corresponding to the target generative AI application based on the graphics card resource priority if there is no first graphics card among the plurality of first graphics cards with a remaining resource amount greater than the resource amount required by the target generative AI application, wherein the resource priority of the second graphics cards is lower than the resource priority of the first graphics cards; and a third determining subunit, configured to determine the target graphics card based on the candidate graphics card if there is a candidate graphics card among the plurality of second graphics cards with a remaining resource amount greater than the resource amount required by the target generative AI application.

[0017] Furthermore, the device further includes: a third determining unit, configured to determine the amount of unused resources of the first graphics cards among the plurality of first graphics cards based on the data processing requests currently being processed by the first graphics cards among the plurality of first graphics cards before detecting whether there is a first graphics card among the plurality of first graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application; a prediction unit, configured to predict the amount of data to be processed by the first graphics cards among the plurality of first graphics cards at the next moment based on the data processing requests currently being processed by the first graphics cards among the plurality of first graphics cards, thereby obtaining a predicted resource amount; and a fourth determining unit, configured to determine the remaining resource amount of the first graphics cards among the plurality of first graphics cards based on the unused resource amount and the predicted resource amount.

[0018] Furthermore, the apparatus further includes: an execution unit, configured to, if there are no candidate graphics cards among the plurality of second graphics cards with remaining resources greater than the resources required by the target generative artificial intelligence application, repeatedly execute the step of determining a third graphics card with a third graphics card resource priority from among the plurality of graphics cards corresponding to the target generative artificial intelligence application based on the graphics card resource priority, until there are candidate graphics cards among the plurality of graphics cards with the current graphics card resource priority with remaining resources greater than the resources required by the target generative artificial intelligence application.

[0019] Furthermore, the third determining subunit includes: an acquisition module, configured to acquire, if there are multiple candidate graphics cards, the amount of resources required by the target generative artificial intelligence application and the initial resource utilization rate of the candidate graphics cards; an estimation module, configured to estimate based on the amount of resources required by the target generative artificial intelligence application and the initial resource utilization rate to obtain the target resource utilization rate of the candidate graphics cards; and a determining module, configured to determine the target graphics card from the candidate graphics cards based on the target resource utilization rate.

[0020] Further, the determining module includes: a determining submodule, used to determine the server where the candidate graphics card is located; a detecting submodule, used to detect whether there is a code package corresponding to the target generative artificial intelligence application cached in the target server among the servers where the candidate graphics card is located; and a determining submodule, used to determine the target graphics card from the candidate graphics cards based on the target resource utilization rate if there is no code package corresponding to the target generative artificial intelligence application cached in the target server.

[0021] Further, the processing unit includes: an acquisition subunit for acquiring parameter information corresponding to the target graphics card; an adjustment subunit for adjusting the target configuration information in the configuration file corresponding to the target generative AI application based on the parameter information corresponding to the target graphics card to obtain an adjusted configuration file, wherein the target configuration information is configuration information related to the parameter information corresponding to the target graphics card; and a running subunit for running the target generative AI application and the adjusted configuration file on the target graphics card to process the data information to be processed through the target generative AI application to obtain the processing result.

[0022] According to another aspect of the present invention, a computer-readable storage medium is also provided, the storage medium storing a program, wherein, when the program is executed, it controls the device where the storage medium is located to perform the data processing method described in any one of the above embodiments.

[0023] According to another aspect of the present invention, an electronic device is also provided, including a memory storing an executable program; and a processor for running the program, wherein the program executes the data processing method described in any one of the preceding embodiments.

[0024] According to another aspect of the present invention, a computer program product is also provided, the computer program product comprising a stored computer program that, when executed by a processor, implements the data processing method described in any one of the preceding embodiments.

[0025] In this embodiment, the following steps are employed: obtaining a target data processing request, wherein the target data processing request carries at least data information to be processed; determining the target generative artificial intelligence application for the target data processing request; determining the target graphics card from multiple graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority and graphics card resource availability, wherein the graphics card among the multiple graphics cards is used to provide resources for the operation of the target generative artificial intelligence application; running the target generative artificial intelligence application on the target graphics card to process the data information to be processed and obtain the processing result. This solves the technical problem in related technologies where multiple generative artificial intelligence applications request the same type of graphics card resources for data processing, resulting in the generative artificial intelligence application being unable to run due to resource limitations. In this solution, upon receiving a target data processing request that needs to be processed by the target generative AI application, the solution determines which graphics card from among multiple graphics cards corresponding to the target generative AI application will provide resources for its operation based on the graphics card resource priority. This avoids the problem of the generative AI application failing to run due to resource limitations. Through the above steps, graphics card resources can be managed and allocated more effectively. Resource configuration can be flexibly adjusted according to the needs of the generative AI application and the priority of graphics card resources to ensure the continuity and performance of the generative AI application, thereby achieving the technical effect of improving the operational stability of the generative AI application. Attached Figure Description

[0026] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0027] Figure 1 This is a hardware structure block diagram of a computer terminal provided according to Embodiment 1 of this application;

[0028] Figure 2 This is the flow chart of the data processing method provided in Embodiment 1 of this application. Figure 1 ;

[0029] Figure 3 This is the flow chart of the data processing method provided in Embodiment 1 of this application. Figure 2 ;

[0030] Figure 4 This is a flowchart of the data processing method provided according to Embodiment 2 of this application;

[0031] Figure 5 This is a schematic diagram of a data processing apparatus provided according to Embodiment 2 of this application;

[0032] Figure 6This is a schematic diagram of the data processing system provided according to Embodiment 3 of this application;

[0033] Figure 7 This is a structural block diagram of a computer terminal provided according to Embodiment 3 of this application. Detailed Implementation

[0034] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0035] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0036] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:

[0037] Serverless is a cloud computing model based on Platform as a Service (PaaS). Serverless computing provides a micro-architecture where end clients do not need to deploy, configure, or manage server services; all server services required for code execution are provided by the cloud platform.

[0038] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0039] Example 1

[0040] According to an embodiment of this application, a data processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0041] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a data processing method is shown. Figure 1 As shown, the computer terminal (or mobile device) 10 may include a processor set 102 (the processor set 102 may include, but is not limited to, processing devices such as microprocessors (MCUs) or programmable logic devices (FPGAs), and the processor set 102 may include a processor set, Figure 1 (Illustrated as 102a, 102b, ..., 102n), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0042] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0043] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the data processing method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby implementing the aforementioned data processing method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0044] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0045] The display may be a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).

[0046] Under the aforementioned operating environment, this application provides the following: Figure 2 The data processing method shown. Figure 2 This is a flowchart of a data processing method according to Embodiment 1 of this application. The method includes:

[0047] Step S201: Obtain the target data processing request, wherein the target data processing request carries at least the data information to be processed.

[0048] Optionally, a server-less architecture can be used to receive target data processing requests initiated by users. It should be noted that the target data processing request must carry at least the data information to be processed.

[0049] Step S202: Determine the target generative artificial intelligence application for the target data processing request.

[0050] Optionally, after receiving a target data processing request, the Serverless architecture determines the target generative AI application that can handle the request from among the multiple generative AI applications currently available in the Serverless architecture. For example, if the target data processing request is an image segmentation request, then the corresponding target generative AI application is an image processing AI application.

[0051] Step S203: Determine the target graphics card from multiple graphics cards corresponding to the target generative artificial intelligence application based on the graphics card resource priority and graphics card resource availability. The graphics card among the multiple graphics cards is used to provide resources for the operation of the target generative artificial intelligence application.

[0052] Optionally, multiple graphics cards that can provide runtime resources for the target generative AI application can be specified. Here, a graphics card refers to a graphics card that can provide runtime resources for the target generative AI application when there are available idle resources.

[0053] It should be noted that the graphics card models, GPU core counts, memory sizes, and vCPU counts mentioned above differ. Developers can configure a set of graphics card configurations for their Serverless AIGC applications (i.e., the aforementioned generative AI applications) through the Resource Configuration Manager, including the model, GPU core count, memory size, and vCPU count, and set the graphics card resource priorities for these configurations. It should also be noted that the graphics card optimization level can be set based on the graphics card model, response speed, or error rate.

[0054] The target graphics card for the generative AI application is determined from the aforementioned pool of graphics cards based on graphics card resource priority and availability. It should be noted that graphics card resource availability refers to whether the remaining resources on the current graphics card are sufficient to run the target generative AI application.

[0055] Step S204: Run the target generative artificial intelligence application on the target graphics card to process the data information to be processed and obtain the processing result.

[0056] Optionally, after determining the target graphics card, if the target generative AI application has already been cached on the target graphics card, the target generative AI application is run directly on the target graphics card, and then the data information to be processed is processed by the target generative AI application to obtain the corresponding processing results.

[0057] In summary, upon receiving a target data processing request that needs to be processed by the target generative AI application, the system determines which graphics card from among multiple graphics cards corresponding to the target generative AI application will provide resources for its operation based on the graphics card resource priority. This avoids the problem of the generative AI application failing to run due to resource limitations. Through the above steps, graphics card resources can be managed and allocated more effectively. Resource configuration can be flexibly adjusted according to the needs of the generative AI application and the graphics card resource priority to ensure the continuity and performance of the generative AI application, thereby achieving the technical effect of improving the operational stability of the generative AI application.

[0058] To ensure the normal operation of the target generative artificial intelligence application, the data processing method provided in Embodiment 1 of this application, which determines the target graphics card from multiple graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority and availability, includes: determining multiple first graphics cards with a first graphics card resource priority from the multiple graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority; detecting whether there is a first graphics card among the multiple first graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application; if there is no first graphics card among the multiple first graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, then determining multiple second graphics cards with a second graphics card resource priority from the multiple graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority, wherein the resource priority of the second graphics cards is lower than the resource priority of the first graphics cards; if there are candidate graphics cards among the multiple second graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, then determining the target graphics card based on the candidate graphics cards.

[0059] Optionally, among the multiple graphics cards corresponding to the target generative AI application, there is a preferred graphics card, that is, multiple first graphics cards with first graphics card resource priority. It should be noted that the graphics cards will be deployed in servers, and generally, servers with the same model of graphics cards are grouped together to form a cluster.

[0060] It should be noted that when the target generative AI application is triggered, it is necessary to first determine whether the first graphics card is available. That is, the above-mentioned detection checks whether there is a first graphics card among multiple first graphics cards with more remaining resources than the resources required by the target generative AI application. If there is a first graphics card with more remaining resources than the resources required by the target generative AI application, then the target generative AI application will run directly through that first graphics card.

[0061] If none of the multiple first graphics cards has a remaining resource amount greater than the resource amount required by the target generative AI application, then the next configuration needs to be tried automatically according to the preset graphics card resource priority order. That is, based on the graphics card resource priority, multiple second graphics cards with second graphics card resource priority are determined from the multiple graphics cards corresponding to the target generative AI application, and it is determined again whether there are candidate graphics cards with a remaining resource amount greater than the resource amount required by the target generative AI application among the second graphics cards with second graphics card resource priority. If there are the aforementioned candidate graphics cards among the second graphics cards, then the target graphics card is selected from the candidate graphics cards.

[0062] In an optional embodiment, the selection of a candidate graphics card as the target graphics card can be based on the amount of remaining resources of the candidate graphics card.

[0063] To ensure the normal operation of the target generative artificial intelligence application, in the data processing method provided in Embodiment 1 of this application, if there are no candidate graphics cards among the multiple second graphics cards with remaining resources greater than the resources required by the target generative artificial intelligence application, the method further includes: repeatedly executing the step of determining the third graphics card with the resource priority of the third graphics card from the multiple graphics cards corresponding to the target generative artificial intelligence application based on the graphics card resource priority, until there are candidate graphics cards among the multiple graphics cards with the current graphics card resource priority with remaining resources greater than the resources required by the target generative artificial intelligence application.

[0064] Optionally, if there is no candidate graphics card among the multiple second graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, then it is necessary to automatically try the next configuration again according to the preset graphics card resource priority order. That is, the third graphics card with the third graphics card resource priority is determined from the multiple graphics cards corresponding to the target generative artificial intelligence application according to the graphics card resource priority, until there is a candidate graphics card with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application.

[0065] By prioritizing graphics card resources as described above, the system can automatically switch to the next priority graphics card resource configuration when the preferred resource becomes unavailable, thereby reducing resource waiting time and waste and optimizing overall resource utilization. Furthermore, it achieves fine-grained control over resource usage, making resource scheduling more flexible and efficient.

[0066] To improve the accuracy and rationality of the remaining resource amount, in the data processing method provided in Embodiment 1 of this application, before detecting whether there is a first graphics card among the multiple first graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, the method further includes: determining the amount of unoccupied resources of the first graphics cards among the multiple first graphics cards based on the data processing requests currently being processed by the first graphics cards among the multiple first graphics cards; predicting the amount of data to be processed by the first graphics cards among the multiple first graphics cards at the next moment based on the data processing requests currently being processed by the first graphics cards among the multiple first graphics cards, to obtain the predicted resource amount; and determining the remaining resource amount of the first graphics cards among the multiple first graphics cards based on the unoccupied resource amount and the predicted resource amount.

[0067] Optionally, in order to improve the accuracy of calculating the remaining resource amount, the remaining resource amount is calculated by the following steps: determining the data processing requests currently being processed by the first graphics card among multiple first graphics cards, and determining the amount of resources already occupied by the first graphics card based on these data processing requests, thereby obtaining the aforementioned amount of unoccupied resources.

[0068] To avoid exceeding the total resources of the graphics card due to the addition of target-based generative AI applications, it is necessary to predict the amount of data to be processed by the first graphics card in the next moment based on the current data processing requests, thus obtaining an estimated resource requirement (i.e., the predicted resource amount mentioned above). Finally, the remaining resource amount is calculated based on the unused resources and the predicted resource amount.

[0069] By combining the currently used resources with the estimated resources, the remaining resources of the graphics card can be calculated more accurately.

[0070] To further improve the rationality of determining the target graphics card and the rationality of resource allocation, in the data processing method provided in Embodiment 1 of this application, determining the target graphics card based on the candidate graphics cards includes: if there are multiple candidate graphics cards, obtaining the amount of resources required by the target generative artificial intelligence application and the initial resource utilization rate of the candidate graphics cards; estimating the target resource utilization rate of the candidate graphics cards based on the amount of resources required by the target generative artificial intelligence application and the initial resource utilization rate; and determining the target graphics card from the candidate graphics cards based on the target resource utilization rate.

[0071] Optionally, after determining the candidate graphics cards, it is necessary to determine the number of candidate graphics cards. If there is only one candidate graphics card, then the candidate graphics card can be directly determined as the target graphics card.

[0072] If there are multiple candidate GPUs, the resource requirements of the target generative AI application and the initial resource utilization of the candidate GPUs are obtained. Then, based on the required resource requirements and the initial resource utilization, an estimate is made to determine the target resource utilization after using the candidate GPU to process the target generative AI application. Finally, the target GPU is determined from the candidate GPUs based on the target resource utilization.

[0073] It should be noted that calculating the resource utilization rate of the graphics card and determining the target graphics card from the candidate graphics cards based on the resource utilization rate is to allocate the resources of the graphics card more rationally, so as to improve the resource utilization rate of the graphics card and improve the rationality of resource allocation.

[0074] To improve the efficiency of subsequent data processing, in the data processing method provided in Embodiment 1 of this application, determining the target graphics card from the candidate graphics cards based on the target resource utilization rate includes: determining the server where the candidate graphics card is located; detecting whether the server where the candidate graphics card is located has a cached code package corresponding to the target generative artificial intelligence application in the target server; if the server does not have a cached code package corresponding to the target generative artificial intelligence application in the target server, then determining the target graphics card from the candidate graphics cards based on the target resource utilization rate.

[0075] Optionally, determining the target graphics card from the candidate graphics cards based on the target resource utilization rate includes the following steps: determining the server where the candidate graphics card is located, and determining whether the server already has the code package corresponding to the target generative artificial intelligence application cached in the target server. If the code package corresponding to the target generative artificial intelligence application is not cached in the target server, the candidate graphics cards are sorted according to the target resource utilization rate, and then the candidate graphics card with the higher target resource utilization rate is determined as the target graphics card based on the sorting result.

[0076] If the target server already has the code package corresponding to the target generative AI application cached, then the candidate graphics card in the target server can be directly identified as the target graphics card. By using the target server that has cached the code package corresponding to the target generative AI application, the cold start time of the application can be effectively reduced, thereby improving the processing efficiency of the target data processing request.

[0077] To ensure the stability of the subsequent target generative AI application on the target graphics card, the data processing method provided in Embodiment 1 of this application involves running a target generative AI application on the target graphics card to process the data information to be processed and obtain the processing result. This includes: obtaining parameter information corresponding to the target graphics card; adjusting the target configuration information in the configuration file corresponding to the target generative AI application based on the parameter information corresponding to the target graphics card to obtain an adjusted configuration file, wherein the target configuration information is configuration information related to the parameter information corresponding to the target graphics card; and running the target generative AI application and the adjusted configuration file on the target graphics card to process the data information to be processed and obtain the processing result.

[0078] Optionally, after determining the target graphics card, the parameter information of the target graphics card is determined, such as the graphics card model and the number of GPU cores. Then, the target configuration information in the configuration file corresponding to the target generative artificial intelligence application is adjusted according to the parameter information of the target graphics card, so that the target generative artificial intelligence application can run stably on the target graphics card and process the target data processing requests.

[0079] In an alternative embodiment, the following can be employed: Figure 3 The flowchart shown illustrates the data processing: Step 1, the developer defines a set of GPU configurations for their Server-less AIGC application through the Resource Configuration Manager, including model, number of GPU cores, memory size, and number of vCPUs, and sets GPU resource priorities for these configurations. Step 2, when the application is triggered to execute, the status detector assesses the current resource usage and application requirements. Step 3, the resource scheduler queries the availability of the preferred GPU configuration. If the preferred resource is available, it is allocated to the application. If it is unavailable, the next configuration is automatically tried according to the preset GPU resource priority order. Step 4, the application adapter adjusts the application to adapt to the new resource configuration, handles possible dependencies and configuration changes, and ensures that application performance is not affected. Step 5, after the application execution is complete, the status detector updates the resource usage, providing data support for the next scheduling.

[0080] Through the steps described above, the continuity and performance of AIGC applications can be guaranteed even in environments with high GPU resource contention. By automatically downgrading resource configuration to the next lower GPU resource priority, applications can continue to run, avoiding service interruptions due to insufficient resources. By allowing applications to automatically switch between different resource levels, users can optimize costs while meeting performance requirements. Users can flexibly configure GPU resource priorities based on the application's actual performance needs and cost sensitivity. Resource scheduling through automatic GPU downgrading optimizes resource usage, improves the resilience, continuity, and performance of AIGC applications in serverless environments, while also optimizing cost control and enhancing user experience and development efficiency.

[0081] In the data processing method provided in Embodiment 1 of this application, a target data processing request is obtained, wherein the target data processing request carries at least data information to be processed; a target generative artificial intelligence application for the target data processing request is determined; a target graphics card is determined from multiple graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority and graphics card resource availability, wherein the graphics card among the multiple graphics cards is used to provide resources for the operation of the target generative artificial intelligence application; the target generative artificial intelligence application is run on the target graphics card to process the data information to be processed and obtain the processing result. This solves the technical problem in related technologies where multiple generative artificial intelligence applications request the same type of graphics card resources for data processing, resulting in the generative artificial intelligence application being unable to run due to resource limitations. In this solution, upon receiving a target data processing request that needs to be processed by the target generative AI application, the solution determines which graphics card from among multiple graphics cards corresponding to the target generative AI application will provide resources for its operation based on the graphics card resource priority. This avoids the problem of the generative AI application failing to run due to resource limitations. Through the above steps, graphics card resources can be managed and allocated more effectively. Resource configuration can be flexibly adjusted according to the needs of the generative AI application and the priority of graphics card resources to ensure the continuity and performance of the generative AI application, thereby achieving the technical effect of improving the operational stability of the generative AI application.

[0082] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0083] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0084] Example 2

[0085] According to embodiments of this application, a data processing method is also provided, such as... Figure 4 The method includes:

[0086] Step S401: Obtain the target data processing request sent by the client, wherein the target data processing request carries at least the data information to be processed;

[0087] Step S402: In the cloud server, the target generative artificial intelligence application for the target data processing request is determined. Based on the priority and availability of graphics card resources, the target graphics card is determined from multiple graphics cards corresponding to the target generative artificial intelligence application. The graphics card among the multiple graphics cards is used to provide resources for the operation of the target generative artificial intelligence application.

[0088] Step S403: Return the identification information corresponding to the target graphics card to the client. In this step, run the target generative artificial intelligence application on the target graphics card to process the data information to be processed and obtain the processing result.

[0089] It should be noted that the specific method for determining the target graphics card from multiple graphics cards corresponding to the target generative artificial intelligence application in the cloud server based on the priority and availability of graphics card resources is the same as in Example 1, and will not be repeated here.

[0090] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0091] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0092] Example 3

[0093] According to embodiments of this application, a data processing apparatus for implementing the above-described data processing method is also provided, such as... Figure 5 As shown, the device includes: an acquisition unit 501, a first determination unit 502, a second determination unit 503, and a processing unit 504.

[0094] The acquisition unit 501 is used to acquire a target data processing request, wherein the target data processing request carries at least data information to be processed;

[0095] The first determining unit 502 is used to determine the target generative artificial intelligence application for the target data processing request to be processed.

[0096] The second determining unit 503 is used to determine the target graphics card from multiple graphics cards corresponding to the target generative artificial intelligence application based on the graphics card resource priority and graphics card resource availability, wherein the graphics card among the multiple graphics cards is used to provide resources for the operation of the target generative artificial intelligence application.

[0097] The processing unit 504 is used to run a target generative artificial intelligence application on the target graphics card to process the data information to be processed and obtain the processing result.

[0098] In the data processing apparatus provided in Embodiment 3, a target data processing request is obtained by an acquisition unit 501, wherein the target data processing request carries at least data information to be processed; a first determining unit 502 determines the target generative artificial intelligence application for the target data processing request; a second determining unit 503 determines the target graphics card from multiple graphics cards corresponding to the target generative artificial intelligence application based on the graphics card resource priority and graphics card resource availability, wherein the graphics card among the multiple graphics cards is used to provide resources for the operation of the target generative artificial intelligence application; and a processing unit 504 runs the target generative artificial intelligence application on the target graphics card to process the data information to be processed and obtain the processing result. This solves the technical problem in the related art where multiple generative artificial intelligence applications request the same type of graphics card resources for data processing, resulting in the generative artificial intelligence application being unable to run due to resource limitations. In this solution, upon receiving a target data processing request that needs to be processed by the target generative AI application, the solution determines which graphics card from among multiple graphics cards corresponding to the target generative AI application will provide resources for its operation based on the graphics card resource priority. This avoids the problem of the generative AI application failing to run due to resource limitations. Through the above steps, graphics card resources can be managed and allocated more effectively. Resource configuration can be flexibly adjusted according to the needs of the generative AI application and the priority of graphics card resources to ensure the continuity and performance of the generative AI application, thereby achieving the technical effect of improving the operational stability of the generative AI application.

[0099] Optionally, in the data processing apparatus provided in Embodiment 3 of this application, the second determining unit 503 includes: a first determining subunit, configured to determine a plurality of first graphics cards with a first graphics card resource priority from a plurality of graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority; a detecting subunit, configured to detect whether there is a first graphics card among the plurality of first graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application; a second determining subunit, configured to determine a plurality of second graphics cards with a second graphics card resource priority from a plurality of graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority if there is no first graphics card among the plurality of first graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, wherein the resource priority of the second graphics cards is lower than the resource priority of the first graphics cards; and a third determining subunit, configured to determine the target graphics card based on the candidate graphics card if there is a candidate graphics card among the plurality of second graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application.

[0100] Optionally, in the data processing apparatus provided in Embodiment 3 of this application, the apparatus further includes: a third determining unit, configured to determine the amount of unused resources of the first graphics cards among the plurality of first graphics cards based on the data processing requests currently being processed by the first graphics cards among the plurality of first graphics cards before detecting whether there is a first graphics card among the plurality of first graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application; a prediction unit, configured to predict the amount of data to be processed by the first graphics cards among the plurality of first graphics cards at the next moment based on the data processing requests currently being processed by the first graphics cards among the plurality of first graphics cards, and obtain a predicted resource amount; and a fourth determining unit, configured to determine the remaining resource amount of the first graphics cards among the plurality of first graphics cards based on the unused resource amount and the predicted resource amount.

[0101] Optionally, in the data processing apparatus provided in Embodiment 3 of this application, the apparatus further includes: an execution unit, configured to repeatedly execute the step of determining a third graphics card with a third graphics card resource priority from the plurality of graphics cards corresponding to the target generative artificial intelligence application based on the graphics card resource priority if there is no candidate graphics card among the plurality of second graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, until there is a candidate graphics card among the plurality of graphics cards with the current graphics card resource priority with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application.

[0102] Optionally, in the data processing apparatus provided in Embodiment 3 of this application, the third determining subunit includes: an acquisition module, configured to acquire the amount of resources required by the target generative artificial intelligence application and the initial resource utilization rate of the candidate graphics cards if there are multiple candidate graphics cards; an estimation module, configured to estimate the target resource utilization rate of the candidate graphics cards based on the amount of resources required by the target generative artificial intelligence application and the initial resource utilization rate; and a determining module, configured to determine the target graphics card from the candidate graphics cards based on the target resource utilization rate.

[0103] Optionally, in the data processing apparatus provided in Embodiment 3 of this application, the determining module includes: a determining submodule, used to determine the server where the candidate graphics card is located; a detecting submodule, used to detect whether there is a code package corresponding to the target generative artificial intelligence application cached in the target server in the server where the candidate graphics card is located; and a determining submodule, used to determine the target graphics card from the candidate graphics cards based on the target resource utilization rate if there is no code package corresponding to the target generative artificial intelligence application cached in the target server.

[0104] Optionally, in the data processing apparatus provided in Embodiment 3 of this application, the processing unit 504 includes: an acquisition subunit, used to acquire parameter information corresponding to the target graphics card; an adjustment subunit, used to adjust the target configuration information in the configuration file corresponding to the target generative artificial intelligence application according to the parameter information corresponding to the target graphics card, to obtain an adjusted configuration file, wherein the target configuration information is configuration information related to the parameter information corresponding to the target graphics card; and a running subunit, used to run the target generative artificial intelligence application and the adjusted configuration file on the target graphics card, so as to process the data information to be processed through the target generative artificial intelligence application and obtain a processing result.

[0105] It should be noted that the acquisition unit 501, the first determination unit 502, the second determination unit 503, and the processing unit 504 mentioned above correspond to steps S201 to S204 in Embodiment 1. The four units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.

[0106] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0107] Example 4

[0108] According to embodiments of this application, a data processing system for implementing the above-described data processing method is also provided, such as... Figure 6 As shown, the system includes: a resource configuration manager 60, a resource scheduler 61, a status detector 62, and an application adapter 63.

[0109] Resource Configuration Manager 60: Responsible for receiving and storing developers' definitions of resource configurations such as graphics card model, number of GPU cores, memory size, and number of vCPUs, as well as graphics card resource priority settings.

[0110] Resource Scheduler 61: Dynamically schedules graphics card resources based on the current cloud platform's resource availability and the application's configured graphics card resource priority. When the preferred resource becomes unavailable, it automatically switches to the next configured graphics card resource priority.

[0111] Status Detector 62: Real-time detection of graphics card resource usage status and application requirements, providing the resource scheduler with accurate resource usage and demand prediction.

[0112] Application Adapter 63: Ensures a smooth transition of applications to new resource configurations, handles dependencies and configuration adjustments related to specific graphics card models, and guarantees application performance.

[0113] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0114] Example 5

[0115] Embodiments of this application may provide a computer terminal, which may be any computer terminal device in a group of computer terminals. Optionally, in this embodiment, the aforementioned computer terminal may also be replaced by a mobile terminal or other terminal device.

[0116] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.

[0117] In this embodiment, the computer terminal described above can execute program code for the following steps in the data processing method: obtaining a target data processing request, wherein the target data processing request carries at least data information to be processed; determining the target generative artificial intelligence application for the target data processing request; determining a target graphics card from multiple graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority and graphics card resource availability, wherein the graphics cards among the multiple graphics cards are used to provide resources for the operation of the target generative artificial intelligence application; running the target generative artificial intelligence application on the target graphics card to process the data information to be processed and obtain a processing result.

[0118] The aforementioned computer terminal can execute program code for the following steps in the data processing method: determining the target graphics card from multiple graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority and graphics card resource availability includes: determining multiple first graphics cards with a first graphics card resource priority from multiple graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority; detecting whether there is a first graphics card among the multiple first graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application; if there is no first graphics card among the multiple first graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, then determining multiple second graphics cards with a second graphics card resource priority from multiple graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority, wherein the resource priority of the second graphics cards is lower than the resource priority of the first graphics cards; if there are candidate graphics cards among the multiple second graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, then determining the target graphics card based on the candidate graphics cards.

[0119] The aforementioned computer terminal can execute program code for the following steps in the data processing method: before detecting whether there is a first graphics card among the multiple first graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, the method further includes: determining the amount of unused resources of the first graphics cards among the multiple first graphics cards based on the data processing requests currently being processed by the first graphics cards among the multiple first graphics cards; predicting the amount of data to be processed by the first graphics cards among the multiple first graphics cards at the next moment based on the data processing requests currently being processed by the first graphics cards among the multiple first graphics cards, and obtaining the predicted resource amount; determining the remaining resource amount of the first graphics cards among the multiple first graphics cards based on the unused resource amount and the predicted resource amount.

[0120] The aforementioned computer terminal can execute program code for the following steps in the data processing method: If there is no candidate graphics card among the plurality of second graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, the method further includes: repeatedly executing the step of determining the third graphics card with a third graphics card resource priority from the plurality of graphics cards corresponding to the target generative artificial intelligence application based on the graphics card resource priority, until there is a candidate graphics card among the plurality of graphics cards with the current graphics card resource priority with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application.

[0121] The aforementioned computer terminal can execute program code for the following steps in the data processing method: determining the target graphics card based on candidate graphics cards includes: if there are multiple candidate graphics cards, obtaining the amount of resources required for the target generative artificial intelligence application and the initial resource utilization rate of the candidate graphics cards; making an estimate based on the amount of resources required for the target generative artificial intelligence application and the initial resource utilization rate to obtain the target resource utilization rate of the candidate graphics cards; and determining the target graphics card from the candidate graphics cards based on the target resource utilization rate.

[0122] The aforementioned computer terminal can execute program code for the following steps in the data processing method: determining the target graphics card from candidate graphics cards based on the target resource utilization rate includes: determining the server where the candidate graphics card is located; detecting whether the target server has cached the code package corresponding to the target generative artificial intelligence application in the target server; if the target server does not have cached the code package corresponding to the target generative artificial intelligence application in the target server, then determining the target graphics card from the candidate graphics cards based on the target resource utilization rate.

[0123] The aforementioned computer terminal can execute program code for the following steps in the data processing method: running a target generative artificial intelligence application on the target graphics card to process the data information to be processed and obtain the processing result, including: obtaining parameter information corresponding to the target graphics card; adjusting the target configuration information in the configuration file corresponding to the target generative artificial intelligence application based on the parameter information corresponding to the target graphics card to obtain an adjusted configuration file, wherein the target configuration information is configuration information related to the parameter information corresponding to the target graphics card; running the target generative artificial intelligence application and the adjusted configuration file on the target graphics card to process the data information to be processed and obtain the processing result.

[0124] Optionally, Figure 7 This is a structural block diagram of a computer terminal according to an embodiment of this application. Figure 7 As shown, the computer terminal 10 may include: one or more ( Figure 7 (Only one is shown in the image) Processor 102 and memory 104. The computing terminal 10 may also include a storage controller to control and manage the memory 104; the computing terminal 10 may also include a peripheral interface to connect to a radio frequency module, an audio module, and a display screen, etc.

[0125] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the data processing method and apparatus in this embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the aforementioned data processing method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0126] The processor can invoke information and application programs stored in memory via a transmission device to perform the following steps: acquiring a target data processing request, wherein the target data processing request carries at least data information to be processed; determining the target generative artificial intelligence application for the target data processing request; determining the target graphics card from multiple graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority and availability, wherein the graphics card among the multiple graphics cards is used to provide resources for the operation of the target generative artificial intelligence application; and running the target generative artificial intelligence application on the target graphics card to process the data information to be processed and obtain the processing result.

[0127] Optionally, the processor may also execute program code that performs the following steps: determining the target graphics card from multiple graphics cards corresponding to the target generative AI application based on graphics card resource priority and graphics card resource availability includes: determining multiple first graphics cards with a first graphics card resource priority from multiple graphics cards corresponding to the target generative AI application based on graphics card resource priority; detecting whether there is a first graphics card among the multiple first graphics cards with a remaining resource amount greater than the resource amount required by the target generative AI application; if there is no first graphics card among the multiple first graphics cards with a remaining resource amount greater than the resource amount required by the target generative AI application, then determining multiple second graphics cards with a second graphics card resource priority from multiple graphics cards corresponding to the target generative AI application based on graphics card resource priority, wherein the resource priority of the second graphics cards is lower than the resource priority of the first graphics cards; if there are candidate graphics cards among the multiple second graphics cards with a remaining resource amount greater than the resource amount required by the target generative AI application, then determining the target graphics card based on the candidate graphics cards.

[0128] Optionally, the processor may also execute program code for the following steps: before detecting whether there is a first graphics card among the plurality of first graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, the method further includes: determining the amount of unused resources of the first graphics cards among the plurality of first graphics cards based on the data processing requests currently being processed by the first graphics cards among the plurality of first graphics cards; predicting the amount of data to be processed by the first graphics cards among the plurality of first graphics cards at the next moment based on the data processing requests currently being processed by the first graphics cards among the plurality of first graphics cards, and obtaining a predicted resource amount; determining the remaining resource amount of the first graphics cards among the plurality of first graphics cards based on the unused resource amount and the predicted resource amount.

[0129] Optionally, the processor may also execute program code for the following steps: if there is no candidate graphics card among the plurality of second graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, the method further includes: repeatedly executing the step of determining the third graphics card with a third graphics card resource priority from the plurality of graphics cards corresponding to the target generative artificial intelligence application based on the graphics card resource priority, until there is a candidate graphics card among the plurality of graphics cards with the current graphics card resource priority with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application.

[0130] Optionally, the processor may also execute program code for the following steps: determining the target graphics card based on candidate graphics cards includes: if there are multiple candidate graphics cards, obtaining the amount of resources required for the target generative artificial intelligence application and the initial resource utilization rate of the candidate graphics cards; estimating the target resource utilization rate of the candidate graphics cards based on the amount of resources required for the target generative artificial intelligence application and the initial resource utilization rate; and determining the target graphics card from the candidate graphics cards based on the target resource utilization rate.

[0131] Optionally, the processor may also execute program code that performs the following steps: determining the target graphics card from candidate graphics cards based on the target resource utilization rate, including: determining the server where the candidate graphics card is located; detecting whether the server where the candidate graphics card is located has a cached code package corresponding to the target generative artificial intelligence application in the target server; if the server does not have a cached code package corresponding to the target generative artificial intelligence application in the target server, then determining the target graphics card from candidate graphics cards based on the target resource utilization rate.

[0132] Optionally, the processor may also execute program code that performs the following steps: running a target generative AI application on the target graphics card to process the data information to be processed and obtain the processing result, including: obtaining parameter information corresponding to the target graphics card; adjusting the target configuration information in the configuration file corresponding to the target generative AI application based on the parameter information corresponding to the target graphics card to obtain an adjusted configuration file, wherein the target configuration information is configuration information related to the parameter information corresponding to the target graphics card; running the target generative AI application and the adjusted configuration file on the target graphics card to process the data information to be processed and obtain the processing result.

[0133] Those skilled in the art will understand that Figure 7 The structure shown is for illustrative purposes only. The computer terminal can also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, and other terminal devices. Figure 7 This does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include components that are more advanced than those described above. Figure 7 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 7 The different configurations shown.

[0134] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0135] Example 6

[0136] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the data processing method provided in Embodiment 1.

[0137] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.

[0138] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining a target data processing request, wherein the target data processing request carries at least data information to be processed; determining the target generative artificial intelligence application for the target data processing request; determining a target graphics card from multiple graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority and graphics card resource availability, wherein the graphics card among the multiple graphics cards is used to provide resources for the operation of the target generative artificial intelligence application; and running the target generative artificial intelligence application on the target graphics card to process the data information to be processed and obtain a processing result.

[0139] The aforementioned storage medium is configured to store program code for performing the following steps: determining a target graphics card from multiple graphics cards corresponding to the target generative AI application based on graphics card resource priority and availability includes: determining multiple first graphics cards with a first graphics card resource priority from multiple graphics cards corresponding to the target generative AI application based on graphics card resource priority; detecting whether there is a first graphics card among the multiple first graphics cards with a remaining resource amount greater than the resource amount required by the target generative AI application; if there is no first graphics card among the multiple first graphics cards with a remaining resource amount greater than the resource amount required by the target generative AI application, then determining multiple second graphics cards with a second graphics card resource priority from multiple graphics cards corresponding to the target generative AI application based on graphics card resource priority, wherein the resource priority of the second graphics cards is lower than the resource priority of the first graphics cards; if there are candidate graphics cards among the multiple second graphics cards with a remaining resource amount greater than the resource amount required by the target generative AI application, then determining the target graphics card based on the candidate graphics cards.

[0140] The aforementioned storage medium is configured to store program code for performing the following steps: before detecting whether there is a first graphics card among the plurality of first graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, the method further includes: determining the amount of unused resources of the first graphics cards among the plurality of first graphics cards based on the data processing requests currently being processed by the first graphics cards among the plurality of first graphics cards; predicting the amount of data to be processed by the first graphics cards among the plurality of first graphics cards at the next moment based on the data processing requests currently being processed by the first graphics cards among the plurality of first graphics cards, to obtain a predicted resource amount; and determining the remaining resource amount of the first graphics cards among the plurality of first graphics cards based on the unused resource amount and the predicted resource amount.

[0141] The aforementioned storage medium is configured to store program code for performing the following steps: if there is no candidate graphics card among the plurality of second graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, the method further includes: repeatedly executing the step of determining the third graphics card with a third graphics card resource priority from the plurality of graphics cards corresponding to the target generative artificial intelligence application based on the graphics card resource priority, until there is a candidate graphics card among the plurality of graphics cards with the current graphics card resource priority with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application.

[0142] The aforementioned storage medium is configured to store program code for performing the following steps: determining the target graphics card based on candidate graphics cards includes: if there are multiple candidate graphics cards, obtaining the amount of resources required by the target generative artificial intelligence application and the initial resource utilization rate of the candidate graphics cards; estimating the target resource utilization rate of the candidate graphics cards based on the amount of resources required by the target generative artificial intelligence application and the initial resource utilization rate; and determining the target graphics card from the candidate graphics cards based on the target resource utilization rate.

[0143] The aforementioned storage medium is configured to store program code for performing the following steps: determining the target graphics card from candidate graphics cards based on the target resource utilization rate includes: determining the server where the candidate graphics card is located; detecting whether the target server has cached the code package corresponding to the target generative artificial intelligence application in the target server; if the target server does not have cached the code package corresponding to the target generative artificial intelligence application in the target server, then determining the target graphics card from the candidate graphics cards based on the target resource utilization rate.

[0144] The aforementioned storage medium is configured to store program code for performing the following steps: running a target generative AI application on the target graphics card to process data information to be processed and obtain processing results, including: obtaining parameter information corresponding to the target graphics card; adjusting the target configuration information in the configuration file corresponding to the target generative AI application based on the parameter information corresponding to the target graphics card to obtain an adjusted configuration file, wherein the target configuration information is configuration information related to the parameter information corresponding to the target graphics card; running the target generative AI application and the adjusted configuration file on the target graphics card to process the data information to be processed and obtain processing results.

[0145] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0146] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0147] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0148] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0149] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0150] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0151] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A data processing method, characterized in that, include: Obtain a target data processing request, wherein the target data processing request carries at least data information to be processed; Identify the target generative artificial intelligence application for which the target data processing request needs to be processed; The target graphics card is determined from multiple graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority and graphics card resource availability, wherein the graphics card among the multiple graphics cards is used to provide resources for the operation of the target generative artificial intelligence application; A target generative artificial intelligence application is run on the target graphics card to process the data information to be processed and obtain the processing result.

2. The method according to claim 1, characterized in that, The target graphics card is determined from multiple graphics cards corresponding to the target generative artificial intelligence application based on graphics card resource priority and availability, including: Based on the graphics card resource priority, a plurality of first graphics cards with a first graphics card resource priority are determined from a plurality of graphics cards corresponding to the target generative artificial intelligence application; Detect whether there is a first graphics card among the plurality of first graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application; If none of the plurality of first graphics cards has a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, then a plurality of second graphics cards with second graphics card resource priority are determined from the plurality of graphics cards corresponding to the target generative artificial intelligence application according to the graphics card resource priority, wherein the resource priority of the second graphics card is lower than the resource priority of the first graphics card. If among the plurality of second graphics cards there is a candidate graphics card with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, then the target graphics card is determined based on the candidate graphics card.

3. The method according to claim 2, characterized in that, Before detecting whether there is a first graphics card among the plurality of first graphics cards with a remaining resource amount greater than the resource amount required by the target generative artificial intelligence application, the method further includes: Based on the data processing request currently being processed by the first graphics card among the plurality of first graphics cards, determine the amount of unused resources of the first graphics card among the plurality of first graphics cards; Based on the current data processing request of the first graphics card among the plurality of first graphics cards, the amount of data to be processed by the first graphics card among the plurality of first graphics cards at the next moment is predicted to obtain the predicted resource amount; Based on the amount of unoccupied resources and the predicted amount of resources, the remaining resources of the first graphics card among the plurality of first graphics cards are determined.

4. The method according to claim 2, characterized in that, If none of the plurality of second graphics cards has a remaining resource quantity greater than that required by the target generative artificial intelligence application, the method further includes: The step of determining the third graphics card with the third graphics card resource priority from multiple graphics cards corresponding to the target generative artificial intelligence application based on the graphics card resource priority is repeated until there is a candidate graphics card among the multiple graphics cards with the current graphics card resource priority whose remaining resources are greater than the resources required by the target generative artificial intelligence application.

5. The method according to claim 3, characterized in that, Determining the target graphics card based on the candidate graphics cards includes: If there are multiple candidate graphics cards, then obtain the amount of resources required for the target generative artificial intelligence application and the initial resource utilization rate of the candidate graphics cards; Based on the resource requirements and initial resource utilization rate of the target generative artificial intelligence application, the target resource utilization rate of the candidate graphics card is estimated. The target graphics card is determined from the candidate graphics cards based on the target resource utilization rate.

6. The method according to claim 5, characterized in that, Determining the target graphics card from the candidate graphics cards based on the target resource utilization includes: Determine the server where the candidate graphics card is located; Detect whether the server where the candidate graphics card is located contains the code package corresponding to the target generative artificial intelligence application that has been cached in the target server; If the target server does not have a cached code package corresponding to the target generative AI application, then the target graphics card is determined from the candidate graphics cards based on the target resource utilization rate.

7. The method according to claim 1, characterized in that, A target generative artificial intelligence application is run on the target graphics card to process the data information to be processed, and the processing results include: Obtain the parameter information corresponding to the target graphics card; The target configuration information in the configuration file corresponding to the target generative artificial intelligence application is adjusted according to the parameter information corresponding to the target graphics card to obtain the adjusted configuration file. The target configuration information is configuration information related to the parameter information corresponding to the target graphics card. The target generative AI application and the adjusted configuration file are run on the target graphics card to process the data information to be processed and obtain the processing result.

8. A data processing method, characterized in that, include: Obtain the target data processing request sent by the client, wherein the target data processing request carries at least the data information to be processed; In a cloud server, a target generative artificial intelligence application is identified to process the target data processing request. Based on the graphics card resource priority and availability, a target graphics card is determined from multiple graphics cards corresponding to the target generative artificial intelligence application. The graphics card among the multiple graphics cards is used to provide resources for the operation of the target generative artificial intelligence application. The identification information corresponding to the target graphics card is returned to the client. A target generative artificial intelligence application is run on the target graphics card to process the data information to be processed and obtain the processing result.

9. A data processing apparatus, characterized in that, include: An acquisition unit is used to acquire a target data processing request, wherein the target data processing request carries at least data information to be processed; The first determining unit is used to determine the target generative artificial intelligence application for the target data processing request to be processed. The second determining unit is used to determine the target graphics card from multiple graphics cards corresponding to the target generative artificial intelligence application based on the graphics card resource priority and graphics card resource availability, wherein the graphics card among the multiple graphics cards is used to provide resources for the operation of the target generative artificial intelligence application; The processing unit is configured to run a target generative artificial intelligence application on the target graphics card to process the data information to be processed and obtain a processing result.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the storage medium is located to perform the data processing method according to any one of claims 1 to 8.

11. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the data processing method according to any one of claims 1 to 8.

12. A computer program product, characterized in that, The computer program product includes a stored computer program that, when executed by a processor, implements the data processing method according to any one of claims 1 to 8.