Serverless computing-based data processing methods and electronic device
By optimizing the execution order based on the performance indicators of the objective function in the serverless architecture, the problem of data transmission time consumed in the Serverless GPU video memory replacement scenario is solved, and more efficient data transmission is achieved.
Patent Information
- Application Number
- PCT/CN2024/122607
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-14
- Filing Date
- 2024-09-30
- Publication Date
- 2025-05-22
AI Technical Summary
In a serverless architecture, in the Serverless GPU video memory replacement scenario, there is a problem of time-consuming data transmission, and the existing technology has not yet proposed an effective solution.
By acquiring multiple function requests from the host, determining the objective function, and determining its execution order based on the performance metrics of the objective function, the data transmission between the host host memory and the image processor video memory is optimized.
This method effectively reduces the time-consuming process in data transmission and solves the technical problem of time-consuming process in data transmission.
Smart Images

Figure CN2024122607_22052025_PF_FP_ABST
Abstract
Description
Data processing method and electronic device based on server-less computing Technical Field
[0001] The present disclosure relates to the field of function computing in serverless computing, and in particular to a data processing method and electronic device based on server-less computing. Background Art
[0002] At present, the serverless graphics processing unit (Serverless GPU) is a new type of emerging cloud computing GPU service with the advantages of free operation and maintenance and pay-as-you-go.
[0003] In related technologies, in the Serverless GPU memory replacement scenario, data can be copied between the GPU memory and the host main memory. However, a GPU can only process one request function. When there are too many request functions, the request functions to be executed will be queued, resulting in a technical problem of long data transmission time.
[0004] To address the above-mentioned problems, no effective solutions have been proposed so far.
[0005] Summary of the Invention
[0006] The embodiments of the present disclosure provide a data processing method and electronic device based on server-unaware computing, so as to at least solve the technical problem of long time consumption during data transmission.
[0007] According to one aspect of an embodiment of the present disclosure, a data processing method based on server-aware computing is provided. The method may include: obtaining multiple function requests from a host computer; determining a target function that any function request requests an image processor to execute, thereby obtaining multiple target functions; and determining an execution order for the multiple target functions based on performance indicators of the multiple target functions, wherein the performance indicators of the target functions are used to characterize the performance of the target functions during execution, and the execution order is used to determine the data transferred between the host computer's main memory and the image processor's video memory.
[0008] According to another aspect of an embodiment of the present disclosure, another data processing method based on server-insensitive computing is provided. The method may include: obtaining multiple data processing requests from a host computer; determining a target function for each data processing request to be executed by an image processor, thereby obtaining multiple target functions, wherein the target function is at least used for data copying or data synchronization; and determining an execution order for the multiple target functions based on performance indicators of the multiple target functions, wherein the performance indicators of the target functions are used to characterize the performance of the target functions during execution, and the execution order is used to determine the data to be copied or synchronized between the host computer's main memory and the image processor's video memory.
[0009] According to another aspect of an embodiment of the present disclosure, an electronic device is also provided. The electronic device may include a memory and a processor: the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions. When the above-mentioned computer-executable instructions are executed by the processor, any one of the above-mentioned data processing methods based on server-imperceptible computing is implemented.
[0010] According to another aspect of an embodiment of the present disclosure, a processor is further provided, which is used to run a program, wherein any one of the above-mentioned data processing methods based on server-unaware computing is executed when the program is running.
[0011] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is also provided, which includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute any of the above-mentioned data processing methods based on server-unaware computing.
[0012] In an embodiment of the present disclosure, multiple function requests are obtained from a host machine; a target function requested by any one of the function requests is determined to be executed by the image processor, thereby obtaining multiple target functions; and an execution order of the multiple target functions is determined based on performance indicators of the multiple target functions, wherein the performance indicators of the target functions are used to characterize the performance of the target functions during execution, and the execution order is used to determine the data to be transferred between the host machine's main memory and the image processor's video memory. That is, in an embodiment of the present disclosure, multiple function requests are obtained from the host machine, the target functions requested by the function requests are determined, and based on the performance indicators of the target functions, the execution order of the multiple target functions is determined, thereby achieving the technical effect of reducing the time consumed during data transmission and resolving the technical problem of long data transmission times.
[0013] It is easy to note that the above general description and the following detailed description are only for the purpose of exemplifying and explaining the present disclosure, and do not constitute a limitation of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of the present disclosure. The exemplary embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:
[0015] FIG1 is a schematic diagram of an application scenario of a data processing method based on server-unaware computing according to an embodiment of the present disclosure;
[0016] FIG2 is a structural block diagram of a computing environment of a data processing method based on server-unaware computing according to an embodiment of the present disclosure;
[0017] FIG3 is a flow chart of a data processing method based on server-unaware computing according to an embodiment of the present disclosure;
[0018] FIG4 is a flow chart of another data processing method based on server-unaware computing according to an embodiment of the present disclosure;
[0019] FIG5 is a schematic diagram of a function computing service architecture according to an embodiment of the present disclosure;
[0020] FIG6 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a data processing method based on server-unaware computing according to an embodiment of the present disclosure;
[0021] FIG7 is a structural block diagram of a service grid according to an embodiment of the present disclosure;
[0022] FIG8 is a schematic diagram of a data processing device based on server-unaware computing according to an embodiment of the present disclosure;
[0023] FIG9 is a schematic diagram of another data processing device based on server-unaware computing according to an embodiment of the present disclosure;
[0024] FIG10 is a structural block diagram of a computer terminal according to an embodiment of the present disclosure;
[0025] FIG11 is a block diagram of an electronic device according to an embodiment of the present disclosure, which is a data processing method based on server-unaware computing. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present disclosure.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0028] First, some nouns or terms that appear in the description of the embodiments of the present disclosure are subject to the following explanations:
[0029] Function as a Service (FaaS) is a cloud computing service that allows developers to build, compute, run, and manage application packages in the form of functions without having to maintain the underlying infrastructure.
[0030] Function Compute is a form of serverless architecture and is set up as a fully managed event-driven computing service.
[0031] Host Virtual Machine (Host VM) is a machine used to run virtual machines;
[0032] A graphics processing unit (GPU) is a processor composed of many smaller, more specialized cores. It is a computing engine. A motherboard expansion card with a graphics processor as its core is also called a graphics card.
[0033] Video memory swap refers to the technology of copying data between GPU video memory and host main memory.
[0034] Example 1
[0035] According to a method according to an embodiment of the present disclosure, a data processing method based on server-insensitive computing is provided. As an optional implementation, the data processing method based on server-insensitive computing may include, but is not limited to, application scenarios as shown in Figure 1. Figure 1 is a schematic diagram of an application scenario of a data processing method based on server-insensitive computing according to an embodiment of the present disclosure. As shown in Figure 1, in the application scenario, a terminal device 102 may, but is not limited to, communicate with a server 106 via a network 104. The server 106 may, but is not limited to, perform operations on a database 108, such as writing or reading data. The terminal device 102 may, but is not limited to, include a human-computer interaction screen, a processor, and memory. The human-computer interaction screen may, but is not limited to, be used to display a virtual machine on the terminal device 102. The processor may, but is not limited to, responding to the human-computer interaction operation, executing the corresponding operation, or generating the corresponding instruction and sending the generated instruction to the server 106. The memory is used to store relevant processing data, such as function requests and data to be transmitted.
[0036] As an optional method, the following steps in the data processing method based on server-imperceptible computing can be executed on the server 106: Step S102, obtaining multiple function requests from the host machine; Step S104, determining the target function that any function request requests the image processor to execute, and obtaining multiple target functions; Step S106, determining the execution order of the multiple target functions based on the performance indicators of the multiple target functions.
[0037] Using the above method, multiple function requests are obtained from a host machine; the target function requested by any function request is determined to be executed by the image processor, thereby obtaining multiple target functions; and the execution order of the multiple target functions is determined based on the performance indicators of the multiple target functions, wherein the performance indicators of the target functions are used to characterize the performance of the target functions during execution, and the execution order is used to determine the data transferred between the host machine's main memory and the image processor's video memory. That is, in the disclosed embodiment, multiple function requests are obtained from the host machine, the target functions requested by the function requests are determined, and the execution order of the multiple target functions is determined based on the performance indicators of the target functions, thereby achieving the technical effect of reducing the time consumed during data transmission and resolving the technical problem of long data transmission times.
[0038] The application scenario diagram shown in FIG1 can serve not only as an exemplary block diagram of a computer terminal (or mobile device), but also as an exemplary block diagram of the server 106 described above. In an optional embodiment, FIG2 shows a block diagram of an embodiment using the server 106 shown in FIG1 as a computing node in a computing environment 201. FIG2 is a structural block diagram of a computing environment for a data processing method based on server-insensitive computing according to an embodiment of the present disclosure. As shown in FIG2, the computing environment 201 includes multiple computing nodes (such as servers) (illustrated as 210-1, 210-2, ...) running on a distributed network. The computing nodes all include local processing and memory resources, and the end user 202 can remotely run applications or store data in the computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 in the computing environment 201, representing services "A," "D," "E," and "H," respectively.
[0039] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or request of end user 202 can be provided to the ingress gateway 230. The ingress gateway 230 may include a corresponding agent to handle the provisioning and / or request for services (one or more services provided in the computing environment 201).
[0040] Services are provided or deployed based on various virtualization technologies supported by the computing environment 201. In some embodiments, services can be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. Virtual machine-based virtualization can be to simulate a real computer by initializing a virtual machine, executing programs and applications without directly contacting any actual hardware resources. While the virtual machine virtualizes the machine, according to container-based virtualization, a container can be started to virtualize the entire operating system (OS) so that multiple workloads can run on a single operating system instance.
[0041] In one embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). For example, as shown in Figure 2, service 220-2 can be equipped with one or more Pods 240-1, 240-2, ..., 240-N (collectively referred to as Pods). A Pod can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively referred to as containers). One or more containers in a Pod process requests related to one or more corresponding functions of the service, and the proxy 245 typically controls network functions related to the service, such as routing, load balancing, etc. Other services can also be equipped with Pods similar to Pods.
[0042] During operation, executing a user request from end user 202 may require invoking one or more services in computing environment 201. Executing one or more functions of one service may require invoking one or more functions of another service. As shown in FIG2 , service "A" 220-1 receives a user request from end user 202 from ingress gateway 230. Service "A" 220-1 may invoke service "D" 220-2, and service "D" 220-2 may request service "E" 220-3 to execute one or more functions.
[0043] This computing environment can be a cloud computing environment, where resource allocation is managed by the cloud service provider, allowing for feature development without having to worry about implementing, adjusting, or scaling servers. This computing environment allows developers to execute code in response to events without building or maintaining complex infrastructure. Services can be partitioned to perform a set of functions that can scale independently and automatically, rather than scaling a single hardware device to handle the potential load.
[0044] In the above operating environment, the present disclosure provides a data processing method based on server-insensitive computing as shown in Figure 3. It should be noted that the data processing method based on server-insensitive computing of this embodiment can be executed by the scheduling device in the server of the embodiment shown in Figure 1, wherein the scheduling device can be a scheduler. Figure 3 is a flow chart of a data processing method based on server-insensitive computing according to an embodiment of the present disclosure. As shown in Figure 3, the method may include the following steps:
[0045] Step S302: Acquire multiple function requests from the host machine.
[0046] In the technical solution provided in step S302 above, multiple function requests can be obtained from the host machine. The function request can be a request for calling a function, and can be a request triggered by a Hypertext Transfer Protocol (HTTP) request, a message queue, an event trigger, a timer trigger, or the like. The host machine can be a tenant using a virtual machine.
[0047] Optionally, the host machine may send the function request to the function computing platform through a hypertext transfer protocol request, a message queue, an event trigger, a timer trigger, etc. The function computing platform receives multiple function requests from the host machine.
[0048] For example, when the host machine wants to synchronize or copy data, it can send a function request to run a function, and the cloud host can receive the function request sent by the host machine.
[0049] Step S304: determining a target function that any function request requests the image processor to execute, and obtaining a plurality of target functions.
[0050] In the technical solution provided in step S304 of the present disclosure, the obtained function request is parsed to determine the target function requested by the function request. The target function can be a function in function computing, a simple function, or a relatively complete function, such as a login function, or a simple module formed by combining several functions, such as a login or registration module, or a complete service. It should be noted that the function in function computing is not a function in the ordinary sense, but a relatively broad concept, and no specific restrictions are placed on the content and type of the function.
[0051] Optionally, when the host machine sends a large number of function requests to the computing node of the same function computing platform (also known as function computing service), the computing node obtains multiple function requests and determines the target function that any function request requests the image processor to execute, and obtains multiple target functions.
[0052] For example, this method can be applied to a function computing platform scenario. In this scenario, the host machine can be the host computer running the function computing platform. The host machine sends a function request to the function computing platform. After receiving the function request, the function computing platform parses the function request to determine the request parameters, loads the target function requested by the function request, and creates an execution environment for the target function, where the execution environment can be a graphics processor. In the function computing platform, each function has corresponding code. Based on the function request, the function computing platform calls the corresponding target function, such as the function code, and passes the request parameters to the target function.
[0053] Step S306, based on the performance indicators of the multiple objective functions, determining the execution order of the multiple objective functions, wherein the performance indicators of the objective functions are used to characterize the performance of the objective functions during the execution process, and the execution order is used to determine the data transmitted between the host machine's main memory and the graphics memory of the image processor.
[0054] In the optional technical solution provided in the above step S306 of the present disclosure, after obtaining multiple objective functions, the performance indicators of all objective functions can be determined, and based on the performance indicators of the objective functions, the execution order of the multiple objective functions can be determined. The performance indicator indication can include a service level objective, wherein the service level objective involved in the above optional embodiment can be a service level objective (Service Level Objective, referred to as SLO). It should be noted that since the service level objective can be characterized by a target value or range value of the service level, at the same time, in an optional example, the service level objective can be composed of one or more service level indicators (Service Level Indicator, referred to as SLI), therefore, the service level objective can also be called a service level objective. During implementation, it can be used to characterize key performance indicators such as availability, response time, and error rate of the objective function.
[0055] When determining the execution order of multiple target functions based on target function performance metrics, a single GPU device can be executing a single function request at any given moment. This single function request can occupy nearly the entire GPU's memory resources, and other function requests routed to the same GPU device must queue. Therefore, an intelligent scheduling solution is needed to reduce the end-to-end latency caused by function request queuing.
[0056] In this embodiment, in order to avoid the end-to-end time consumption caused by queuing function requests, when a large number of function requests are sent to the computing node of the same function computing service, the target functions requested by multiple function requests can be determined, thereby determining the performance indicators of the target functions. Therefore, the present disclosure can determine the execution order of multiple target functions based on the respective performance indicators of multiple target functions, thereby achieving the purpose of satisfying the performance indicators of each function as much as possible through intelligent scheduling and queuing strategies, thereby reducing queuing delays, achieving the technical effect of reducing the time consumption in the data transmission process, and solving the technical problem of long time consumption in the data transmission process.
[0057] In an optional technical solution of the present disclosure, since the performance indicators of different objective functions can be different, when multiple function requests are obtained, multiple objective functions can be obtained by determining the objective functions requested by the multiple function requests. Performance indicators of the multiple objective functions can then be further determined, such as the execution efficiency and execution speed of the objective functions. This is for illustrative purposes only and does not specifically limit the type of performance indicator. The execution order of the multiple objective functions is determined based on the performance indicators of the objective functions. The objective function can be a service.
[0058] For example, the performance indicators of different objective functions can be used in different application scenarios. When multiple function requests are obtained, multiple objective functions can be obtained by determining the objective functions requested by the multiple function requests, and then the performance indicators of the multiple objective functions can be determined. The performance indicators determined here may include: determining the response time required by the objective function. If the required response time is longer, it means that the objective function can be responded to later, so that the evaluation result corresponding to the objective function can be determined, and then the execution order of the objective function can be determined to be relatively later. It should be noted that this is only an example, and there is no specific restriction on the type of performance indicators and the method of determining the execution order.
[0059] Optionally, embodiments of the present disclosure may also implement sequential execution of multiple target functions according to their execution order. The multiple target functions executed sequentially may be used to schedule data transfer between the host computer's main memory and the video memory used for image processing. The video memory configured for image processing may be the video memory of an image processor.
[0060] In this embodiment, when multiple function requests are obtained on a computing node, since the number of image processors on the computing node is limited, all function requests cannot be executed at the same time, which will cause the function requests to be executed to be queued. When a large number of function requests to be executed are queued, the end-to-end performance for the user will be greatly extended. To solve the above problem, in this embodiment, when the function requests on the current computing node are queued, the execution order of the requested target functions is intelligently adjusted and distributed based on the performance indicators of the target functions requested by the function requests. Based on the adjusted execution order, multiple target functions are executed in sequence, so that the execution results of the target functions meet the performance indicators of the target functions, thereby achieving the purpose of reducing the queuing time of function requests and improving the end-to-end performance of the user.
[0061] Optionally, based on the performance indicators of multiple objective functions, the execution order of the multiple objective functions can be determined, and the objective functions can be executed sequentially in the execution order. The executed objective function can further control the data transfer process between the main memory of the host machine and the video memory used for image processing. Among them, the video memory set for image processing can be the video memory of the image processor. For example, if the objective function is used to copy data in the main memory to the video memory, the objective function is executed to transfer the data in the main memory to the video memory; if the objective function is used to move data in the video memory to the main memory, the objective function can be executed to transfer the data in the video memory to the main memory.
[0062] For example, the function computing platform may include multiple computing nodes, and the computing nodes may include workers and schedulers for image processors. One worker corresponds to a real GPU device, and the GPU device can be time-shared and multiplexed. At the same time, a GPU device can only serve one function request. Under the above conditions, when the function computing platform obtains multiple function requests from the host machine, it determines the multiple target functions corresponding to the multiple function requests and obtains target functions within the range of cluster numbers. In order to avoid the waiting execution time of the target function, the execution order of the multiple target functions can be determined based on the performance indicators of the target function. After the execution order of the multiple target functions is determined, the scheduler in the computing node can be used to schedule the multiple target functions according to the execution order, and the image processor can execute the target function to control the data transmission process between the host machine's main memory and the image processor's video memory.
[0063] It should be noted that the execution object of this method can be a function computing platform, or a computing node, scheduler, cloud host, etc. in the function computing platform. No specific restrictions are placed on the execution subject here.
[0064] Through steps S302 to S308 of the present disclosure, multiple function requests are obtained from the host machine; the target function that any function request requests the image processor to execute is determined, thereby obtaining multiple target functions; and the execution order of the multiple target functions is determined based on the performance indicators of the multiple target functions, wherein the performance indicators of the target functions are used to characterize the performance of the target functions during execution, and the execution order is used to determine the data transmitted between the host machine's main memory and the image processor's video memory. That is, in an embodiment of the present disclosure, multiple function requests are obtained from the host machine, the target functions requested by the function requests are determined, and the execution order of the multiple target functions is determined based on the performance indicators of the target functions, thereby achieving the technical effect of reducing the time consumed during data transmission and resolving the technical problem of long data transmission time.
[0065] The above method of this embodiment is further introduced below.
[0066] As an optional implementation, step S306 determines the execution order of multiple objective functions based on the performance indicators of the multiple objective functions, including: determining the evaluation results of the objective functions based on the performance indicators of the objective functions; and determining the execution order based on the evaluation results.
[0067] In this embodiment, after obtaining multiple function requests from the host machine, the target function requested by the function request and the performance index of the target function are determined. Based on the performance index of the target function, the evaluation result of the target function can be determined, and based on the evaluation result, the execution order of the target function can be determined. Among them, the evaluation result can be used to characterize the priority of the execution of the target function or the importance level of the execution, and can be a measurement indicator. For example, the evaluation result can be a calculated score. The larger the score, the higher the priority of the target function. It should be noted that this is only for example, and does not impose specific restrictions on the form and meaning of the target function.
[0068] Optionally, in order to meet the performance indicators of multiple objective functions, this embodiment defines a metric, which can be the required request count (RRC) corresponding to the objective function. This metric can be used to determine how many times the objective function must be executed to meet the performance indicator of the objective function. As can be seen from the above, in this embodiment, the execution priority of all objective functions to be executed on the computing node is intuitively divided by the metric to determine the execution order of the objective functions.
[0069] For example, assuming that in the current execution of the objective function, the first objective function processes 10 function requests per second, and the second objective function processes 20 function requests per second, and the performance indicator of the first objective function requires the first objective function to process 100 function requests per second, and the performance indicator of the second objective function requires the second objective function to process 80 function requests per second, therefore, it can be determined that the evaluation result of the first objective function is 90, and the evaluation result of the second objective function is 60. It should be noted that the above-mentioned method of determining the evaluation results is only for illustration, and the method of determining the evaluation results can be changed according to actual conditions, and no specific restrictions are made here.
[0070] As an optional implementation, based on the performance indicator of the objective function, the evaluation result of the objective function is determined, including: determining the first execution number of the objective function that has been executed within a historical time period, and the second execution number of the executed objective function whose execution results meet the performance indicator threshold; based on the first execution number, the second execution number and the performance indicator, determining the evaluation result of the objective function at the current moment.
[0071] In this embodiment, the process of determining the evaluation result of the objective function based on the performance indicator of the objective function may include determining a first number of completed executions of the objective function within a historical time period, and a second number of completed executions of the objective function whose execution results meet a performance indicator threshold. Based on the first number of executions, the second number of executions, and the performance indicator, the evaluation result of the objective function at the current moment may be determined. The historical time period may be a predetermined period of time before the current moment, or a historical time range. For example, the historical time period may be three hours or three days before the current moment. This is for illustrative purposes only, and no specific restrictions are placed on the time or length of the historical time period. The execution results may be used to characterize execution speed, latency, efficiency, and other factors. This is for illustrative purposes only, and no specific restrictions are placed on the type of execution results. The performance indicator threshold may be a predetermined value, such as a predetermined maximum execution speed value or a predetermined minimum latency value. This is for illustrative purposes only, and no specific restrictions are placed on the source of the performance indicator threshold.
[0072] Optionally, since functions in the function computing platform can be called repeatedly, the target function may have been executed before the current moment. A historical time period is set, and the number of target functions that have been executed within the historical time period is determined to obtain a first execution count. Among the target functions that have been executed, the number of execution results that meet the performance indicator threshold is determined to obtain a second execution count. Based on the first and second execution counts and the performance indicator, an evaluation result of the target function at the current moment is determined. Based on the evaluation result, the execution order of the target functions at the current moment can be determined. Based on the execution order, the target functions are executed sequentially at the current moment.
[0073] For example, in a technical solution for determining the evaluation result of an objective function at the current moment based on a first execution number, a second execution number, and a performance index, the following technical solutions can be used: Solution 1: By setting a ratio of the first execution number to the second execution number, the difference between this ratio and the performance index can be used as the evaluation result of the objective function; Solution 2: Alternatively, the weighted average of the first execution number, the second execution number, and the performance index can be set as the evaluation result. It should be noted that this is merely an example and does not impose a specific limitation on the method for determining the evaluation result. Any method for determining the evaluation result based on the first execution number, the second execution number, and the performance index should be within the scope of protection of this disclosure.
[0074] As another optional implementation, determining the evaluation result of the objective function at the current moment based on the first execution number, the second execution number and the performance indicator can also be achieved through the following optional embodiments: obtaining the service level target in the performance indicator, wherein the service level target is used to characterize the execution status of the objective function; determining the evaluation result based on the first execution number, the second execution number and the service level target.
[0075] In this embodiment, the service target level in the performance indicator can be obtained, and the evaluation result can be determined based on the first execution quantity, the second execution quantity, and the service level target, wherein the service level target can be used to characterize the execution status of the objective function.
[0076] The current function computing service utilizes a memory replacement method, which can record the metadata and data of the inference application on the GPU memory. However, in this method, only one inference application is allowed to be executed on a GPU card at the same time, which causes other inference applications to be queued. However, the queuing of function requests to be executed will greatly prolong the end-to-end performance for users, resulting in a technical problem of long time consumption during data transmission. In order to solve the above problem, in this embodiment, the queue of the target function is adjusted and the function request is distributed intelligently according to the queuing status of the function request on the current computing node and the SLO status of the target function on the computing node. The SLO-sensitive scheduling strategy in this embodiment can greatly improve the processing efficiency of the function request, thereby achieving the technical effect of reducing the time consumption during data transmission and solving the technical problem of long time consumption during data transmission.
[0077] For example, based on the first execution count, the second execution count, and the service level objective, it is possible to determine how many times the objective function needs to be executed to meet the service level objective, thereby obtaining an evaluation result. Based on the evaluation result, the execution order of the multiple objective functions is determined. An SLO is a target for performance, availability, or other metrics that defines the minimum requirement or expected level that a service or objective function should achieve. It can refer to the number of requests processed by a function within a certain period of time or the response time requirement.
[0078] As an optional implementation, determining the evaluation result based on the first execution number, the second execution number and the service level target in the performance indicator includes: determining the tail delay value in the service level target; determining the evaluation result based on the first execution number, the second execution number and the tail delay value.
[0079] In this embodiment, a tail latency value in a service level objective can be determined, and an evaluation result can be determined based on the first execution number, the second execution number, and the tail latency value. The tail latency value can be a pre-set value, which can refer to the latency of a small number of requests that remain after the function has processed most requests. It can also be a threshold, such as the 99.9th percentile or 0.98th percentile. This is for illustration only and does not impose a specific limit on the tail latency value.
[0080] Alternatively, if the tail latency value is 0.98, this means that when the target function executes the corresponding function request, it must ensure that the latency of the 0.98 function request does not exceed a set threshold, where the threshold can be a specified time, such as 100 milliseconds. During data processing, using tail latency, the latency of relatively slow function requests executed by the target function can be controlled within an acceptable range, providing a better user experience.
[0081] Optionally, a computing node may include a scheduler having counting logic and timing logic for tracking the latency of each target function and the number of times the function has been executed, thereby determining the first execution count, the second execution count, and the latency of each function request for executing the target function.
[0082] Optionally, the evaluation result (RRC) can be calculated using the following formula: RRC = (P*NM / 1-P)
[0083] Among them, P can be used to represent the tail latency value corresponding to the SLO; N can be used to represent the number of times a certain target function has been executed, that is, the first execution number, which can be accumulated during the function execution process; M can be used to represent the number of execution results of the executed target function within the target time period that meet the SLO, that is, the second execution number.
[0084] In this embodiment, an evaluation result may be determined based on the first execution number, the second execution number, and the tail latency value. The evaluation result may be a predefined evaluation indicator, such as a required number of requests (RRC), which may be used to determine how many function requests the target function needs to execute in order to meet the SLO of a target function.
[0085] In this embodiment, the following steps are used to determine a reasonable execution order to address the time-consuming data transmission process. First, the evaluation results are determined based on the SLO, and second, the execution order is determined based on the evaluation results. The above steps mainly explain how to calculate the evaluation results based on the SLO. The following steps further explain how to determine the execution order based on the evaluation results.
[0086] As an optional implementation, determining the execution order based on the evaluation result includes: sorting the multiple objective functions based on the evaluation result to obtain a sorting result; and determining the execution order of the multiple objective functions according to the sorting result.
[0087] In this embodiment, according to the evaluation results of the objective functions, multiple objective functions can be sorted according to priority levels (which can be simply referred to as priorities) to obtain sorting results. According to the sorting results, the execution order of the multiple objective functions can be determined.
[0088] As an optional implementation, a scheme for determining the execution order of multiple objective functions according to the sorting results may include the following implementation steps: first, according to the sorting results, multiple objective functions are divided to obtain a first objective function group and a second objective function group; then, the execution order of the objective functions in the first objective function group and the execution order of the objective functions in the second objective function group are determined, wherein the execution order of the objective functions in the first objective function group is before the execution order of the objective functions in the second objective function group.
[0089] In this embodiment, multiple objective functions can be divided according to the sorting results to obtain a first objective function group and a second objective function group. The execution order of the objective functions in the first objective function group and the second objective function group is determined respectively, thereby obtaining the execution order of all objective functions. Among them, the first objective function group can be a high priority queue (High priority), which can also be called a first priority queue. The second objective function group can be a low priority queue (Low priority), which can also be called a second priority queue.
[0090] Optionally, the execution priority of the objective function in the first objective function group is higher than the execution priority of the objective function in the second objective function group. Therefore, the execution order of the objective function in the first objective function group is before the execution order of the objective function in the second objective function group. Multiple objective functions in the first objective function group can be executed first. After the execution of multiple objective functions in the first objective function group is completed, the multiple objective functions in the second objective function group can be executed.
[0091] Optionally, multiple target functions are divided into low-priority queues and high-priority queues, and the high-priority queue can be dequeued first and then the low-priority queue. After determining the execution order of the target functions, the target functions can be executed to issue a scheduling request to the image processor video memory, thereby completing the data replacement between the main memory and the video memory.
[0092] In this embodiment, in order to meet the SLO of multiple objective functions, the indicator of the required number of requests is defined. Through this indicator, all the objective functions to be executed on the computing node can be intuitively allocated to the high and low priority queues, thereby obtaining the first objective function group to be executed first, and the second objective function group to be executed later. For example, the objective functions (which can be simply referred to as functions) corresponding to the smaller RRC and the negative RRC can be placed in the high priority queue, and the functions corresponding to the larger RRC can be placed in the low priority queue to obtain the first objective function group and the second objective function group. After determining the first objective function group and the second objective function group, the execution order of the objective functions in the first objective function group and the second objective function group can be determined respectively.
[0093] In this embodiment, there are multiple methods for dividing multiple objective functions to obtain the first objective function group and the second objective function group. The multiple objective functions can be divided by using objective rankings or based on evaluation result thresholds. The following further describes two methods for determining the first objective function group and the second objective function group. However, it should be noted that the following two methods are only examples and do not specifically limit the methods for determining the first objective function group and the second objective function group.
[0094] As an optional implementation, multiple objective functions are divided according to the sorting results to obtain a first objective function group and a second objective function group, including: dividing the objective functions that are located before the target ranking in the sorting results into the first objective function group; dividing the objective functions that are located at the target ranking and after the target ranking in the sorting results into the second objective function group.
[0095] In this embodiment, multiple objective functions can be divided based on target ranking. Based on the scoring results, multiple objective functions can be sorted in order of execution priority from high to low to obtain a sorting result. The target ranking can be set. The objective functions that are located before the target ranking in the sorting result are all objective functions with high execution priority. Therefore, the objective functions that are located before the target ranking in the sorting result can be divided into the first objective function group, and the objective functions that are located before the target ranking and after the target ranking can be divided into the second objective function group. Among them, the target ranking can be a ranking set in advance according to actual conditions, for example, it can be fifth, sixth, etc. This is only for example, and no specific limitation is made to the size of the target ranking. The size of the target ranking can change according to actual conditions or changes in the number of objective functions in the first objective function group.
[0096] Optionally, when multiple objective functions are sorted in order of execution priority from low to high based on the scoring results, after obtaining the sorting results, the target ranking can be set. The objective functions at the target ranking and before the target ranking in the sorting results are all objective functions with low execution priority. Therefore, the objective functions at the target ranking and before the target ranking in the sorting results can be divided into the second objective function group, and the objective functions after the target ranking can be divided into the first objective function group.
[0097] For example, assuming that a computing node currently has N objective functions, and the N objective functions are sorted from small to large according to the size of RRC, the first A objective functions can be placed in the high priority queue, and the Ath and subsequent objective functions can be placed in the low priority queue.
[0098] For another example, suppose a computing node currently has N objective functions, and the N objective functions are sorted from small to large in RRC, where RRC i It can be used to represent the RRC corresponding to the objective function i. The first K objective functions can be placed in a high priority queue, where k can be the largest integer that satisfies the following formula:
[0099] Among them, RRC j It can be used to represent the RRC corresponding to the objective function j; K can be changed according to actual conditions.
[0100] As another optional implementation, based on the sorting results, multiple objective functions are divided to obtain a first objective function group and a second objective function group, including: dividing the objective functions in the sorting results whose evaluation results are less than the evaluation result threshold into the first objective function group; dividing the objective functions in the sorting results whose evaluation results are greater than or equal to the evaluation result threshold into the second objective function group.
[0101] In this embodiment, multiple objective functions can be sorted based on the scoring results to obtain a sorting result. In the sorting result, objective functions whose evaluation results are less than the evaluation result threshold can be divided into a first objective function group, and objective functions whose evaluation results are greater than or equal to the evaluation result threshold can be divided into a second objective function group. The evaluation result threshold can be a pre-set value, such as an RRC threshold, for example, 1 or 0. This is only for example purposes and does not impose any specific restrictions on the size of the evaluation result threshold.
[0102] Optionally, this embodiment may define that when the RRC values are not less than the RRC threshold, all target functions that are not less than the RRC threshold are placed in a low priority queue; when the RRC values are less than the RRC threshold, all target functions that are less than the RRC threshold are placed in a high priority queue. Alternatively, the target functions may be sorted from small to large according to the RRC value, with the functions before the target ranking placed in a high priority queue, and the target functions after the target ranking placed in a low priority queue. It should be noted that this is merely an example and does not impose any specific restrictions on the method for determining high and low priority queues.
[0103] In this embodiment, RRC can be used to guide the execution of queued function requests. If the target functions requested by function requests are simply sorted based on RRC, for example, the target functions corresponding to function requests with smaller RRCs will be executed first. This will lead to a snowball effect, that is, making good SLOs better and bad SLOs worse. To avoid this problem, two self-regulating priority queues can be used to solve this problem. Different priority queues will execute in different orders, thereby avoiding the snowball effect during the execution of the target functions.
[0104] In this embodiment, multiple methods are proposed for determining the first and second objective function groups. Different methods are used to determine the execution order for different objective function groups, thereby reducing the time consumed during data transmission and avoiding a snowball effect. The following further describes the two resulting self-adjusting priority queues, specifically the execution order of the first and second objective function groups.
[0105] As an optional implementation, in response to the number of objective functions in the first objective function group being less than a quantity threshold, the evaluation result threshold is adjusted.
[0106] In this embodiment, the distribution of objective functions within the high and low priority queues can be dynamically adjusted by automatically adjusting the size of the evaluation result threshold. That is, when the number of objective functions in the first objective function group is less than the quantity threshold, the evaluation result threshold can be adjusted, for example, the evaluation result threshold can be increased or decreased. No specific restrictions are placed on the method of adjusting the evaluation result threshold. The quantity threshold can be used to limit the number of objective functions in the first objective function group and can be a pre-set value. The quantity threshold can be changed according to actual conditions.
[0107] Optionally, the quantity threshold may be changed according to the change in the quantity of the objective functions in the first objective function group, or may be fixed, and may be set and selected according to actual needs.
[0108] Optionally, the RRC threshold can fluctuate with the number of requested target functions to dynamically adjust the number of target functions in the high- and low-priority queues. For example, when there are fewer target functions in the high-priority queue, the RRC threshold can be increased to allow more target functions to enter the high-priority queue.
[0109] In this embodiment, the evaluation result threshold or the target ranking may be changed according to the change in the number of the target functions in the first target function group, thereby achieving dynamic adjustment of the distribution of the target functions in the priority queue.
[0110] As an optional implementation, the execution order of the objective functions in the first objective function group is to execute the objective functions in the first objective function group in descending order according to the evaluation results.
[0111] In this embodiment, the objective functions in the first objective function group may be executed sequentially according to the size of the evaluation results, from large to small.
[0112] Optionally, this embodiment can determine the execution level of all objective functions in the first objective function group based on the evaluation result, and execute the objective functions in the first objective function group in order from large to small execution levels, or when the evaluation result is RRC, the objective functions in the first objective function group can be executed in order from large to small RRC.
[0113] In this embodiment, in order to meet the SLOs of multiple target functions, an indicator is defined. This indicator can be the required number of requests. This indicator can be used to determine how many function requests the target function needs to execute in order to meet the SLO of a certain target function. Therefore, for target functions that are very close to or have already met the SLO, the number of function requests that need to be executed is very small, or even unnecessary; for target functions that do not meet the SLO, the opposite is true. In other words, the closer the RRC corresponding to the target function is to 0, the more likely the target function can achieve the SLO of the function with a smaller number of function requests, and therefore the execution priority of the target function can be increased; the larger the RRC, the further the target function is from the SLO, the greater the number of function requests required, and the lower the priority of this target function.
[0114] Optionally, in a high-priority queue, to avoid a snowball effect, it is set that the target function with a larger RRC in the high-priority queue should be executed first. Therefore, the function with a larger RRC is dequeued first; and the target functions corresponding to smaller RRC or even negative RRC are very close to their own SLO and have a low priority, so they will not be executed first. After the target function with a larger RRC is executed, the target function with a smaller RRC or even a negative RRC will be executed.
[0115] As an optional implementation, the execution order of the objective functions in the second objective function group is to execute the objective functions in the second objective function group in ascending order according to the evaluation results.
[0116] In this embodiment, the objective functions in the second objective function group are executed in ascending order based on the evaluation results. That is, in the low-priority queue, objective functions with smaller RRCs are dequeued first because they are closer to their own SLOs, and objective functions with larger RRCs are executed later.
[0117] Optionally, multiple objective functions are first divided based on the RRC size to obtain a first objective function group and a second objective function group. If the objective function corresponding to the function request with the smaller RRC is simply set to be executed first, this will bring about a snowball effect, that is, it will make the good SLO better and the bad SLO worse. Therefore, in order to avoid the snowball effect, different execution rules are set for the objective functions in the first objective function group and the objective functions in the second objective function group. That is, in the first objective function group, the objective function that requires more effort to meet the SLO can be executed first, while the objective function that requires less effort to meet the SLO can be executed later; in the second objective function group, since both require greater effort to meet the SLO, the objective function with the smaller RRC can be executed first, and the objective function with the larger RRC can be executed later.
[0118] As an optional implementation, the method may further include: in the second objective function group, in response to the evaluation result of the first objective function being greater than a target result threshold, prohibiting execution of the first objective function.
[0119] In this embodiment, it is determined whether the evaluation result of the first objective function in the second objective function group is greater than the target result threshold. If the evaluation result of the first objective function is greater than the target result threshold, execution of the first objective function may be prohibited. The target result threshold may be a pre-set value, such as 5, 6, or level 1, level 2, etc. This is for illustration only and does not impose any specific limitation on the target result threshold.
[0120] Optionally, in the second objective function group, if the evaluation result of the first objective function is greater than the target result threshold, it means that executing the first objective function requires a lot of time or memory. Therefore, in order to avoid the problem of excessive time consumption in the process of processing data, the execution of the first objective function can be prohibited.
[0121] In this embodiment, multiple function requests are obtained from the host machine, the target function requested by the function request is determined, and based on the performance indicators of the target function, the execution order of the multiple target functions is determined, thereby achieving the technical effect of reducing the time consumed in the data transmission process and solving the technical problem of long time consumed in the data transmission process.
[0122] The embodiment of the present disclosure also provides a data processing method based on server-unaware computing, which can be applied to a video memory replacement scenario. The video memory replacement scenario can be a video memory replacement scenario of a server-free architecture image processor, or can be a way of using GPU resources for computing in a server-free computing environment. In this environment, the video memory replacement of the GPU can be used to solve the scenario of insufficient video memory. It should be noted that this is only an example and does not impose specific restrictions on the scenario of video memory replacement. Figure 4 is a flowchart of another data processing method based on server-unaware computing according to an embodiment of the present disclosure. As shown in Figure 4, the method may include the following steps:
[0123] Step S402: Acquire multiple data processing requests from the host machine.
[0124] Step S404: determining a target function that any data processing request requests the image processor to execute, and obtaining a plurality of target functions, wherein the target function is at least used for data copying or data synchronization.
[0125] In the technical solution provided in step S404 of the present disclosure, multiple data processing requests from the host machine can be obtained, and the target function requested by any of the data processing requests can be determined to obtain multiple target functions. The data processing request can be a request for data copying or data synchronization. The target function can at least be used for data copying or data synchronization. The data can be a model, an image, or other data. This is for illustrative purposes only and does not impose any specific restrictions on the data type.
[0126] Alternatively, data copying may involve copying data from the host computer's main memory to the graphics processor's video memory, or copying data from the video memory to the main memory. Data synchronization may involve synchronizing data from the video memory to the main memory, or copying data from the main memory to the video memory. This is merely an example and does not impose any specific limitation on the direction of data synchronization or data copying.
[0127] Step S406, based on the performance indicators of the multiple objective functions, determining the execution order of the multiple objective functions, wherein the performance indicators of the objective functions are used to characterize the performance of the objective functions during the execution process, and the execution order is used to determine the data copied or synchronized between the main memory of the host machine and the video memory of the image processor.
[0128] With respect to the intelligent scheduling scheme in the context of video memory replacement, at each moment, only one function request will be executed on a GPU device, and the execution of each function request will occupy almost the entire video memory resources of the card, and other function requests routed to the GPU device need to queue and wait. Therefore, when there are multiple function requests that need to be executed, an intelligent scheduling scheme is needed to reduce the end-to-end time consumption caused by the queuing of function requests. In this embodiment, when there are a large number of function requests sent to the back-end computing node of the same function computing service, the target function requested by any function request is determined, and multiple target functions are obtained. Through intelligent scheduling and queuing strategies, the SLO of each target function can be met as much as possible, thereby determining the execution order of multiple target functions, and executing multiple target functions in sequence according to the execution order, thereby reducing the queuing delay in the process of executing the target function.
[0129] For example, when multiple function requests arrive at the Function Compute Platform (referred to as the platform), the platform checks whether the function requests require a swap between host main memory and GPU memory. The Function Compute Platform assigns the function requests to compute nodes to execute the target functions requested by the function requests. After receiving the function requests, the compute nodes invoke the scheduler to determine the execution order of the multiple target functions and execute them sequentially according to the execution order. During the execution of the target functions, function instances are started and the function code is loaded. The function instances perform initialization operations such as loading models and reading data. During function execution, if a swap between host main memory and GPU memory is required, the compute nodes perform the corresponding data transfers as needed. The compute nodes transfer the required data from host main memory to GPU memory, or vice versa. This process may involve operations such as data copying, memory management, and synchronization. Once the data swap is complete, the function executes the corresponding computational operations on the GPU according to the execution logic. This may include image processing, model inference, rendering, and so on. After the function is executed, the compute node returns the result to the function computing platform, which then returns the result to the caller. The compute node can be a virtual machine instance with certain computing resources and GPU devices.
[0130] It should be noted that the above video memory replacement process is only an example and is not specifically limited here.
[0131] In this embodiment, the Function Compute Platform is responsible for scheduling and managing the execution of the target function, while the compute nodes are responsible for actual function execution and data transfer. Function developers only need to focus on the function's functional logic and do not need to directly handle the swapping of host main memory and GPU memory. The Function Compute Platform automatically manages resources and data transfer to provide flexible and efficient Function Compute services. Function developers can also be users of the host machine.
[0132] Optionally, the video memory can be the memory space in the GPU used to store the data and intermediate results required for calculation. In some computationally intensive tasks, when the amount of data is large or the computational requirements are high, the video memory may be insufficient. At this time, video memory replacement can move some infrequently used data out of the video memory to make room for more important data. For example, when performing large-scale image processing tasks, it may be necessary to process multiple high-resolution images at the same time, which takes up a large amount of video memory. If the video memory is insufficient, some images can be moved out of the video memory and moved back when needed to ensure that the calculation is carried out; or when training a deep learning model, the model parameters and intermediate calculation results may take up a large amount of video memory. When the model is large or the batch size is large, the video memory may be insufficient. Video memory replacement can move some infrequently used intermediate calculation results out of the video memory to make room for more important data. This is just an example and does not impose specific restrictions on the scenarios where video memory replacement is required.
[0133] The current Function Compute service integrates memory swap technology, which records the metadata and data of inference applications in GPU memory. However, this technology only allows one inference application to execute on a GPU card at a time, and other inference applications will be queued. The queuing of function requests to be executed will significantly delay end-to-end performance for users, resulting in a technical problem of long data transmission time. To address the above problem, this embodiment proposes an intelligent memory swap scheduling method for achieving request SLOs in the serverless GPU memory swap scenario. This method combines the characteristics of the GPU inference scenario and introduces multiple priority queues to intelligently adjust the queue of functions and distribute function requests based on the queue status of the current node and the SLO status of the function on the node. In this embodiment, the SLO-sensitive scheduling strategy can greatly improve the efficiency of memory swap, thereby reducing the end-to-end request latency of serverless GPU users, achieving the technical effect of reducing the time consumption during data transmission, and solving the technical problem of long data transmission time.
[0134] Through the above-mentioned steps S402 to S410 of the present invention, multiple data processing requests are obtained from the host machine; the target function requested by any data processing request is determined to obtain multiple target functions, wherein the target function is at least used for data copying or data synchronization; based on the performance indicators of the multiple target functions, the execution order of the multiple target functions is determined, wherein the performance indicators of the target functions are used to characterize the performance of the target functions during the execution process; the multiple target functions are executed in sequence according to the execution order; and using the executed target functions, data transmission between the main memory of the host machine and the video memory used for image processing is scheduled, thereby achieving the technical effect of reducing the time consumed in the data transmission process and solving the technical problem of long time consumed in the data transmission process.
[0135] Example 2
[0136] Currently, in machine learning fields such as artificial intelligence (AI), in order to quickly and agilely deploy trained models on cloud servers, serverless architecture image processors are usually used to achieve fast and agile deployment of trained models on cloud servers.
[0137] In one implementation, Serverless GPU is a new and emerging cloud computing GPU service with advantages such as no maintenance and pay-as-you-go. However, its relatively high usage cost and unpredictable model cold start time make it difficult for many users to quickly migrate existing applications.
[0138] Function Compute can optionally use GPU resources on a pay-as-you-go basis, lowering the barrier to entry for developers and engineers. The current Function Compute service integrates memory swapping technology, which records the metadata and data of inference applications in GPU memory. However, this technology only allows one inference application to execute on a GPU at a time, forcing other inference applications to queue. This queuing of pending requests significantly reduces end-to-end performance for users, leading to technical issues such as lengthy data transfers.
[0139] To solve the above problems, this embodiment proposes an intelligent memory replacement scheduling method for achieving request SLO in the memory replacement scenario of Serverless GPU. This method combines the characteristics of GPU inference scenarios and introduces multiple priority queues to intelligently adjust the queue of functions and distribute function requests based on the queuing status of the current node and the SLO status of the function on the node. In this embodiment, the SLO-sensitive scheduling strategy can greatly improve the efficiency of memory replacement, thereby reducing the end-to-end request delay of Serverless GPU users, achieving the technical effect of reducing the time consumption in the data transmission process, and solving the technical problem of long time consumption in the data transmission process.
[0140] The following further introduces the intelligent memory replacement scheduling method for achieving the requested SLO.
[0141] FIG5 is a schematic diagram of a function computing service architecture according to an embodiment of the present disclosure. The function computing service architecture in the embodiment of the present disclosure may include multiple computing nodes. As shown in FIG5 , cluster-wide function invocation utilizes the function computing service. The function computing service architecture may include a computing node 501 (FC Node). On a computing node 501, there are four GPU workers, for example, worker 502, worker 503, worker 504, and worker 505. One worker corresponds to a real GPU device. The GPU device can be time-divided and reused. At the same time, one GPU device can only serve one function request.
[0142] In this embodiment, a computing node 501 further includes a scheduler 506 , which can execute counting logic and timing logic, and can be used to track the delay of each target function and the number of times the function has been executed.
[0143] Optionally, there are two queues on the computing node 501, namely a high priority queue 507 and a low priority queue 508. Each GPU worker can be understood as a consumer, and the priority queue can be understood as a producer.
[0144] In this embodiment, after the function computing platform receives a function request from the host machine, it determines the target function corresponding to the function request and obtains the cluster-wide target function. The cluster-wide target function is scheduled using the scheduler in the computing node. In response to the need to schedule multiple target functions, the multiple functions can be divided into high-priority queues and low-priority queues. Functions in the high-priority queue are dequeued first, and functions in the low-priority queue are dequeued later. After determining the priority of function scheduling, a scheduling request can be issued to control the graphics processor to execute the target function. Based on the target function, the graphics processor video memory can be controlled to complete video memory replacement.
[0145] In this embodiment, for the intelligent scheduling solution in the context of video memory swapping, only one function request is executed on a GPU device at any given moment, and the execution of each function request consumes almost the entire video memory resources of the graphics processor. Other function requests routed to the GPU device must wait in queue. Therefore, an intelligent scheduling solution is needed to reduce the end-to-end time caused by queuing.
[0146] Optionally, when a large number of function requests are sent to the backend computing nodes of the same function computing service, intelligent scheduling and queuing strategies can be used to meet the SLO of each function as much as possible, thereby reducing queuing delays.
[0147] In this embodiment, a metric is defined to meet the SLOs of multiple objective functions. This metric can be the number of required requests. This metric can be used to intuitively allocate all objective functions to be executed on the compute node into high- and low-priority queues. For example, objective functions corresponding to smaller and negative RRCs can be placed in the high-priority queue, while functions corresponding to larger RRCs can be placed in the low-priority queue.
[0148] In this embodiment, the function distribution in the high and low priority queues can be dynamically adjusted through automatic adjustment of RRC.
[0149] Optionally, this embodiment may define that when the RRC values are not less than the RRC threshold, all target functions that are not less than the RRC threshold may be placed in a low priority queue; when the RRC values are less than the RRC threshold, all target functions that are less than the RRC threshold may be placed in a high priority queue. Alternatively, the target functions may be sorted from small to large according to the RRC value, with the functions before the target ranking placed in a high priority queue, and the target functions after the target ranking placed in a low priority queue. It should be noted that this is merely an example and does not impose any specific restrictions on the method for determining high and low priority queues.
[0150] For example, an RRC threshold (β) can be defined between 0 and 1 (i.e., 0 ≤ β ≤ 1). When β = 0, all target functions are placed in the high-priority queue. When β = 1, all target functions are placed in the low-priority queue. Furthermore, the β value can be dynamically adjusted based on load fluctuations and the number of function requests in the high- and low-priority queues. When there are fewer target functions in the high-priority queue, β can be increased to allow more functions to enter the high-priority queue.
[0151] For another example, suppose a computing node currently has N objective functions, and the N objective functions are sorted from small to large in RRC, where RRC i It can be used to represent the RRC corresponding to the objective function i. The first K objective functions can be placed in a high priority queue, where k can be the largest integer that satisfies the following formula:
[0152] Among them, RRC j It can be used to represent the RRC corresponding to the objective function j.
[0153] Optionally, the RRC threshold can fluctuate with the number of requested target functions to dynamically adjust the number of target functions in the high- and low-priority queues. For example, when there are fewer target functions in the high-priority queue, the RRC threshold can be increased to allow more target functions to enter the high-priority queue.
[0154] In this embodiment, in order to meet the SLOs of multiple objective functions, an indicator is defined. This indicator can be the required number of requests. This indicator can be used to determine how much "effort" is needed to meet the SLO of a certain objective function. Therefore, for objective functions that are very close to or have already met the SLO, this "effort" is very small or even unnecessary; the opposite is true for objective functions that are not met.
[0155] Optionally, RRC can be used to represent how many function requests need to be executed to meet the SLO of the target function. The RRC can be determined by the following formula: RRC = (P*NM / 1-P)
[0156] Among them, P can be used to represent the tail latency value corresponding to the SLO; N can be used to represent the number of times a certain target function has been executed, that is, the first execution number, which can be accumulated during the function execution process; M can be used to represent the number of execution results of the executed target function within the target time period that meet the SLO, that is, the second execution number.
[0157] Optionally, the closer RRC is to 0, the more the objective function can achieve its SLO with less effort, and therefore the execution priority of this objective function should be increased; the larger the RRC is, the further the objective function is from the SLO, the more effort is required, and the priority of this objective function will be lowered.
[0158] In this embodiment, RRC can guide how to execute queued function requests. If the target functions requested by function requests are simply sorted based on RRC, for example, the target functions corresponding to function requests with smaller RRCs will be executed first. This will lead to a snowball effect, that is, making good SLOs better and bad SLOs worse. To avoid this problem, two self-regulating priority queues can be used to solve this problem. Different priority queues have different execution orders, thus avoiding the snowball effect during the execution of the target functions.
[0159] Optionally, in a high-priority queue, because the objective function with a larger RRC is very close to meeting the SLO, it should be executed first. Therefore, the function with a larger RRC is dequeued first; and the objective functions corresponding to those with smaller RRC or even negative RRC are already very close to their own SLO and have a low priority, so they will not be executed first. After the objective function with a larger RRC is executed, the objective function with a smaller RRC or even a negative RRC will be executed.
[0160] Optionally, in the low-priority queue, objective functions with smaller RRCs are dequeued first, because these objective functions are closer to their own SLOs, and objective functions with larger RRCs are more likely to be discarded here.
[0161] In this embodiment, the intelligent scheduling capability based on the memory replacement technology can ensure the SLO of function requests. By integrating memory replacement into the production environment through the intelligent scheduling system, it significantly reduces costs for users while also ensuring end-to-end performance, and also improves the platform resource sales rate on the server side. Based on the time-sharing reuse of GPU devices, this embodiment uses a priority queue to solve the scheduling problems in the two major intersection areas of FaaS and GPU inference, further improving the utilization rate of GPU devices and correspondingly reducing the request queue backlog caused by overselling of memory.
[0162] The method embodiment provided in Example 1 of the present disclosure can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 6 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a data processing method based on server-less computing according to an embodiment of the present disclosure. As shown in Figure 6, the computer terminal 60 (or mobile device) may include one or more (602a, 602b, ..., 602n are used in the figure to illustrate) processors 602 (the processor 602 may include but is not limited to a microprocessor (Microcontroller Unit, referred to as MCU) or a programmable logic device (Field Programmable Gate Array, referred to as FPGA) and other processing devices), a memory 604 for storing data, and a transmission device 606 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that the structure shown in Figure 1 is only for illustration and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 60 may also include more or fewer components than shown in FIG. 6 , or have a configuration different from that shown in FIG. 6 .
[0163] The hardware structure block diagram shown in Figure 6 can not only serve as an exemplary block diagram of the above-mentioned computer terminal 60 (or mobile device), but also as an exemplary block diagram of the above-mentioned server. In an optional embodiment, Figure 2 shows in a block diagram an embodiment of using the computer terminal 60 (or mobile device) shown in Figure 6 as a computing node in the computing environment 201.
[0164] The memory 604 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the data processing method based on server-insensitive computing in the embodiment of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 604, that is, realizing the above-mentioned data processing method based on server-insensitive computing. The memory 604 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 604 may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal 60 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0165] Transmission device 606 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of computer terminal 60. In one embodiment, transmission device 606 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 606 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0166] The display may be, for example, a touch screen liquid crystal display (LCD), which enables a user to interact with a user interface of the computer terminal 60 (or mobile device).
[0167] In another alternative embodiment, FIG7 shows a block diagram of an embodiment using the computer terminal 60 (or mobile device) shown in FIG6 as a service grid. FIG7 is a block diagram of a service grid structure according to an embodiment of the present disclosure. As shown in FIG7, the service grid 700 is primarily used to facilitate secure and reliable communication between multiple microservices. Microservices refer to the decomposition of an application into multiple smaller services or instances, which are distributed and run on different clusters / machines.
[0168] As shown in Figure 7, microservices may include application service instance A and application service instance B, which form the functional application layer of service grid 700. In one embodiment, application service instance A runs as a container / process 708 on a machine / workload container group 714 (POD), and application service instance B runs as a container / process 710 on a machine / workload container group 716 (POD).
[0169] In one embodiment, application service instance A may be a data copy service, and application service instance B may be a data transfer service.
[0170] As shown in Figure 7 , application service instance A and grid proxy (sidecar) 703 coexist in machine workload container group 614, while application service instance B and grid proxy 705 coexist in machine workload container 714. Grid proxy 703 and grid proxy 705 form the data plane layer (dataplane) of service grid 700. Grid proxy 703 and grid proxy 705 each run as container / process 704, which can receive requests 712 for product query services, and grid proxy 706. Bidirectional communication is possible between grid proxy 703 and application service instance A, and between grid proxy 705 and application service instance B. Furthermore, bidirectional communication is also possible between grid proxy 703 and grid proxy 705.
[0171] In one embodiment, all traffic for application service instance A is routed to the appropriate destination via grid proxy 703, and all network traffic for application service instance B is routed to the appropriate destination via grid proxy 705. It should be noted that network traffic referred to herein includes, but is not limited to, Hypertext Transfer Protocol (HTTP), Representational State Transfer (REST), the high-performance, general-purpose open source framework (Google Remote Procedure Call, gRPC), and the open source in-memory data structure storage system (Redis).
[0172] In one embodiment, the data plane layer's functionality can be extended by writing custom filters for the proxy (Envoy) in service mesh 700. The service mesh proxy configuration can be designed to enable the service mesh to correctly proxy service traffic, enabling service interoperability and service governance. Mesh proxy 703 and mesh proxy 705 can be configured to perform at least one of the following functions: service discovery, health checking, routing, load balancing, authentication and authorization, and observability.
[0173] As shown in Figure 7 , service grid 700 also includes a control plane layer. The control plane layer can be comprised of a set of services running in a dedicated namespace, hosted by a managed control plane component 701 within a machine / workload container group (machine / pod) 702. As shown in Figure 7 , managed control plane component 701 communicates bidirectionally with mesh proxy 703 and mesh proxy 705. Managed control plane component 701 is configured to perform certain control and management functions. For example, managed control plane component 701 receives telemetry data transmitted by mesh proxy 703 and mesh proxy 705 and can further aggregate this telemetry data. Managed control plane component 701 can also provide user-oriented application programming interfaces (APIs) for these services, making it easier to manipulate network behavior and provide configuration data to mesh proxy 703 and mesh proxy 705.
[0174] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present disclosure.
[0175] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present disclosure.
[0176] Example 3
[0177] According to an embodiment of the present disclosure, a data processing device based on server-insensitive computing is also provided for implementing the data processing method based on server-insensitive computing shown in FIG. 3 above.
[0178] Figure 8 is a schematic diagram of a data processing device based on server-insensitive computing according to an embodiment of the present disclosure. As shown in Figure 8, the data processing device 800 based on server-insensitive computing may include: a first acquisition component 802, a first determination component 804 and a second determination component 806.
[0179] The first acquisition component 802 is configured to acquire multiple function requests from the host machine.
[0180] The first determining component 804 is configured to determine a target function that any function request requests the image processor to execute, and obtain multiple target functions.
[0181] The second determination component 806 is configured to determine the execution order of multiple objective functions based on the performance indicators of multiple objective functions, wherein the performance indicators of the objective functions are used to characterize the performance of the objective functions during the execution process, and the execution order is used to determine the data transmitted between the main memory of the host machine and the video memory of the image processor.
[0182] Here, the first acquisition component 802, the first determination component 804, and the second determination component 806 correspond to steps S302 to S306 in Example 1. The three components and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in Example 1. It should be noted that the above components can be hardware components or software components stored in a memory (e.g., memory 604) and processed by one or more processors (e.g., processors 602a, 602b..., 602n). The above components can also be part of the device and can be run in the computer terminal 60 provided in Example 2.
[0183] According to an embodiment of the present disclosure, a data processing device based on server-insensitive computing is also provided for implementing the data processing method based on server-insensitive computing shown in FIG. 4 above.
[0184] Figure 9 is a schematic diagram of another data processing device based on server-insensitive computing according to an embodiment of the present disclosure. As shown in Figure 9, the data processing 900 based on server-insensitive computing may include: a second acquisition component 902, a third determination component 904 and a fourth determination component 906.
[0185] The second acquisition component 902 is configured to acquire multiple data processing requests from the host machine.
[0186] The third determining component 904 is configured to determine a target function that any data processing request requests the image processor to execute, and obtain multiple target functions, wherein the target function is at least used for data copying or data synchronization.
[0187] The fourth determination component 906 is configured to determine the execution order of multiple target functions based on the performance indicators of multiple target functions, wherein the performance indicators of the target functions are used to characterize the performance of the target functions during the execution process, and the execution order is used to determine the data copied or synchronized between the main memory of the host machine and the video memory of the image processor.
[0188] It should be noted that the second acquisition component 902, the third determination component 904, and the fourth determination component 906 correspond to steps S402 to S406 in Example 1. The three components and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in Example 1. It should be noted that the above components can be hardware components or software components stored in a memory (e.g., memory 604) and processed by one or more processors (e.g., processors 602a, 602b..., 602n). The above components can also be part of the device and can be run in the computer terminal 60 provided in Example 2.
[0189] In this data processing device based on server-imperceptible computing, multiple function requests are obtained from the host machine, the target function requested by the function request is determined, and the execution order of the multiple target functions is determined based on the performance indicators of the target function, thereby achieving the technical effect of reducing the time consumed in the data transmission process and solving the technical problem of long time consumed in the data transmission process.
[0190] Example 5
[0191] The embodiment of the present disclosure may provide a computer terminal, which may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal may also be replaced by a terminal device such as a mobile terminal.
[0192] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of a computer network.
[0193] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the data processing method based on server-imperceptible computing: obtaining multiple function requests from the host machine; determining the target function that any function request requests the image processor to execute, and obtaining multiple target functions; based on the performance indicators of the multiple target functions, determining the execution order of the multiple target functions, wherein the performance indicators of the target functions are used to characterize the performance of the target functions during the execution process, and the execution order is used to determine the data transmitted between the main memory of the host machine and the video memory of the image processor.
[0194] Optionally, Figure 10 is a structural block diagram of a computer terminal according to an embodiment of the present disclosure. As shown in Figure 10, the computer terminal A may include: one or more (only one is shown in the figure) processors 1002, a memory 1004 and a transmission device 1006.
[0195] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the data processing method and device based on server-insensitive computing in the embodiments of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned data processing method based on server-insensitive computing. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0196] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain multiple function requests from the host machine; determine the target function that any function request requests the image processor to execute, and obtain multiple target functions; based on the performance indicators of the multiple target functions, determine the execution order of the multiple target functions, wherein the performance indicators of the target functions are used to characterize the performance of the target functions during the execution process, and the execution order is used to determine the data transmitted between the host machine's main memory and the image processor's video memory.
[0197] Optionally, the processor may further execute program codes of the following steps: determining an evaluation result of the objective function based on a performance indicator of the objective function; and determining an execution order based on the evaluation result.
[0198] Optionally, the processor may also execute the program code of the following steps: determining the first number of executions of the target function that has been completed within a historical time period, and the second number of executions of the target function that has been executed, the execution results of which meet the performance indicator threshold; and determining the evaluation result of the target function at the current moment based on the first number of executions, the second number of executions and the performance indicator.
[0199] Optionally, the processor may also execute the program code of the following steps: obtaining a service level target in the performance indicator, wherein the service level target is used to characterize the execution of the objective function; and determining an evaluation result based on the first execution number, the second execution number and the service level target.
[0200] Optionally, the processor may further execute program code of the following steps: determining a tail delay value in the service level objective; and determining an evaluation result based on the first execution quantity, the second execution quantity, and the tail delay value.
[0201] Optionally, the processor may also execute the program code of the following steps: based on the evaluation results, sorting multiple objective functions to obtain a sorting result; dividing multiple objective functions according to the sorting result to obtain a first objective function group and a second objective function group; determining the execution order of the objective functions in the first objective function group and the execution order of the objective functions in the second objective function group, wherein the execution order of the objective functions in the first objective function group is before the execution order of the objective functions in the second objective function group.
[0202] Optionally, the processor may also execute the following program code: dividing the objective functions that are before the target ranking in the sorting results into a first objective function group; dividing the objective functions that are before and after the target ranking in the sorting results into a second objective function group.
[0203] Optionally, the above-mentioned processor can also execute the program code of the following steps: dividing the objective functions in the sorting results whose evaluation results are less than the evaluation result threshold into the first objective function group; dividing the objective functions in the sorting results whose evaluation results are greater than or equal to the evaluation result threshold into the second objective function group.
[0204] Optionally, the processor may further execute program code of the following steps: in response to the number of objective functions in the first objective function group being less than a number threshold, adjusting an evaluation result threshold.
[0205] Optionally, the processor may further execute program code of the following steps: in the second objective function group, in response to an evaluation result of the first objective function being greater than a target result threshold, prohibiting execution of the first objective function.
[0206] The processor can call information and applications stored in the memory through a transmission device to perform the following steps: obtain multiple data processing requests from the host machine; determine the target function that any data processing request requests the image processor to execute, and obtain multiple target functions, wherein the target function is at least used for data copying or data synchronization; based on the performance indicators of the multiple target functions, determine the execution order of the multiple target functions, wherein the performance indicators of the target functions are used to characterize the performance of the target functions during the execution process, and the execution order is used to determine the data to be copied or synchronized between the host machine's main memory and the image processor's video memory.
[0207] By adopting the embodiment of the present disclosure, multiple function requests from the host machine are obtained, the target functions requested by the function requests are determined, and the execution order of the multiple target functions is determined based on the performance indicators of the target functions, thereby achieving the technical effect of reducing the time consumed in the data transmission process and solving the technical problem of long time consumed in the data transmission process.
[0208] Those skilled in the art will appreciate that the structure shown in FIG10 is merely illustrative, and that computer terminal A may also be a smartphone (e.g., an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal device. FIG10 does not limit the structure of the aforementioned computer terminal A. For example, computer terminal A may include more or fewer components (e.g., a network interface, a display device, etc.) than those shown in FIG10 , or may have a configuration different from that shown in FIG10 .
[0209] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0210] Example 6
[0211] The embodiment of the present disclosure further provides a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the data processing method based on server-unaware computing provided in the first embodiment.
[0212] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0213] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: obtaining multiple function requests from the host machine; determining a target function that any function request requests the image processor to execute, to obtain multiple target functions; and determining an execution order of the multiple target functions based on performance indicators of the multiple target functions, wherein the performance indicators of the target functions are used to characterize the performance of the target functions during execution, and the execution order is used to determine the data transmitted between the host machine's main memory and the image processor's video memory.
[0214] Optionally, the computer-readable storage medium may further execute program code for the following steps: determining an evaluation result of the objective function based on a performance indicator of the objective function; and determining an execution order based on the evaluation result.
[0215] Optionally, the above-mentioned computer-readable storage medium can also execute the program code of the following steps: determining the first execution number of the target function that has been executed within a historical time period, and the second execution number of the target function that has been executed, the execution results of which meet the performance indicator threshold; based on the first execution number, the second execution number and the performance indicator, determining the evaluation result of the target function at the current moment.
[0216] Optionally, the above-mentioned computer-readable storage medium can also execute the program code of the following steps: obtaining the service level target in the performance indicator, wherein the service level target is used to characterize the execution of the objective function; and determining the evaluation result based on the first execution number, the second execution number and the service level target.
[0217] Optionally, the computer-readable storage medium may further execute program code for the following steps: determining a tail delay value in a service level objective; and determining an evaluation result based on the first execution quantity, the second execution quantity, and the tail delay value.
[0218] Optionally, the above-mentioned computer-readable storage medium can also execute the program code of the following steps: based on the evaluation results, sorting multiple objective functions to obtain sorting results; dividing multiple objective functions according to the sorting results to obtain a first objective function group and a second objective function group; determining the execution order of the objective functions in the first objective function group and the execution order of the objective functions in the second objective function group, wherein the execution order of the objective functions in the first objective function group is before the execution order of the objective functions in the second objective function group.
[0219] Optionally, the computer-readable storage medium can also execute the program code of the following steps: dividing the objective functions that are located before the target ranking in the sorting results into the first objective function group; dividing the objective functions that are located before the target ranking and after the target ranking in the sorting results into the second objective function group.
[0220] Optionally, the above-mentioned computer-readable storage medium can also execute the program code of the following steps: dividing the objective functions in the sorting results whose evaluation results are less than the evaluation result threshold into the first objective function group; dividing the objective functions in the sorting results whose evaluation results are greater than or equal to the evaluation result threshold into the second objective function group.
[0221] Optionally, the computer-readable storage medium may further execute program code for the following steps: in response to the number of objective functions in the first objective function group being less than a number threshold, adjusting an evaluation result threshold.
[0222] Optionally, the computer-readable storage medium may further execute program code for the following steps: in the second objective function group, in response to an evaluation result of the first objective function being greater than a target result threshold, prohibiting execution of the first objective function.
[0223] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: obtaining multiple data processing requests from a host machine; determining a target function that any one of the data processing requests requests the image processor to execute, to obtain multiple target functions, wherein the target function is at least used for data copying or data synchronization; based on performance indicators of the multiple target functions, determining an execution order of the multiple target functions, wherein the performance indicators of the target functions are used to characterize the performance of the target functions during execution, and the execution order is used to determine the data to be copied or synchronized between the host machine's main memory and the image processor's video memory.
[0224] In the embodiment of the present disclosure, multiple function requests are obtained from the host machine, the target function requested by the function request is determined, and based on the performance indicators of the target function, the execution order of the multiple target functions is determined, thereby achieving the technical effect of reducing the time consumed in the data transmission process and solving the technical problem of long time consumed in the data transmission process.
[0225] Example 7
[0226] An embodiment of the present disclosure may provide an electronic device, which may include a memory and a processor.
[0227] 11 is a block diagram of an electronic device according to an embodiment of the present disclosure, which is based on a data processing method for server-insensitive computing. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0228] As shown in FIG11 , the device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. Various programs and data required for the operation of the device 1100 may also be stored in the RAM 1103. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0229] Various components in device 1100 are connected to I / O interface 1105, including: an input unit 1106, such as a keyboard, mouse, etc.; an output unit 1104, such as various types of displays, speakers, etc.; a storage unit 1108, such as a magnetic disk, optical disk, etc.; and a communication unit 1109, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1109 allows device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0230] The computing unit 1101 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as the data verification method. For example, in some embodiments, the data verification method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the data verification method described above can be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured to execute the data verification method in any other appropriate manner (for example, by means of firmware).
[0231] According to an embodiment of the present disclosure, a data processing method based on server-imperceptible computing is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0232] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0233] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0234] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0235] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or an LCD (liquid crystal display, monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0236] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0237] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0238] It should be noted that the serial numbers of the above-mentioned embodiments of the present disclosure are only for description and do not represent the advantages or disadvantages of the embodiments.
[0239] In the above embodiments of the present disclosure, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0240] In the several embodiments provided in this disclosure, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0241] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0242] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0243] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the various embodiments of the present disclosure. The aforementioned storage medium includes: U disk, read-only memory (ROM), random access memory (RAM), mobile hard disk, magnetic disk or optical disk, etc., various media that can store program code.
[0244] The above is only a preferred embodiment of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present disclosure. These improvements and modifications should also be regarded as within the scope of protection of the present disclosure. Industrial Applicability
[0245] The solution provided by the embodiment of the present disclosure can be applied to the process of processing multiple request functions simultaneously during data transmission. Multiple function requests from the host machine can be obtained; the target function requested by any function request to the image processor for execution is determined to obtain multiple target functions; and the execution order of the multiple target functions is determined based on the performance indicators of the multiple target functions, wherein the performance indicators of the target functions are used to characterize the performance of the target functions during execution, and the execution order is used to determine the data transmitted between the host machine's main memory and the image processor's video memory. That is, in the embodiment of the present disclosure, multiple function requests from the host machine are obtained, the target functions requested by the function requests are determined, and based on the performance indicators of the target functions, the execution order of the multiple target functions is determined, thereby achieving the technical effect of reducing the time consumed during the data transmission process and solving the technical problem of long time consumed during the data transmission process.
Claims
1. A data processing method for non-perceptual computing, wherein: include: Get multiple function requests from the host; Determine a target function that any one of the function requests requests the image processor to execute, and obtain a plurality of the target functions; Based on the performance indicators of the multiple objective functions, the execution order of the multiple objective functions is determined, wherein the performance indicators of the objective functions are used to characterize the performance of the objective functions during the execution process, and the execution order is used to determine the data transmitted between the main memory of the host machine and the video memory of the image processor.
2. The method according to claim 1, wherein: Determining the execution order of the plurality of objective functions based on the performance indicators of the plurality of objective functions comprises: Determining an evaluation result of the objective function based on the performance indicator of the objective function; Based on the evaluation result, the execution order is determined.
3. The method according to claim 2, wherein: Determining an evaluation result of the objective function based on the performance indicator of the objective function includes: Determine a first number of executions of the target function that have been completed within a historical time period, and a second number of executions of the target function that have been completed and whose execution results meet a performance indicator threshold; Based on the first execution number, the second execution number and the performance indicator, an evaluation result of the objective function at a current moment is determined.
4. The method according to claim 3, wherein: Determining the evaluation result of the objective function at a current moment based on the first execution quantity, the second execution quantity, and the performance indicator includes: Acquire a service level target in the performance indicator, wherein the service level target is used to characterize the execution status of the objective function; The evaluation result is determined based on the first execution quantity, the second execution quantity, and the service level target.
5. The method according to claim 4, wherein: Determining the evaluation result based on the first execution quantity, the second execution quantity, and the service level target in the performance indicator includes: determining a tail latency value in the service level objective; The evaluation result is determined based on the first execution number, the second execution number, and the tail delay value.
6. The method according to claim 2, wherein: Determining the execution order based on the evaluation result includes: Based on the evaluation result, the plurality of objective functions are sorted to obtain a sorting result; According to the sorting result, the plurality of objective functions are divided to obtain a first objective function group and The second objective function group; Determine an execution order of objective functions in the first objective function group and an execution order of objective functions in the second objective function group, wherein the execution order of objective functions in the first objective function group is before the execution order of objective functions in the second objective function group.
7. The method according to claim 6, wherein: According to the sorting result, the plurality of objective functions are divided to obtain the first objective function group and the second objective function group, including: Classify the objective functions that are located before the target ranking in the sorting results into the first objective function group; The objective functions located at the target ranking and after the target ranking in the sorting result are divided into the second objective function group.
8. The method according to claim 6, wherein: Based on the sorting result, the plurality of objective functions are divided to obtain the first objective function group and the second objective function group, including: In the sorting results, the objective functions whose evaluation results are less than the evaluation result threshold are divided into the first objective function group; In the sorting results, objective functions whose evaluation results are greater than or equal to the evaluation result threshold are divided into the second objective function group.
9. The method according to claim 8, wherein: The method further comprises: In response to the number of the objective functions in the first objective function group being less than a number threshold, adjusting the evaluation result threshold.
10. The method according to any one of claims 7 to 9, wherein: The execution order of the objective functions in the first objective function group is to execute the objective functions in the first objective function group in order from large to small according to the evaluation result, and the execution order of the objective functions in the second objective function group is to execute the objective functions in the second objective function group in order from small to large according to the evaluation result.
11. The method according to claim 10, wherein: The method further comprises: In the second objective function group, in response to the evaluation result of the first objective function being greater than a target result threshold, execution of the first objective function is prohibited.
12. A method for data processing based on server-unaware computing, wherein: include: Get multiple data processing requests from the host; Determine a target function that any one of the data processing requests requests the image processor to execute, and obtain a plurality of the target functions, wherein the target function is at least used for data copying or data synchronization; Based on the performance indicators of the multiple objective functions, the execution order of the multiple objective functions is determined, wherein the performance indicators of the objective functions are used to characterize the performance of the objective functions during the execution process, and the execution order is used to determine the data to be copied or synchronized between the main memory of the host machine and the video memory of the image processor.
13. An electronic device, wherein: include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method described in any one of claims 1 to 12 are implemented.
14. A computer-readable storage medium, the computer-readable storage medium comprising a stored program, wherein: When the program is running, the device where the storage medium is located is controlled to execute the steps of any one of the methods described in claims 1 to 12.
15. A processor for running a program, wherein: When the program is run, the steps of the method according to any one of claims 1 to 12 are executed.
Citation Information
Patent Citations
Performance optimization method and device of deep learning algorithm based on GPU
CN114418827A
Calculation network scheduling service method and system based on network performance comprehensive weight decision
CN116708446A
Data processing method based on non-perceptual calculation of server and electronic equipment
CN117499494A
Scheduling function calls of a transactional application programming interface (API) protocol based on argument dependencies
US20230214284A1