Chain resource decoupling method for distributed machine learning inference
By monitoring and analyzing load data in real time, and dynamically adjusting task loads to match the load characteristic threshold range, the problem of low resource utilization in distributed machine learning inference is solved, chain decoupling of resources is achieved, and system efficiency and response speed is improved.
Patent Information
- Application Number
- CN202411410063.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-10-10
AI Technical Summary
In distributed machine learning inference, the system is inefficient, especially when the context of long and shorter request processing is caused by the tight coupling of fixed resources and variable resources, low resource utilization and increased latency.
By monitoring and analyzing the current and historical load data of the inference task in real time, the task load is dynamically adjusted to match the load characteristic threshold range of each task processor, achieving chain decoupling of resources.
It improves resource utilization, optimizes system performance, enhances system flexibility and adaptability, and improves overall reasoning efficiency and response speed.
Smart Images

Figure CN119356867B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of distributed machine learning, and particularly relates to a chained resource decoupling method, device, electronic device, and computer-readable storage medium for distributed machine learning inference. Background Art
[0002] As the scale of machine learning models continues to expand, especially with the widespread application of large language models (LLMs), the number of model parameters ranges from billions to hundreds of billions, and distributed processing is required during the inference process. Large language models adopt an autoregressive generation method based on a context window, which enables them to process longer conversations, documents, and code blocks, thus significantly improving the interaction quality and application scope.
[0003] Currently, distributed inference mainly adopts two strategies: model parallelism and data parallelism. Model parallelism divides the parameters of the model among different computing devices, and each device processes a part of the model parameters. Specifically, it is further divided into tensor parallelism and pipeline parallelism. Tensor parallelism equally divides and distributes the weights of each layer to each computing device, while pipeline parallelism divides the model weights by layer. Data parallelism runs multiple model replicas on multiple computing devices, and each request is inferred by a specific model replica.
[0004] However, as the context window increases, the iteration latency also increases. Under the existing distributed strategies, each model replica serves a fixed number of requests in a fixed model parallelism mode, resulting in a tight coupling between the allocation of fixed resources and variable resources: that is, when requests with longer contexts and requests with shorter contexts are processed mixedly, the variable resource requirements of the batch tasks depend on the longest context request, resulting in insufficient resources for requests with longer contexts and low resource utilization for requests with shorter contexts. This increases the overall system latency and reduces resource utilization. Summary of the Invention
[0005] This application aims to provide a chained resource decoupling method, device, electronic device, and computer-readable storage medium for distributed machine learning inference, at least solving the problem of low system efficiency during the distributed machine learning inference process.
[0006] In a first aspect, an embodiment of this application also discloses a chained resource decoupling method for distributed machine learning inference, which is applied to a task processor of a distributed machine learning system and includes:
[0007] Obtain the current load data of the inference task being executed; the current load data is used to characterize the task load generated by the inference task;
[0008] Send the current load data to the task dispatcher of the distributed machine learning system, so that after obtaining the historical load data and the current load data sent from all the task processors in the distributed system, the task dispatcher determines the data processing decision of the distributed machine learning system according to all the current load data and the historical load data, and sends the data processing decision to each task processor respectively; the data processing decision includes the load characteristic threshold range defined for each task processor in the distributed machine learning system; the historical load data is used to represent the task load generated by all the historical inference tasks in the distributed machine learning system.
[0009] Receive the data processing decision sent from the task dispatcher, and adjust the inference task being executed according to the data processing decision, so that the task load generated by the adjusted inference task matches the load characteristic threshold range corresponding to the task processor.
[0010] In a second aspect, an embodiment of the present application discloses a chained resource decoupling method for distributed machine learning inference, which is applied to the task dispatcher of a distributed machine learning system, and includes:
[0011] Obtain the historical load data and the current load data sent from each task processor in the distributed system; the current load data is used to represent the task load generated by the inference task being executed by the corresponding task processor; the historical load data is used to represent the task load generated by all the historical inference tasks in the distributed machine learning system.
[0012] Determine the data processing decision of the distributed machine learning system according to all the current load data and the historical load data; the data processing decision includes the load characteristic threshold range defined for each task processor in the distributed machine learning system.
[0013] Send the data processing decision to each task processor respectively, so that after receiving the data processing decision, the task processor adjusts the inference task being executed by the task processor according to the data processing decision, so that the task load generated by the adjusted inference task of the task processor matches the load characteristic threshold range corresponding to the task processor.
[0014] In a third aspect, an embodiment of the present application further discloses a chained resource decoupling device for distributed machine learning inference, which is applied to the task processor of a distributed machine learning system, and includes:
[0015] A monitoring module, configured to obtain current load data of an ongoing inference task; the current load data is used to characterize the task load generated by the inference task.
[0016] A monitoring data sending module, configured to send the current load data to a task dispatcher of a distributed machine learning system, so that after obtaining historical load data and the current load data sent from all task processors in the distributed system, the task dispatcher determines a data processing decision of the distributed machine learning system according to all the current load data and the historical load data, and sends the data processing decision to each task processor respectively; the data processing decision includes a load characteristic threshold range defined for each task processor in the distributed machine learning system; the historical load data is used to characterize the task load generated by all historical inference tasks in the distributed machine learning system.
[0017] A decision execution module, configured to receive the data processing decision sent from the task dispatcher, and adjust the ongoing inference task according to the data processing decision, so that the task load generated by the adjusted inference task matches the load characteristic threshold range corresponding to the task processor.
[0018] In a fourth aspect, an embodiment of the present application further discloses a chained resource decoupling device for distributed machine learning inference, which is applied to a task dispatcher of a distributed machine learning system, and includes:
[0019] A monitoring data acquisition module, configured to obtain historical load data and the current load data sent from each task processor in the distributed system; the current load data is used to characterize the task load generated by the ongoing inference task of the corresponding task processor; the historical load data is used to characterize the task load generated by all historical inference tasks in the distributed machine learning system.
[0020] A decision module, configured to determine a data processing decision of the distributed machine learning system according to all the current load data and the historical load data; the data processing decision includes a load characteristic threshold range defined for each task processor in the distributed machine learning system.
[0021] A decision sending module, configured to send the data processing decision to each task processor respectively, so that after receiving the data processing decision, the task processor adjusts the ongoing inference task according to the data processing decision, so that the task load generated by the adjusted inference task of the task processor matches the load characteristic threshold range corresponding to the task processor.
[0022] Fifth aspect, an embodiment of the present application further discloses an electronic device, including a processor and a memory. The memory stores a program or instructions that can run on the processor. When the program or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.
[0023] Sixth aspect, an embodiment of the present application further discloses a readable storage medium. A program or instructions are stored on the readable storage medium. When the program or instructions are executed by a processor, the steps of the method described in the first aspect are implemented.
[0024] In summary, in the embodiment of the present application, by combining and analyzing the load data and historical data of the real-time monitoring inference task, a more optimized resource allocation decision is comprehensively made to predict the future load trend, and a reasonable resource allocation strategy is formulated to adjust the current inference task load, ensuring that the load of each task processor runs within a reasonable range, avoiding overload or resource idleness, improving resource utilization, realizing dynamic scheduling of resources, improving the flexibility and adaptability of the system, and being able to better respond to requests with different voluntary requirements. Therefore, based on the method of the embodiment of the present application, by dynamically monitoring and adjusting the task load, the system can allocate resources more efficiently, improve the overall inference efficiency and response speed. The chained decoupling of resources in the distributed machine learning inference process is realized, and the problem of low system efficiency in the distributed machine learning inference process is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In the drawings:
[0026] Figure 1 is the architecture diagram of the distributed machine learning system in the embodiment of the present application;
[0027] Figure 2 is the flowchart of the steps of a chained resource decoupling method for distributed machine learning inference provided by the embodiment of the present application;
[0028] Figure 3 is the flowchart of the steps of another chained resource decoupling method for distributed machine learning inference provided by the embodiment of the present application;
[0029] Figure 4 is the data flow process based on the method of the embodiment of the present application;
[0030] Figure 5 is the overall steps of the method of the embodiment of the present application;
[0031] Figure 6 is the block diagram of a chained resource decoupling device for distributed machine learning inference provided by the embodiment of the present application;
[0032] Figure 7It is a block diagram of another chain resource decoupling device for distributed machine learning inference provided by an embodiment of the present application;
[0033] Figure 8 It is a block diagram of an electronic device according to an embodiment provided by an embodiment of the present application;
[0034] Figure 9 It is a block diagram of an electronic device according to another embodiment provided by an embodiment of the present application. Detailed implementation manners
[0035] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0036] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally represents an "or" relationship between the associated objects before and after.
[0037] Figure 1 It is a distributed machine learning system in an embodiment of the present application. The number of task processors is only a reference, and in actual situations, it may be more or less. Among them, the task dispatcher is connected to each task processor one by one, and the task processors are connected to each other in pairs. Based on Figure 1 the system architecture, Figure 2 It is a chain resource decoupling method for distributed machine learning inference provided in this embodiment.
[0038] Among them, step 101, step 102, and step 106 are applied to the task processors of the distributed machine learning system; step 103, step 104, and step 105 are applied to the task dispatcher of the distributed machine learning system.
[0039] As Figure 2 shown, a chain resource decoupling method for distributed machine learning inference provided in this embodiment specifically includes the following steps:
[0040] Step 101, obtain the current load data of the inference task being executed.
[0041] Among them, the current load data is used to characterize the task load generated by the inference task.
[0042] In some embodiments of the present application, the current load data of the inference task being executed is obtained. The current load data refers to the data used to characterize the task load generated by the inference task. These data include, but are not limited to, the number of requests processed by the task processor within a specific time window, the context length of each request, the iteration arrival rate (the number of iterations arriving per unit time), and the output length of the completed requests, etc. By collecting these data, the current load situation of the task processor can be understood, providing a basis for subsequent load prediction and resource allocation. It should be emphasized that due to the persistence of distributed machine learning inference, when the inference task involved in the present application is being executed, it is bounded by the time window of the task, that is, all inference tasks within the current time window will be recognized as the inference tasks being executed.
[0043] For example, in a distributed machine learning system, assume that task processor A is processing multiple inference tasks. The runtime monitoring module will count the context length and iteration arrival rate of all requests processed by task processor A within a time window. For example, in the past 10 minutes, the task processor processed two requests. The context length of ten requests increased from 50 to 150, and the context length of another ten requests increased from 100 to 200. Then the iteration arrival rate of the context length of 50 - 100 is 1 time per minute, the iteration arrival rate of the context length of 100 - 150 is 2 times per minute, and the iteration arrival rate of the context length of 150 - 200 is 1 time per minute. This represents that there are several iterations within each time window, so each context length has a corresponding number of iterations. The runtime monitoring module also records the output length of each request and summarizes these data into a load histogram, which is submitted to the task dispatcher. Through these current load data, the system can accurately understand the load situation of task processor A, providing a basis for subsequent load prediction and resource allocation.
[0044] Step 102, send the current load data to the task dispatcher of the distributed machine learning system.
[0045] Among them, after obtaining the historical load data and the current load data sent from all task processors in the distributed system, the task dispatcher will determine the data processing decision of the distributed machine learning system based on all the current load data and the historical load data, and send the data processing decision to each task processor respectively.
[0046] Among them, the data processing decision includes the range of load characteristic thresholds defined for each task processor in the distributed machine learning system; the historical load data is used to characterize the task loads generated by all historical inference tasks in the distributed machine learning system.
[0047] In some embodiments of the present application, the current load data will be sent to the task dispatcher of the distributed machine learning system. After collecting the current load data, the task processor will send this data to the task dispatcher. After receiving the current load data and the historical load data from all task processors, the task dispatcher will comprehensively analyze this data to determine the overall data processing decision of the system. The data processing decision includes the range of load characteristic thresholds defined for each task processor to ensure that each task processor operates within a reasonable load range and avoid overload or resource idleness.
[0048] For example, in a distributed machine learning system, assume that task processor A and task processor B respectively collect the current load data of the system. Each task processor sends this current load data to the task dispatcher. The task dispatcher also has the total historical load data in the distributed machine learning system, such as the average load of task processor A in the past hour. The task dispatcher comprehensively analyzes this data and determines that in the next time window, the range of load characteristic thresholds for task processor A is requests with a character length of 800 to 1200, and the range of load characteristic thresholds for task processor B is requests with a character length of 1201 to 1600. The task dispatcher sends these data processing decisions back to each task processor so that they can adjust the current inference tasks according to the new range of load characteristic thresholds.
[0049] Step 103, obtain the historical load data and the current load data sent from each task processor in the distributed system.
[0050] Among them, the current load data is used to characterize the task load generated by the inference task being executed by the corresponding task processor; the historical load data is used to characterize the task loads generated by all historical inference tasks in the distributed machine learning system.
[0051] In some embodiments of the present application, the historical load data and the current load data sent from each task processor in the distributed system will be obtained. The current load data is used to characterize the task load generated by the inference task being executed by the current task processor, while the historical load data is used to characterize the task loads generated by the inference tasks executed in the past in the entire distributed machine learning system. By combining these two types of data, the system can more comprehensively understand the load situation of the task processors and provide a basis for subsequent load prediction and resource allocation.
[0052] For example, in a distributed machine learning system, assume that task processor A and task processor B have respectively collected the current load data of the system. Each task processor sends this current load data to the task dispatcher. Meanwhile, the task dispatcher also obtains the historical load data of the distributed machine learning system, that is, within a past time window, the cumulative iteration arrival rate for each context length, and the average output length requested for each context length. By combining the current load data and the historical load data, the task dispatcher can more accurately predict the future load trend and formulate a reasonable resource allocation strategy.
[0053] Step 104, determine the data processing decision of the distributed machine learning system according to all the current load data and the historical load data.
[0054] Among them, the data processing decision includes the range of load characteristic thresholds defined for each task processor in the distributed machine learning system.
[0055] In some embodiments of the present application, the data processing decision of the distributed machine learning system will be determined according to all the current load data and the historical load data. The data processing decision includes the range of load characteristic thresholds defined for each task processor to ensure that each task processor operates within a reasonable load range. By comprehensively analyzing the current load data and the historical load data, the system can predict the future load trend and formulate a reasonable resource allocation strategy, thereby optimizing the system performance and avoiding overloading or resource idling.
[0056] For example, in a distributed machine learning system, the task dispatcher receives the current load data and the historical load data from task processor A and task processor B. Task processor A has processed requests with a character length of 1000 in the current time window, and the historical data shows that its average load in the past hour was requests with a character length of 900. Task processor B has processed requests with a character length of 800 in the current time window, and the historical data shows that its average load in the past hour was requests with a character length of 750. The task dispatcher comprehensively analyzes these data and predicts that in the next time window, the load of task processor A may increase to requests with a character length of 1100 per minute, while the load of task processor B may remain at around requests with a character length of 800 per minute. Based on these predictions, the task dispatcher determines the data processing decision. The range of load characteristic thresholds defined for task processor A is from 800 to 1200 requests, and the range of load characteristic thresholds defined for task processor B is from 1201 to 1600 requests. The task dispatcher sends these data processing decisions back to each task processor so that they can adjust the current inference tasks according to the new range of load characteristic thresholds.
[0057] Step 105: Send the data processing decisions to each task processor respectively.
[0058] Among them, after receiving the data processing decision, the task processor will adjust the inference task in execution according to the data processing decision, so that the task load generated by the adjusted inference task of the task processor matches the load characteristic threshold range corresponding to the task processor.
[0059] In some embodiments of the present application, the data processing decisions will be sent to each task processor respectively. After determining the overall data processing decisions of the system, the task dispatcher will send these decisions to each task processor. After receiving the data processing decision, the task processor will compare and adjust it according to its current load data to ensure that the task load generated by the currently executed inference task runs within the load characteristic threshold range corresponding to the task processor. This process helps to optimize the system performance and avoid overloading of task processors or idling of resources.
[0060] For example, in a distributed machine learning system, the task dispatcher has determined the load characteristic threshold ranges of task processor A and task processor B and sent these data processing decisions to them. The decision received by task processor A is that its load characteristic threshold range is requests with 800 to 1200 token counts, and the load characteristic threshold range of task processor B is requests with 1201 to 1600 character lengths. After receiving the data processing decision, task processor A checks its current load data and finds that the context length of a currently processed request has changed from 1200 character lengths to 1201, exceeding the threshold range. So, task processor A adjusts the current inference task according to the data processing decision and transfers the inference task to other task processors to bring its load back within 1200 character lengths. Similarly, after receiving the data processing decision, task processor B finds that the current character length count being processed is 1201, which is within the threshold range, so no adjustment is needed. In this way, the system ensures that the load of each task processor runs within a reasonable range, improving the efficiency and stability of the overall system.
[0061] Step 106: Receive the data processing decision sent from the task dispatcher and adjust the inference task in execution according to the data processing decision, so that the task load generated by the adjusted inference task matches the load characteristic threshold range corresponding to the task processor.
[0062] In some embodiments of the present application, the task processor will receive the data processing decision sent from the task dispatcher, and adjust the ongoing inference task according to the current load data and the data processing decision. The purpose of the adjustment is to ensure that the task load generated by the adjusted inference task runs within the load characteristic threshold range corresponding to the task processor. This process helps to optimize the system performance, avoid overloading of the task processor or idling of resources, thereby improving the overall efficiency and stability of the system.
[0063] For example, in a distributed machine learning system, task processor A receives the data processing decision sent by the task dispatcher, and the decision stipulates that its load characteristic threshold range is requests with a character length of 800 to 1200 characters. Task processor A checks the current load data and finds that the context length of a certain request being processed has changed from 1200 character lengths to 1201, and at this time the character length exceeds the threshold range. To meet the requirements of the data processing decision, task processor A needs to adjust the ongoing inference task. Task processor A first identifies which requests can be transferred to other task processors. Then, it migrates the metadata of these requests and the associated key-value cache to other task processors to reduce its own load. After the adjustment, the context lengths of all tasks on task processor A are within the range of 800 to 1200, meeting the requirements of the data processing decision. In this way, task processor A ensures that its load runs within a reasonable range, avoids overloading, and improves the stability and response speed of the system.
[0064] In summary, in the embodiments of the present application, by combining the analysis of the load data and historical data of the real-time monitored inference task, more optimized resource allocation decisions are comprehensively made to predict future load trends, and a reasonable resource allocation strategy is formulated to adjust the current inference task load, ensuring that the load of each task processor runs within a reasonable range, avoiding overloading or resource idling, improving resource utilization, realizing dynamic scheduling of resources, improving the flexibility and adaptability of the system, and being able to better handle requests with different voluntary requirements. Thus, based on the method of the embodiments of the present application, by dynamically monitoring and adjusting the task load, the system can allocate resources more efficiently, improve the overall inference efficiency and response speed. It realizes the chain decoupling of resources in the distributed machine learning inference process and solves the problem of low system efficiency in the distributed machine learning inference process.
[0065] Figure 3 Another chain resource decoupling method for distributed machine learning inference provided by the application embodiments.
[0066] Among them, steps 201, 202, and 206 are applied to the task processors of the distributed machine learning system, and steps 203, 204, and 205 are applied to the task dispatcher of the distributed machine learning system.
[0067] Specifically, it includes the following steps:
[0068] Step 201: Obtain the current load data of the inference task being executed.
[0069] Among them, the current load data is used to characterize the task load generated by the inference task.
[0070] The method shown in this step has been described in step 101 and will not be elaborated here.
[0071] Step 202: Send the current load data to the task dispatcher of the distributed machine learning system.
[0072] Among them, after obtaining the historical load data and the current load data sent from all the task processors in the distributed system, the task dispatcher will determine the data processing decision of the distributed machine learning system based on all the current load data and the historical load data, and send the data processing decision to each task processor respectively.
[0073] Among them, the data processing decision includes the load characteristic threshold range defined for each task processor in the distributed machine learning system; the historical load data is used to characterize the task load generated by all the historical inference tasks in the distributed machine learning system.
[0074] The method shown in this step has been described in step 102 and will not be elaborated here.
[0075] Step 203: Obtain the historical load data and the current load data sent from each task processor in the distributed system.
[0076] Among them, the current load data is used to characterize the task load generated by the inference task being executed by the corresponding task processor; the historical load data is used to characterize the task load generated by all the historical inference tasks in the distributed machine learning system.
[0077] The method shown in this step has been described in step 103 and will not be elaborated here.
[0078] Step 204: Determine the data processing decision of the distributed machine learning system based on all the current load data and the historical load data.
[0079] Among them, the data processing decision includes the load characteristic threshold range defined for each task processor in the distributed machine learning system.
[0080] The method shown in this step has been described in step 104 and will not be elaborated here.
[0081] Optionally, the task load includes the text length of the data processed by the corresponding task processor, and the historical load data includes the historical length of the data processed by the distributed machine learning system; step 204 includes the following sub-steps:
[0082] Sub-step 2041, determine the text length interval range corresponding to the maximum text length, and divide the text length interval range into multiple consecutive and equally wide text length interval sub-ranges according to the preset number of intervals.
[0083] In some embodiments of the present application, the text length interval range corresponding to the maximum text length will be determined, and the text length interval range will be divided into multiple consecutive and equally wide text length interval sub-ranges according to the preset number of intervals. The text length interval range refers to the maximum text length configuration of the machine learning model. By dividing this range into multiple equally wide sub-ranges, different lengths of text data can be analyzed and processed more carefully, thereby optimizing the system's resource allocation and load management.
[0084] For example, in a distributed machine learning system, assume that the text length range processed by task processor A and task processor B ranges from 50 to 500. To perform more refined load management, the system presets to divide the text length interval range into 10 equally wide sub-ranges. First, determine that the maximum text length is 500, so the text length interval range is from 0 to 500. Next, divide this interval range into 10 equally wide sub-ranges, and the width of each sub-range is (500 - 0) / 10 = 50. Thus, the text length interval range is divided into the following sub-ranges: 0-50, 51-100, 101-150, …, 401-450, 451-500. Through this division, the system can analyze the load conditions of different text lengths more precisely, providing a basis for subsequent resource allocation and optimization.
[0085] Sub-step 2042, use the number of text length interval sub-ranges corresponding to the test interval extracted from the text length interval range as the traversal parameter, and the text length and historical length as the statistical parameters, and calculate the standardized iteration delay evaluation value corresponding to the test interval when the traversal parameter takes different values.
[0086] Among them, the text length interval sub-ranges in the test interval are continuous; the first text length interval sub-range in the continuous text length interval sub-ranges in the test interval is the first text length interval sub-range in the remaining text length interval sub-ranges in the text length interval range.
[0087] In some embodiments of the present application, the number of sub - ranges of the text length interval corresponding to the test interval extracted from the text length interval range is used as the traversal parameter, and the text length and the historical length are used as statistical parameters to calculate the standardized iterative delay evaluation values corresponding to the test interval when the traversal parameter takes different values. The sub - ranges of the text length interval in the test interval must be continuous, and the first sub - range of the text length interval is the first remaining sub - range of the text length interval range. In this way, the system performance under different text length interval configurations can be evaluated to find the optimal configuration scheme.
[0088] For example, in a distributed machine learning system, assume that the text length interval range has been divided into 10 equal - width sub - ranges. To evaluate the system performance under different configurations, the system extracts test intervals from these sub - ranges to calculate the standardized iterative delay evaluation values of these test intervals. In this way, the system can evaluate the performance under different text length interval configurations, find the optimal configuration scheme to optimize the system's resource allocation and load management.
[0089] Sub - step 2043: Use the test interval corresponding to the minimum standardized iterative delay evaluation value among all the standardized iterative delay evaluation values as the target load feature threshold range.
[0090] In some embodiments of the present application, the test interval corresponding to the minimum standardized iterative delay evaluation value among all the standardized iterative delay evaluation values is used as the target load feature threshold range. The standardized iterative delay evaluation value is an index for evaluating the system performance under different test interval configurations. By comparing the standardized iterative delay evaluation values of all test intervals, the test interval corresponding to the minimum evaluation value is selected as the target load feature threshold range to optimize the system performance and ensure the efficient use of resources.
[0091] For example, in a distributed machine learning system, assume that the standardized iterative delay evaluation values of multiple test intervals have been calculated. For example, assume that during the test, it is found that the standardized iterative delay evaluation value is the minimum when the test interval contains three sub - ranges of the text length interval. Then, the test interval composed of the first three sub - ranges of the text length interval is selected as the target load feature threshold range.
[0092] Sub - step 2044: Assign the target load feature threshold range to a task processor that has not been assigned a load feature threshold range.
[0093] In some embodiments of the present application, a target load feature threshold range is assigned to a task processor that has not been assigned a load feature threshold range. The target load feature threshold range refers to the load range determined in the previous step that can optimize the system performance. Assigning this range to a task processor that has not been assigned a load feature threshold range can ensure that each task processor in the system has a clear load management goal, thereby optimizing resource utilization and system performance.
[0094] For example, in a distributed machine learning system, assume that the target load feature threshold range determined in the previous step is a test range composed of the first three text length interval sub-ranges. There are multiple task processors in the system, and task processor A has not been assigned a load feature threshold range. Therefore, the system assigns the target load feature threshold range to task processor A, making it responsible for managing and processing the load within this range. Through this assignment, task processor A can perform resource management and task processing according to the preset load feature threshold range, ensuring that the system operates under optimal configuration and improving overall performance and resource utilization.
[0095] Sub-step 2045: Remove the target load feature threshold range from the text length interval range to update the text length interval range, and return to sub-step 2042.
[0096] In some embodiments of the present application, the assigned target load feature threshold range is removed from the text length interval range, the text length interval range is updated, and then the allocation process is restarted from the process of sub-step 2042. By removing the assigned target load feature threshold range, it can be ensured that the remaining text length interval range only contains unassigned parts, so as to continue optimization and allocation in subsequent steps until all intervals are reasonably assigned to task processors.
[0097] For example, in a distributed machine learning system, assume that the text length interval range of 0 - 1200 has been divided into 3 equally wide sub - ranges. Through the previous step, the target load feature threshold range is determined to be the test interval composed of the first three text length interval sub - ranges, and has been assigned to task processor A for the text length range from 0 to 400. Then the system removes the target load feature threshold range of 0 - 400 from the text length interval range, and the updated text length interval range only contains the latter two text length interval sub - ranges (401 to 800 and 801 to 1200). Then, the system will return to sub - step 2042, continue to extract a new test interval from the remaining text length interval range, and conduct a standardized iterative latency evaluation. At this time, since task processor A has already been assigned, subsequent assignments will be carried out between task processors B and C until all intervals are reasonably assigned to task processors. In this way, the system can gradually optimize resource allocation to ensure that the load of each task processor runs within a reasonable range.
[0098] Optionally, in the above - mentioned sub - step, the standardized iterative latency evaluation value is obtained according to the following formula:
[0099] ,
[0100] where, represents the set composed of the historical lengths of multiple task processors, represents the set composed of the text lengths of multiple task processors; represents the interval of the interval iterative latency; represents in the inference task, the output length of the request ; represents in the historical inference task, the estimated value of the output length of the request ;
[0101] At this time, the formula for the target load feature threshold range is:
[0102] ;
[0103] where, is the traversal parameter corresponding to the minimum standardized iterative latency evaluation value, n represents the number of task processors with unassigned load feature threshold ranges, and m is the number of text length interval sub - ranges that have not yet been included in the load feature threshold range.
[0104] In the solution of the embodiment of the present application, in order to find the optimal interval configuration, is assigned to a task processor, and Allocate to the other n - 1 task processors. If n - 1 > M - p is satisfied, there are some task processors that do not have corresponding sub - ranges of the text length interval to process, and the iteration delay of these task processors is recorded as 0. By trying all possible values of p, the optimal value that minimizes the objective can be found. . The correctness of this algorithm is guaranteed by mathematical induction.
[0105] Optionally, in the above sub - step, the interval iteration delay is obtained according to the following formula:
[0106] In the case of p = 1:
[0107] ;
[0108] where represents the text length, represents the preset length threshold for the text length, and the parameters a, b, and c respectively represent the preset model coefficients;
[0109] In the case of p > 1:
[0110] ;
[0111] where represents the probability that the data corresponding to the maximum text length belongs to the test interval :
[0112] ,
[0113] λ represents the total number of inference tasks in the same batch of inference tasks.
[0114] Since in the solution of the embodiment of the present application, the formal system optimization objective is: given N task processors, equally divide the context length interval into M sub - ranges of text length intervals. Let b i represent the i - th sub - range of the text length interval, represent the interval of the text length interval sub - range . Configuring a chained interval for the task processor is equivalent to allocating these M intervals to N task processors. Using to represent the interval allocated to the i - th task processor, so represents a chained interval configuration scheme. Let represent the iteration delay of iteration e running in the i - th task processor. Among them, e represents an iteration, and q(e) represents the request ID to which e belongs.
[0115] Then its normalized iteration delay is , where is the output length of the request , is the delay of iteration e affected by the context length. For request q, the sum of all its iterations divided by a constant is the average iteration delay of the system. This module obtains the iterations arriving within the time window from the load histogram, and denotes the sets of arriving iterations of completed and uncompleted requests as E and U respectively. For the iterations in U, this paper uses the expected output length based on history for estimation. Then, the interval configuration problem can be formalized as:
[0116] ,
[0117] In the above formula, there is .
[0118] Therefore, to solve the interval configuration problem, it is necessary to estimate the delay of the iterations running on a specific task processor, that is, . The interval iteration delay depends on the maximum context length in the batch. Let represent the average context length in the text length interval sub-range . The iteration delay in the interval containing only one text length interval sub-range is jointly affected by the fixed resource requirements of inference and the variable resource requirements related to the context length. Among them, the first case represents the constant overhead when the context length is small (less than or equal to θ), denoted as c, and the second case represents the linear increase in delay when the context length is large (greater than θ), and the delay has a linear relationship with the context length (with coefficients a and b). In the embodiments of this application, specific parameters a, b, c, and threshold θ can be obtained by offline analysis of each model, and the specific values of the relevant data are determined by the specific distributed machine learning system.
[0119] Step 205, send the data processing decisions to each task processor respectively.
[0120] Among them, after receiving the data processing decision, the task processor will adjust the inference task being executed according to the data processing decision, so that the task load generated by the adjusted inference task of the task processor matches the load characteristic threshold range corresponding to the task processor.
[0121] The method shown in this step has been described in step 105 and will not be elaborated here.
[0122] Step 206, receive the data processing decision sent from the task dispatcher, and adjust the inference task being executed according to the data processing decision, so that the task load generated by the adjusted inference task matches the load characteristic threshold range corresponding to the task processor.
[0123] The method shown in this step has been described in step 106 and will not be elaborated here.
[0124] Optionally, in order to adjust the ongoing inference task according to the current load data and data processing decisions, so that the task load generated by the adjusted inference task matches the load characteristic threshold range corresponding to the task processor, step 206 includes the following sub-steps:
[0125] Sub-step 2061, in the case where the current load data does not match the load characteristic threshold range corresponding to the task processor, determine the first target task processor among the task processors of the distributed machine learning system.
[0126] Among them, the first target task processor is the task processor corresponding to the first target load characteristic threshold range that matches the current load data in the distributed machine learning system.
[0127] In some embodiments of the present application, in the case where the current load data does not match the load characteristic threshold range corresponding to the task processor, the first target task processor will be determined among the task processors of the distributed machine learning system according to the current load data and data processing decisions. The first target task processor refers to the task processor corresponding to the first target load characteristic threshold range that matches the current load data. By determining the first target task processor, tasks that exceed the load characteristic threshold range of the current task processor can be transferred to a suitable task processor, thereby optimizing system performance and avoiding overload or resource idleness.
[0128] For example, in a distributed machine learning system, assume that the load characteristic threshold range of task processor A is requests with a character length of 800 to 1200. If the current load data indicates that the character length of the request being processed by task processor A will change from 1200 to 1201, which is beyond its load characteristic threshold range at this time. According to the current load data and data processing decisions, the system needs to determine a new task processor to receive part of the requests. The local task processor will check the load conditions of other task processors and find that the load characteristic threshold range of task processor B is 1201 to 1600 character lengths. The load of task processor B is within its threshold range. Therefore, the system determines task processor B as the first target task processor and transfers the task request from task processor A to task processor B to ensure that the load of task processor A returns to a reasonable range and optimize system performance.
[0129] Sub-step 2062, send the task content of the ongoing inference task, the task metadata of the inference task, and the key-value cache data of the inference task to the first target task processor.
[0130] Among them, the first target task processor will re - establish the inference task according to the task content of the inference task, and continue to process the task metadata of the inference task based on the key - value cache data of the inference task, so that the inference task can continue to execute.
[0131] Among them, the key - value cache data is used to record the processing progress of the inference task for the corresponding task metadata.
[0132] In some embodiments of the present application, the task content, task metadata, and key - value cache data of the inference task being executed will be sent to the first target task processor. After receiving these data, the first target task processor will re - establish the inference task according to the task content, and continue to process the task metadata using the key - value cache data to ensure that the inference task can continue to execute seamlessly. The key - value cache data is used to record the processing progress of the inference task to ensure that the progress information of the task will not be lost during the migration process.
[0133] For example, in a distributed machine - learning system, assume that the current load of task processor A exceeds its load characteristic threshold range, and part of the tasks need to be migrated to task processor B. Task processor A packages and sends the task content, task metadata, and key - value cache data of the inference task being executed to task processor B. After receiving these data, task processor B re - establishes the inference task according to the task content. For example, the task content may include an inference request for a large - language model, the task metadata contains the context information and parameter settings of the request, and the key - value cache data records the processing progress of the task. Task processor B uses these data to continue processing the inference task to ensure that the task can seamlessly continue to execute on the new processor. In this way, the system realizes the dynamic migration of tasks and load balancing, improving the overall performance and resource utilization rate.
[0134] Optionally, as the task executed on the opposite side of sub - steps 2061 and 2062, when the task processor is determined to be the first target task processor, for the task processors of the distributed machine - learning system, the method of the embodiments of the present application further includes the following additional steps:
[0135] Step 2071, obtain the task content of the inference task, the task metadata of the inference task, and the key - value cache data of the inference task sent from other task processors in the distributed machine - learning system.
[0136] Among them, the key - value cache data is used to record the processing progress of the inference task for the corresponding task metadata.
[0137] In some embodiments of the present application, when the task processor is determined to be the first target task processor, the task content, task metadata, and key-value cache data of the inference task sent from other task processors in the distributed machine learning system are obtained. The key-value cache data is used to record the processing progress of the inference task. After receiving these data, the first target task processor will re-establish the inference task according to the task content of the inference task, and continue to process the task metadata using the key-value cache data to ensure that the inference task can continue to execute seamlessly.
[0138] For example, in a distributed machine learning system, assume that task processor B is determined to be the first target task processor and needs to receive a partial inference task from task processor A. Task processor A sends the task content, task metadata, and key-value cache data of the ongoing inference task to task processor B. After receiving these data, task processor B re-establishes the inference task according to the task content. For example, the task content may include an inference request for a large language model, the task metadata contains the context information and parameter settings of the request, and the key-value cache data records the processing progress of the task. Task processor B uses these data to continue processing the inference task to ensure that the task can seamlessly continue to execute on the new processor. In this way, the system realizes the dynamic migration of tasks and load balancing, improving the overall performance and resource utilization rate.
[0139] Step 2072, re-establish the inference task according to the task content of the inference task, and continue to process the task metadata of the inference task according to the key-value cache data of the inference task, so that the inference task continues to execute.
[0140] The first target task processor is the task processor corresponding to the range of the first target load feature thresholds that matches the current load data in the distributed machine learning system.
[0141] In some embodiments of the present application, after the task processor receives the task content, task metadata, and key-value cache data of the inference task from other task processors, it will re-establish the inference task according to the task content of the inference task, and continue to process the task metadata using the key-value cache data. The key-value cache data records the processing progress of the inference task, ensuring that the progress information will not be lost during the migration process, so that the inference task can continue to execute seamlessly.
[0142] For example, in a distributed machine learning system, assume that task processor B receives inference task migration data from task processor A. The data sent by task processor A includes the task content of the inference task (e.g., an inference request for a large language model), task metadata (e.g., request context information and parameter settings), and key-value cache data (recording the processing progress of the task). After receiving this data, task processor B reconstructs the inference task based on the task content. For example, if the task content is a request to generate text, task processor B will re-initialize the inference task according to the request context information and parameter settings. Then, task processor B uses the key-value cache data to continue processing the task metadata, ensuring that the task can seamlessly continue execution on the new processor. In this way, the system realizes dynamic task migration and load balancing, improving overall performance and resource utilization.
[0143] Optionally, for the task dispatcher of the distributed machine learning system, the method of this embodiment of the present application further includes the following additional interaction steps, where the steps:
[0144] Step 2081, obtain the inference task to be executed and the task metadata of the inference task to be executed.
[0145] In some embodiments of the present application, the inference task to be executed and the task metadata of the inference task to be executed will be obtained. The inference task to be executed refers to an inference request that has not yet started to be processed, and the task metadata refers to auxiliary information related to the inference task, such as request context information, parameter settings, etc. By obtaining this information, the system can prepare for the inference task to be executed, ensuring that the task can be smoothly allocated and executed.
[0146] For example, in a distributed machine learning system, assume that a user submits a new inference request that requires generating a piece of text. The task dispatcher receives this inference task to be executed and obtains its task metadata. The task metadata includes request context information (e.g., existing text fragments), parameter settings (e.g., the length of the generated text, model configuration), etc. By obtaining this information, the task dispatcher can allocate appropriate resources for the inference task and ensure that the task can be smoothly executed.
[0147] Step 2082, determine the initial load of the inference task to be executed according to the task metadata, and determine the second target task processor among all task processors according to the initial load and the data processing decision.
[0148] Among them, the second target task processor is the task processor corresponding to the second target load feature threshold range that matches the initial load in the distributed machine learning system.
[0149] In some embodiments of the present application, the initial load of the inference task to be executed is determined based on the task metadata, and the second target task processor is determined from all the task processors according to the initial load and the data processing decision. The initial load refers to the amount of load that the inference task is expected to generate when it starts to execute. By analyzing the task metadata, the initial load can be estimated, and according to the system's load distribution policy, a suitable task processor (i.e., the second target task processor) is selected to process the task. The second target task processor refers to the task processor corresponding to the load characteristic threshold range that matches the initial load.
[0150] For example, in a distributed machine learning system, assume that the task dispatcher receives a new inference request, and the task metadata of this request includes that the length of the generated text is 500 characters. The task dispatcher estimates the initial load of this request according to the task metadata. For example, the initial load is 150 character lengths. The load distribution policy of the system stipulates that the load characteristic threshold ranges of each task processor are different. For example, the threshold range of task processor A is from 100 to 200 character lengths, and the threshold range of task processor B is from 201 to 400 character lengths. According to the initial load of 150 computing units, the task dispatcher determines that the load range of task processor A matches. Therefore, the task dispatcher selects task processor A as the second target task processor and assigns this inference task to task processor A to ensure the load balance and resource optimization of the system.
[0151] Step 2083: Send the task content and task metadata of the inference task to be executed to the second target task processor.
[0152] Among them, the second target task processor will establish the inference task to be executed to process the task metadata, and establish the key-value cache data of the inference task to be executed to record the processing progress of the inference task to be executed on the task metadata.
[0153] In some embodiments of the present application, the task content and task metadata of the inference task to be executed are sent to the second target task processor. After receiving these data, the second target task processor will establish the inference task to be executed and process it according to the task metadata. At the same time, the second target task processor will establish key-value cache data for recording the processing progress of the inference task. The key-value cache data ensures that the task can accurately track and manage its progress during execution, avoiding data loss or processing interruption.
[0154] For example, in a distributed machine learning system, the task dispatcher determines that task processor B is the second target task processor and sends the task content and task metadata of the inference task to be executed to task processor B. The task content may include a request to generate text, and the task metadata contains the context information and parameter settings of the request. After receiving this data, task processor B first creates the inference task to be executed based on the task content. For example, if the task content requires generating a text with a specified meaning, task processor B will initialize the corresponding inference task. Then, task processor B processes according to the task metadata (such as context information and parameter settings) and creates key-value cache data to record the processing progress of the task.
[0155] Optionally, as the task executed on the opposite side of steps 2081, 2082, and 2083, when the task processor is determined to be the second target task processor, for the task processor of the distributed machine learning system, the method of the embodiments of the present application further includes the following additional steps:
[0156] Step 2091, obtain the task content of the inference task to be executed sent from the task dispatcher, and the task metadata of the inference task to be executed.
[0157] In some embodiments of the present application, the task processor will obtain the task content and task metadata of the inference task to be executed sent from the task dispatcher. The task metadata may protect specific inference requests, such as the text to be generated or the computational task to be performed, as well as information related to the inference task, such as the context information of the request, parameter settings, etc. By obtaining this information, the task processor can prepare for the inference task to be executed and ensure that the task can be smoothly allocated and executed.
[0158] For example, in a distributed machine learning system, assume that the task dispatcher determines that task processor B is the second target task processor and sends the task content and task metadata of the inference task to be executed to task processor B. The task content may include a request to generate text, and the task metadata contains the context information and parameter settings of the request. After receiving this data, task processor B first creates the inference task to be executed based on the task content. For example, if the task content requires generating a text with a specified meaning, task processor B will initialize the corresponding inference task. Then, task processor B processes according to the task metadata (such as context information and parameter settings) and creates key-value cache data to record the processing progress of the task. In this way, task processor B can accurately track the execution of the inference task and ensure that the task can be successfully completed.
[0159] Step 2092: Create a to-be-executed inference task according to the task content of the to-be-executed inference task to process the task metadata of the to-be-executed inference task, and create key-value cache data for the to-be-executed inference task to record the processing progress of the to-be-executed inference task on the task metadata.
[0160] Among them, the second target task processor is the task processor corresponding to the range of second target load characteristics matching the initial load of the to-be-executed inference task in the distributed machine learning system.
[0161] In some embodiments of the present application, a to-be-executed inference task is created according to the task content of the to-be-executed inference task to process the task metadata, and key-value cache data is created to record the processing progress of the task. The second target task processor refers to the task processor that matches the initial load of the to-be-executed inference task. By creating the inference task and key-value cache data, the task processor can accurately track and manage the execution progress of the task to ensure the smooth completion of the task.
[0162] For example, in a distributed machine learning system, assume that task processor B is determined as the second target task processor and receives the task content and task metadata of the to-be-executed inference task. The task content may include a request to generate text, and the task metadata contains the context information and parameter settings of the request. Task processor B creates a to-be-executed inference task according to the task content. For example, if the task content requires generating a text with a specified meaning, task processor B initializes the corresponding inference task. Then, task processor B processes according to the task metadata (such as context information and parameter settings) and creates key-value cache data to record the processing progress of the task. The key-value cache data ensures that the task can accurately track and manage its progress during execution, avoiding data loss or processing interruption. In this way, task processor B can efficiently execute the inference task to ensure the stability and performance of the system.
[0163] Figure 4 It is the data flow mode when the distributed machine learning system is divided into a task processor cluster and a task dispatcher based on the method of the embodiments of the present application:
[0164] Step S1: Current load data transmission: The operation monitoring module in the task processor cluster monitors the current load situation and transmits the current load data to the task dispatcher. This step ensures that the task dispatcher can understand the load status of each task processor in real time.
[0165] Step S2: Data Processing and Decision-Making: After receiving the current load data, the task dispatcher predicts the future load situation through the load prediction module, and the interval configuration module generates system optimization objectives according to the prediction results through chained interval configuration. The task dispatcher makes decision instructions based on this data.
[0166] Step S3: Task Dispatching and Migration: According to the decision instructions, the task dispatcher distributes the task metadata and task content submitted by the user to the appropriate task processors. If the task needs to be migrated, the task dispatcher also controls the migration of the task metadata.
[0167] Step S4: Task Execution and Data Reading / Writing: After receiving the task, the task processor executes the inference task through the distributed inference engine and uses the key-value cache to record the processing progress of the task. The task processor performs data reading and writing operations during the execution process to ensure the smooth completion of the task.
[0168] As Figure 5 shown, this solution can generally be summarized into the following three major steps:
[0169] Step R1: Load Monitoring and Prediction: The operation monitoring module in the task processor cluster collects the current load data and sends it to the task dispatcher. The task dispatcher uses this data for load prediction and evaluates the future load situation to make corresponding resource allocation decisions.
[0170] Step R2: Data Processing Decision Configuration: Based on the current load data and prediction results, the task dispatcher performs data processing and decision configuration. Specifically, the task dispatcher sets the chained interval configuration policy through the interval configuration module and generates system optimization objectives. Then, the task dispatcher adjusts the load distribution of the task processors according to these decision-making information to ensure the efficient operation of the system.
[0171] Step R3: Task Dispatching and Migration: According to the decision configuration, the task dispatcher distributes the tasks submitted by the user to the appropriate task processors. If the load of the task processor exceeds the threshold range, the task dispatcher also controls the migration of the task and transfers some tasks to other task processors to achieve load balancing and resource optimization.
[0172] In summary, in the embodiments of the present application, by combining and analyzing the load data and historical data of real-time monitoring inference tasks, more optimized resource allocation decisions are comprehensively made to predict future load trends, and reasonable resource allocation strategies are formulated to adjust the current inference task load, ensuring that the load of each task processor runs within a reasonable range, avoiding overload or resource idleness, improving resource utilization, realizing dynamic scheduling of resources, enhancing the flexibility and adaptability of the system, and being able to better handle requests with different voluntary requirements. Thus, based on the method of the embodiments of the present application, by dynamically monitoring and adjusting the task load, the system can allocate resources more efficiently, improve the overall inference efficiency and response speed. The chained decoupling of resources in the distributed machine learning inference process is achieved, and the problem of low system efficiency in the distributed machine learning inference process is solved.
[0173] Reference Figure 6 , which shows a chained resource decoupling device 30 for distributed machine learning inference provided by the embodiments of the present application, applied to a task processor of a distributed machine learning system, including:
[0174] A monitoring module 301, configured to obtain the current load data of the inference task being executed; the current load data is used to characterize the task load generated by the inference task;
[0175] A monitoring data sending module 302, configured to send the current load data to the task dispatcher of the distributed machine learning system, so that after obtaining the historical load data and the current load data sent from all task processors in the distributed system, the task dispatcher determines the data processing decision of the distributed machine learning system according to all the current load data and the historical load data, and sends the data processing decision to each task processor respectively; the data processing decision includes the load characteristic threshold range defined for each task processor in the distributed machine learning system; the historical load data is used to characterize the task load generated by all historical inference tasks in the distributed machine learning system;
[0176] A decision execution module 303, configured to receive the data processing decision sent from the task dispatcher, and adjust the inference task being executed according to the data processing decision, so that the task load generated by the adjusted inference task matches the load characteristic threshold range corresponding to the task processor.
[0177] Optionally, the decision execution module 303 includes:
[0178] The first target sub-module is used to determine a first target task processor in the task processors of the distributed machine learning system when the current load data does not match the load feature threshold range corresponding to the task processor; the first target task processor is the task processor corresponding to the first target load feature threshold range that matches the current load data in the distributed machine learning system.
[0179] The first sending sub-module is used to send the task content of the inference task being executed, the task metadata of the inference task, and the key-value cache data of the inference task to the first target task processor, so that the first target task processor re-establishes the inference task according to the task content of the inference task, and continues to process the task metadata of the inference task according to the key-value cache data of the inference task, so that the inference task continues to execute; the key-value cache data is used to record the processing progress of the inference task for the corresponding task metadata.
[0180] Optionally, when the task processor is determined to be the first target task processor, the chained resource decoupling device 30 for distributed machine learning inference further includes:
[0181] The first acquisition module is used to acquire the task content of the inference task, the task metadata of the inference task, and the key-value cache data of the inference task sent from other task processors in the distributed machine learning system; the key-value cache data is used to record the processing progress of the inference task for the corresponding task metadata.
[0182] The first execution module is used to re-establish the inference task according to the task content of the inference task, and continue to process the task metadata of the inference task according to the key-value cache data of the inference task, so that the inference task continues to execute; the first target task processor is the task processor corresponding to the first target load feature threshold range that matches the current load data in the distributed machine learning system.
[0183] Optionally, when the task processor is determined to be the second target task processor, the chained resource decoupling device 30 for distributed machine learning inference further includes:
[0184] The second acquisition module is used to acquire the task content of the inference task to be executed and the task metadata of the inference task to be executed sent from the task dispatcher.
[0185] A second execution module, configured to establish a to-be-executed inference task according to the task content of the to-be-executed inference task, process the task metadata of the to-be-executed inference task, and establish key-value cache data of the to-be-executed inference task to record the processing progress of the to-be-executed inference task on the task metadata; the second target task processor is a task processor corresponding to the second target load feature threshold range matched by the initial load of the to-be-executed inference task in the distributed machine learning system.
[0186] Reference Figure 7 , which shows a chained resource decoupling device 40 for distributed machine learning inference provided by an embodiment of the present application, applied to a task dispatcher of a distributed machine learning system, including:
[0187] A monitoring data acquisition module 401, configured to acquire historical load data and current load data sent from each task processor in the distributed system; the current load data is used to characterize the task load generated by the inference task being executed by the corresponding task processor; the historical load data is used to characterize the task load generated by all historical inference tasks in the distributed machine learning system;
[0188] A decision-making module 402, configured to determine a data processing decision for the distributed machine learning system according to all the current load data and the historical load data; the data processing decision includes the load feature threshold range defined for each task processor in the distributed machine learning system;
[0189] A decision sending module 403, configured to send the data processing decision to each task processor respectively, so that after receiving the data processing decision, the task processor adjusts the inference task being executed by the task processor according to the data processing decision, so that the task load generated by the adjusted inference task of the task processor matches the load feature threshold range corresponding to the task processor.
[0190] Optionally, the task load includes the text length of the data processed by the corresponding task processor, and the historical load data includes the historical length of the data processed by the distributed machine learning system; the decision-making module 402 includes:
[0191] An interval division sub-module, configured to determine the text length interval range corresponding to the maximum text length, and divide the text length interval range into multiple consecutive and equally wide text length interval sub-ranges according to a preset number of intervals;
[0192] A traversal sub-module, which is used to take the number of sub-ranges of the text length interval corresponding to the test interval extracted from the text length interval range as the traversal parameter, the text length and the historical length as the statistical parameters, and calculate the standardized iterative delay evaluation value corresponding to the test interval when the traversal parameter takes different values; the sub-ranges of the text length interval in the test interval are continuous; the first sub-range of the text length interval in the continuous sub-ranges of the text length interval in the test interval is the first sub-range of the remaining text length interval sub-ranges in the text length interval range;
[0193] A selection sub-module, which is used to take the test interval corresponding to the minimum standardized iterative delay evaluation value among all the standardized iterative delay evaluation values as the target load feature threshold range;
[0194] An allocation sub-module, which is used to allocate the target load feature threshold range to a task processor that has not been allocated a load feature threshold range;
[0195] An iteration sub-module, which is used to remove the target load feature threshold range from the text length interval range to update the text length interval range, and then return to the step of taking the number of sub-ranges of the text length interval corresponding to the test interval extracted from the text length interval range as the traversal parameter, the text length and the historical length as the statistical parameters, and calculate the standardized iterative delay evaluation value corresponding to the test interval when the traversal parameter takes different values.
[0196] Optionally, the chained resource decoupling device 40 for distributed machine learning inference further includes:
[0197] A new task acquisition module, which is used to acquire the inference task to be executed and the task metadata of the inference task to be executed;
[0198] A second selection module, which is used to determine the initial load of the inference task to be executed according to the task metadata, and determine the second target task processor among all the task processors according to the initial load and the data processing decision; the second target task processor is the task processor corresponding to the second target load feature threshold range that matches the initial load in the distributed machine learning system;
[0199] A second sending module, which is used to send the task content and task metadata of the inference task to be executed to the second target task processor, so that the inference task to be executed is established in the second target task processor to process the task metadata, and the key-value cache data of the inference task to be executed is established to record the processing progress of the inference task to be executed for the task metadata.
[0200] In summary, in the embodiments of the present application, by combining and analyzing the load data and historical data of the real-time monitoring inference task, a more optimized resource allocation decision is comprehensively made to predict the future load trend, and a reasonable resource allocation strategy is formulated to adjust the current inference task load, ensuring that the load of each task processor runs within a reasonable range, avoiding overload or resource idleness, improving resource utilization, realizing dynamic scheduling of resources, enhancing the flexibility and adaptability of the system, and being able to better handle requests with different voluntary requirements. Thus, based on the method of the embodiments of the present application, by dynamically monitoring and adjusting the task load, the system can allocate resources more efficiently, improve the overall inference efficiency and response speed. The chained decoupling of resources in the distributed machine learning inference process is achieved, and the problem of low system efficiency in the distributed machine learning inference process is solved.
[0201] Referring to Figure 8 , the electronic device 500 may include one or more of the following components: a processing component 502, a memory 504, a power supply component 506, a multimedia component 508, an audio component 510, an input / output (I / O) interface 512, a sensor component 514, and a communication component 516.
[0202] The processing component 502 generally controls the overall operation of the electronic device 500, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 502 may include one or more processors 520 to execute instructions to complete all or part of the steps of the above-mentioned chained resource decoupling method for distributed machine learning inference. In addition, the processing component 502 may include one or more modules to facilitate the interaction between the processing component 502 and other components. For example, the processing component 502 may include a multimedia module to facilitate the interaction between the multimedia component 508 and the processing component 502.
[0203] The memory 504 is used to store various types of data to support the operation of the electronic device 500. Examples of these data include instructions for any application or method operating on the electronic device 500, contact data, phone book data, messages, pictures, multimedia, etc. The memory 504 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0204] The power supply component 506 provides power for various components of the electronic device 500. The power supply component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 500.
[0205] The multimedia component 508 includes an interface that provides an output interface between the electronic device 500 and the user. In some embodiments, the interface may include a liquid crystal display (LCD) and a touch panel (TP). If the interface includes a touch panel, the interface can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 508 includes a front camera and / or a rear camera. When the electronic device 500 is in an operating mode, such as a shooting mode or a multimedia mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0206] The audio component 510 is used to output and / or input audio signals. For example, the audio component 510 includes a microphone (MIC) that is used to receive external audio signals when the electronic device 500 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 504 or transmitted via the communication component 516. In some embodiments, the audio component 510 further includes a speaker for outputting audio signals.
[0207] The input / output (I / O) interface 512 provides an interface between the processing component 502 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.
[0208] The sensor assembly 514 includes one or more sensors for providing a status assessment of various aspects of the electronic device 500. For example, the sensor assembly 514 can detect the on / off state of the electronic device 500, the relative positioning of components, such as the display and keypad of the electronic device 500. The sensor assembly 514 can also detect a change in the position of the electronic device 500 or a component of the electronic device 500, the presence or absence of user contact with the electronic device 500, the orientation or acceleration / deceleration of the electronic device 500, and a change in the temperature of the electronic device 500. The sensor assembly 514 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 514 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 514 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0209] The communication component 516 is used to facilitate communication between the electronic device 500 and other devices in a wired or wireless manner. The electronic device 500 can access a wireless network based on communication standards, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 516 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 516 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0210] In an exemplary embodiment, the electronic device 500 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for implementing the chained resource decoupling method of distributed machine learning inference provided in the embodiments of the present application.
[0211] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 504 including instructions, and the above instructions can be executed by the processor 520 of the electronic device 500 to complete the above-mentioned chained resource decoupling method of distributed machine learning inference. For example, the non-transitory storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0212] Figure 9is a block diagram of an electronic device 600 according to another embodiment of the present invention. For example, the electronic device 600 may be provided as a server. Referring to Figure 9 , the electronic device 600 includes a processing component 622, which further includes one or more processors, and memory resources represented by a memory 632 for storing instructions executable by the processing component 622, such as application programs. The application programs stored in the memory 632 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 622 is configured to execute instructions to perform the chained resource decoupling method for distributed machine learning inference provided in the embodiments of the present application.
[0213] The electronic device 600 may also include a power component 626 configured to perform power management of the electronic device 600, a wired or wireless network interface 650 configured to connect the electronic device 600 to a network, and an input / output (I / O) interface 658. The electronic device 600 may operate based on an operating system stored in the memory 632, such as WindowsServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
[0214] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0215] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A chained resource decoupling method for distributed machine learning inference, applied to a task processor of a distributed machine learning system, characterized in that Including: Obtain the current load data of the ongoing inference task; the current load data is used to characterize the task load generated by the inference task; Send the current load data to the task dispatcher of the distributed machine learning system, so that after the task dispatcher obtains the historical load data and the current load data sent from all the task processors in the distributed machine learning system, according to all the current load data and the historical load data, determine the data processing decision of the distributed machine learning system, and send the data processing decision to each of the task processors respectively; the data processing decision includes the load characteristic threshold range defined for each task processor in the distributed machine learning system; the historical load data is used to characterize the task load generated by all the historical inference tasks in the distributed machine learning system; Receive the data processing decision sent from the task dispatcher, and adjust the ongoing inference task according to the data processing decision, so that the task load generated by the adjusted inference task matches the load characteristic threshold range corresponding to the task processor; When the task processor is determined to be the second target task processor, obtain the task content of the to-be-executed inference task sent from the task dispatcher, and the task metadata of the to-be-executed inference task; Establish the to-be-executed inference task according to the task content of the to-be-executed inference task to process the task metadata of the to-be-executed inference task, and establish the key-value cache data of the to-be-executed inference task to record the processing progress of the to-be-executed inference task for the task metadata; the second target task processor is the task processor in the distributed machine learning system corresponding to the second target load characteristic threshold range that matches the initial load of the to-be-executed inference task.
2. The method according to claim 1, wherein The adjusting the ongoing inference task according to the data processing decision so that the task load generated by the adjusted inference task matches the load characteristic threshold range corresponding to the task processor includes: When the current load data does not match the load characteristic threshold range corresponding to the task processor, determine a first target task processor among the task processors of the distributed machine learning system; the first target task processor is the task processor in the distributed machine learning system corresponding to the first target load characteristic threshold range that matches the current load data; Send the task content of the ongoing inference task, the task metadata of the inference task, and the key-value cache data of the inference task to the first target task processor, so that the first target task processor re-establishes the inference task according to the task content of the inference task, and continues to process the task metadata of the inference task according to the key-value cache data of the inference task, so that the inference task continues to execute; the key-value cache data is used to record the processing progress of the inference task for the corresponding task metadata.
3. A chained resource decoupling method for distributed machine learning inference, applied to the task dispatcher of a distributed machine learning system, characterized in that Including: Obtain historical load data and current load data sent from each task processor in the distributed machine learning system; The current load data is used to characterize the task load generated by the inference task being executed by the corresponding task processor; the historical load data is used to characterize the task load generated by all historical inference tasks in the distributed machine learning system; Determine the data processing decision of the distributed machine learning system according to all the current load data and the historical load data; the data processing decision includes the load characteristic threshold range defined for each task processor in the distributed machine learning system; Send the data processing decision to each of the task processors respectively, so that after receiving the data processing decision, the task processor adjusts the inference task being executed by the task processor according to the data processing decision, so that the task load generated by the adjusted inference task of the task processor matches the load characteristic threshold range corresponding to the task processor; Obtain the inference task to be executed and the task metadata of the inference task to be executed; Determine the initial load of the inference task to be executed according to the task metadata, and determine the second target task processor among all the task processors according to the initial load and the data processing decision; the second target task processor is the task processor corresponding to the second target load characteristic threshold range that matches the initial load in the distributed machine learning system; Send the task content and task metadata of the inference task to be executed to the second target task processor, so that the inference task to be executed is established in the second target task processor to process the task metadata, and the key-value cache data of the inference task to be executed is established to record the processing progress of the inference task to be executed on the task metadata.
4. The method according to claim 3, characterized in that, The task load includes the text length of the data processed by the corresponding task processor, and the historical load data includes the historical length of the data processed by the distributed machine learning system; the determining of the data processing decision of the distributed machine learning system according to all the current load data and the historical load data includes: Determine the text length interval range corresponding to the maximum text length, and divide the text length interval range into multiple consecutive and equally wide text length interval sub-ranges according to the preset number of intervals; Use the number of text length interval sub-ranges corresponding to the test interval extracted from the text length interval range as the traversal parameter, and the text length and the historical length as the statistical parameters, and calculate the standardized iterative delay evaluation value corresponding to the test interval when the traversal parameter takes different values; the text length interval sub-ranges in the test interval are continuous; the first text length interval sub-range in the continuous text length interval sub-ranges in the test interval is the first text length interval sub-range in the remaining text length interval sub-ranges in the text length interval range; Among all the standardized iterative delay evaluation values, the test interval corresponding to the minimum standardized iterative delay evaluation value is used as the target load feature threshold range; The target load feature threshold range is assigned to a task processor that has not been assigned a load feature threshold range; Remove the target load feature threshold range from the text length interval range to update the text length interval range, and return to the step of calculating the standardized iterative delay evaluation value corresponding to the test interval with the number of text length interval sub-ranges corresponding to the test interval extracted from the text length interval range as the traversal parameter, the text length and the historical length as the statistical parameters, when the traversal parameter takes different values.
5. A chained resource decoupling device for distributed machine learning inference, which is applied to a task processor of a distributed machine learning system, and is characterized in that, Including: A monitoring module for obtaining the current load data of the inference task being executed; the current load data is used to characterize the task load generated by the inference task; A monitoring data sending module for sending the current load data to the task dispatcher of the distributed machine learning system, so that after the task dispatcher obtains the historical load data and the current load data sent from all the task processors in the distributed machine learning system, it determines the data processing decision of the distributed machine learning system according to all the current load data and the historical load data, and sends the data processing decision to each task processor respectively; the data processing decision includes the load feature threshold range delimited for each task processor in the distributed machine learning system; the historical load data is used to characterize the task load generated by all the historical inference tasks in the distributed machine learning system; A decision execution module for receiving the data processing decision sent from the task dispatcher and adjusting the inference task being executed according to the data processing decision, so that the task load generated by the adjusted inference task matches the load feature threshold range corresponding to the task processor; A second acquisition module for obtaining the task content of the inference task to be executed and the task metadata of the inference task to be executed sent from the task dispatcher when the task processor is determined to be the second target task processor; A second execution module for establishing the inference task to be executed according to the task content of the inference task to be executed, processing the task metadata of the inference task to be executed, and establishing the key-value cache data of the inference task to be executed to record the processing progress of the inference task to be executed on the task metadata; the second target task processor is the task processor corresponding to the second target load feature threshold range that matches the initial load of the inference task to be executed in the distributed machine learning system.
6. A chained resource decoupling device for distributed machine learning inference, which is applied to a task dispatcher of a distributed machine learning system, and is characterized in that, Including: A monitoring data acquisition module, configured to acquire historical load data and current load data sent from each task processor in the distributed machine learning system; the current load data is used to characterize the task load generated by the inference task being executed by the corresponding task processor; the historical load data is used to characterize the task load generated by all historical inference tasks in the distributed machine learning system; A decision-making module, configured to determine a data processing decision for the distributed machine learning system according to all the current load data and the historical load data; the data processing decision includes a load characteristic threshold range defined for each task processor in the distributed machine learning system; A decision sending module, configured to send the data processing decision to each of the task processors respectively, so that after receiving the data processing decision, the task processor adjusts the inference task being executed by the task processor according to the data processing decision, so that the task load generated by the adjusted inference task of the task processor matches the load characteristic threshold range corresponding to the task processor; A new task acquisition module, configured to acquire an inference task to be executed and task metadata of the inference task to be executed; A second selection module, configured to determine an initial load of the inference task to be executed according to the task metadata, and determine a second target task processor from all the task processors according to the initial load and the data processing decision; the second target task processor is the task processor corresponding to the second target load characteristic threshold range that matches the initial load in the distributed machine learning system; A second sending module, configured to send the task content and task metadata of the inference task to be executed to the second target task processor, so that the inference task to be executed is established on the second target task processor to process the task metadata, and key-value cache data of the inference task to be executed is established to record the processing progress of the inference task to be executed on the task metadata.
7. An electronic device, characterized in that, Comprising: A processor and a memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the instructions to implement the chained resource decoupling method for distributed machine learning inference according to any one of claims 1 to 4.
8. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can execute the chained resource decoupling method for distributed machine learning inference according to any one of claims 1 to 4.
Citation Information
Patent Citations
Task scheduling method and device
CN113254172A
Parameter management system and related method
CN117632673A