Data reasoning method and device, equipment, storage medium and program product
By dynamically selecting server nodes and transmission paths in a distributed inference system, the problems of low resource utilization and instability in data transmission caused by static resource allocation strategies are solved, and more efficient and reliable data processing and transmission are achieved.
Patent Information
- Application Number
- CN202510096477.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-27
AI Technical Summary
The existing distributed inference system adopts a static resource allocation strategy, and cannot dynamically adjust the load and network conditions of real-time server nodes, resulting in low computing resource utilization and idle resources and wasted resources. At the same time, data transmission relies on a single path, which can easily lead to transmission interruption or failure in network congestion or server node failure, affecting the availability and stability of the system.
When the initial data entered by the user reaches the first server node, it is determined whether to receive the data based on the busyness of the node, and dynamically select other server nodes as the target node when it cannot be received. At the same time, the target transmission path is determined based on the transmission path quality between the first server node and the target server node, and the data is transmitted in priority order.
It realizes dynamic allocation of data based on real-time load and network conditions, avoids idle resources and waste, ensures priority processing of critical data, and improves system response speed and overall performance. Through multi-path transmission and dynamic path selection, the reliability and stability of data transmission are improved, and transmission failure or data loss caused by path quality problems is reduced.
Smart Images

Figure CN120046728A_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the field of artificial intelligence, and in particular, to a method, apparatus, device, storage medium, and program product for data inference. Background Art
[0002] The distributed inference server (NVIDIA Triton Inference Server) supports multiple deep learning frameworks and models, and can perform efficient distributed inference on a Graphics Processing Unit (GPU) cluster. It improves inference performance through technologies such as dynamic batching and concurrent execution, and also supports functions such as model version management and load balancing, facilitating the deployment and update of models.
[0003] In related technologies, distributed inference usually adopts a static resource allocation strategy, which cannot be dynamically adjusted according to the real-time load and network conditions of server nodes, resulting in low utilization of computing resources and situations of resource idling and waste. At the same time, the data transmission method often relies on a single path, which is prone to transmission interruption or failure in case of network congestion or server node failure, affecting the availability and stability of the system. Summary of the Invention
[0004] In view of this, at least one embodiment of this application provides a method, apparatus, device, storage medium, and program product for data inference.
[0005] The technical solution of the embodiment of this application is implemented as follows:
[0006] In a first aspect, an embodiment of this application provides a method for data inference. The method includes: when the initial data input by a user arrives at a first server node, determining whether the first server node can receive the initial data based on the busy degree of the first server node at the current moment; when the first server node cannot receive the initial data, determining a target server node from other server nodes based on the busy degree of other server nodes except the first server node at the current moment; determining a target transmission path based on the quality of the transmission path between the first server node and the target server node; transmitting the initial data of the first server node to the target server node through the target transmission path in the order of priority; and performing inference on the initial data by using the model corresponding to the target server node to obtain an inference result.
[0007] In a second aspect, an embodiment of this application provides a device for data inference. The device includes:
[0008] A first determination module, configured to determine whether the first server node can receive the initial data based on the busy degree of the first server node at the current moment when the initial data input by the user arrives at the first server node;
[0009] A second determination module, configured to determine a target server node from the other server nodes based on the busy degree of the other server nodes except the first server node at the current moment when the first server node cannot receive the initial data;
[0010] A third determination module, configured to determine a target transmission path based on the quality of the transmission path between the first server node and the target server node;
[0011] A transmission module, configured to transmit the initial data of the first server node to the target server node through the target transmission path in the order of priority;
[0012] An inference module, configured to perform inference on the initial data by using the model corresponding to the target server node to obtain an inference result.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor, where the memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements some or all of the steps in the above method.
[0015] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, it implements some or all of the steps in the above method.
[0016] In the embodiments of the present application, on the one hand, data is dynamically allocated according to the busy degrees of the first server node and other server nodes at the current moment. When the first server node is busy, the data can be quickly transferred to other available server nodes, thereby avoiding data accumulation and delay; on the other hand, by transmitting the initial data in the order of priority, it can ensure that critical or urgent data is preferentially processed, which helps to improve the response speed and overall performance of the system; on the further hand, determining the target transmission path according to the quality of the transmission path between the first server node and the target server node helps to ensure the stable transmission of data and reduce transmission failures or data losses caused by path quality problems.
[0017] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, rather than limiting the technical solutions of this application. Brief Description of the Drawings
[0018] The drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with this application and, together with the specification, are used to explain the technical solutions of this application.
[0019] Figure 1 It is a schematic flowchart of a method for reasoning about data provided by an embodiment of this application;
[0020] Figure 2 It is a schematic flowchart of a distributed reasoning and fine-tuning method provided by an embodiment of this application;
[0021] Figure 3 It is a schematic flowchart of a task distribution method provided by an embodiment of this application;
[0022] Figure 4 It is a schematic flowchart of a task processing method provided by an embodiment of this application;
[0023] Figure 5 It is a schematic diagram of an inference device for data provided by an embodiment of this application;
[0024] Figure 6 It is a schematic diagram of the hardware entity of an electronic device provided by an embodiment of this application. Detailed Description of the Embodiments
[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some, rather than all, of the embodiments of this application. The following embodiments are used to illustrate this application, but do not limit the scope of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts fall within the scope of protection of this application.
[0026] In the following description, reference is made to "some embodiments" which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0027] It should be noted that the terms "first", "second", and "third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first", "second", and "third" can be interchanged in a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0028] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art in the field to which the embodiments of the present application belong. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0029] The distributed inference server (NVIDIA Triton Inference Server, NVIDIA Triton) supports multiple deep learning frameworks and models, can perform efficient distributed inference on a Graphics Processing Unit (GPU) cluster, improve the inference performance through technologies such as dynamic batching and concurrent execution, and also supports functions such as model version management and load balancing, facilitating the deployment and update of models.
[0030] In the related art, distributed inference usually adopts a static resource allocation strategy, which cannot be dynamically adjusted according to the real-time load of server nodes and network conditions, resulting in low utilization of computing resources and situations of resource idling and waste. At the same time, the data transmission method often relies on a single path, which is prone to transmission interruption or failure in case of network congestion or server node failure, affecting the availability and stability of the system.
[0031] Based on this, the embodiments of the present application provide a method for inferring data, as Figure 1 shown, the method for inferring data may include the following steps S101 to step S105, where:
[0032] Step S101, when the initial data input by the user arrives at the first server node, determine whether the first server node can receive the initial data based on the busy degree of the first server node at the current moment;
[0033] Here, the initial data input by the user can be data submitted by the user to the system for processing or analysis. These data can be any form of information, such as text, numbers, images, audio, etc. In a distributed system, a server node is a computing unit responsible for processing data or providing services.
[0034] The first server node can be the server node where the initial data first arrives.
[0035] The degree of busyness can refer to the amount of data that the server node is currently processing or the load level. A busy server node may not have enough resources (such as a Central Processing Unit (CPU), memory, network bandwidth, etc.) to receive and process additional data.
[0036] In some embodiments, when the initial data input by the user arrives at the first server node of the system, the system first evaluates the degree of busyness of the first server node at the current moment. If the first server node is idle or relatively not busy at the current moment, that is, it has enough resources to process additional data, then the system will consider that the first server node can receive and process the initial data input by the user. However, if the first server node is busy at the current moment, that is, the resources are close to saturation, then the system will consider that the first server node cannot receive and process more data.
[0037] Step S102, in the case where the first server node cannot receive the initial data, determine a target server node from the other server nodes based on the degree of busyness of the other server nodes except the first server node at the current moment;
[0038] Here, the target server node can be a suitable server node selected by the system according to the degree of busyness of other server nodes when the first server node cannot process the initial data, for receiving and processing the initial data.
[0039] In some embodiments, when the system determines that the first server node is too busy at the current moment to receive and process the initial data input by the user, it will look for other available server nodes as alternatives. These alternative nodes refer to all other server nodes in the system that can process data except the first server node. Next, the system will evaluate the degree of busyness of these alternative nodes at the current moment. This also involves checking the resource usage of each server node, such as CPU usage, memory occupancy, network bandwidth, etc. The purpose is to find a relatively idle or low-resource-utilization node to ensure that it can efficiently receive and process the initial data. Based on the evaluation results, the system will select the most suitable node from all alternative nodes as the target server node. This target server node will be responsible for receiving the initial data forwarded from the first server node and performing subsequent processing.
[0040] Step S103, determine a target transmission path based on the quality of the transmission path between the first server node and the target server node;
[0041] Here, the quality of the transmission path may refer to the quality of the network connection involved in the process of data transmission from the first server node to the target server node. Quality factors may include network latency, packet loss rate, bandwidth, etc. A high-quality transmission path can ensure that data reaches the target server node quickly and accurately.
[0042] The target transmission path may refer to the best transmission path selected by the system from multiple possible transmission paths according to the quality of the transmission path. The initial data will be transmitted from the first server node to the target server node through the best transmission path.
[0043] In some embodiments, after determining the target server node, the system evaluates the quality of all possible transmission paths between the first server node and the target server node. Here, the quality generally refers to the performance metrics of the network connection, including but not limited to network latency (i.e., the time required for data to travel from the sender to the receiver), bandwidth (i.e., the transmission capacity of the network link), packet loss rate (i.e., the proportion of data packets lost during network transmission), and stability (i.e., the reliability of the network connection), etc. Based on the quality of the transmission path, the system selects a best transmission path as the target transmission path. This best transmission path can provide a fast enough transmission speed, sufficient bandwidth, a low enough packet loss rate, and a high enough stability to ensure that data can reach the target server node accurately and quickly.
[0044] Step S104, transmit the initial data of the first server node to the target server node through the target transmission path in the order of priority;
[0045] Here, the order of priority may refer to the priority ranking of the initial data during the transmission process. The system may sort the data according to the urgency, importance, or other factors of the initial data. Some data may be more critical or urgent and thus need to be processed and transmitted immediately; while other data may be less urgent and can be processed later.
[0046] In some embodiments, after determining the target transmission path, the system considers the order of priority of the data and, according to this order of priority, gradually sends the initial data to the target server node through the target transmission path. This means that high-priority data will be transmitted first, and low-priority data will be transmitted after high-priority data.
[0047] Step S105, use the model corresponding to the target server node to perform inference on the initial data to obtain an inference result.
[0048] Here, the model corresponding to the target server node may refer to the model running on the target server node for processing or analyzing data. The model may be a machine learning model for classifying, predicting, or regressing input data; it may also be a deep learning model for tasks such as image recognition, speech recognition, or natural language processing; or it may be other types of data processing algorithms for data mining, statistical analysis, or data visualization, etc.
[0049] The inference result may refer to the conclusion or prediction obtained after the model on the target server node processes or analyzes the input initial data.
[0050] In some embodiments, after the initial data is successfully transmitted to the target server node through the target transmission path, the target server node will use the configured model to perform a series of calculation and inference steps on the initial data, and finally generate an inference result. This inference result is the understanding and analysis of the initial data by the model, and it may be a classification label, a regression value, a prediction probability, an image recognition result, or other forms of data output.
[0051] In some embodiments, assume that the user uploads a picture as the initial data and hopes that the system can recognize the objects or scenes in the picture. In this way, the model will identify the objects or scenes in the picture through a series of calculation and inference steps (such as feature extraction, classifier judgment, etc.). After the model finishes reasoning, it will output an inference result, which may be one or more labels indicating the recognized objects or scenes in the picture. For example, if the picture is a photo of a cat, the inference result may be the label "cat".
[0052] In the embodiments of the present application, on the one hand, data is dynamically allocated according to the busy degrees of the first server node and other server nodes at the current moment. When the first server node is busy, data can be quickly transferred to other available server nodes, thus avoiding data accumulation and delay; on the other hand, by transmitting the initial data in the order of priority, it can ensure that key or urgent data is processed preferentially, which helps to improve the response speed and overall performance of the system; on the still other hand, determining the target transmission path according to the transmission path quality between the first server node and the target server node helps to ensure the stable transmission of data and reduce transmission failures or data losses caused by path quality problems.
[0053] In some embodiments, the busy degree includes the load degree and / or the congestion degree. The implementation of "determining whether the first server node can receive the initial data based on the busy degree of the first server node at the current moment" in step S101 may include the following steps S111 to step S113, where:
[0054] Step S111: Determine the load level of the first server node at the current moment based on the utilization rate of the first processor, the utilization rate of the second processor, and the data throughput of the first server node obtained at the current moment;
[0055] Here, the utilization rate of the processor can refer to the percentage of the processor on the server node that is occupied to execute tasks within a specific time period.
[0056] The data throughput can refer to the amount of data that the server node can process or transmit per unit time, which can reflect the data processing ability of the server node. For example, if the data throughput has reached or is close to its limit, then the server node may not be able to process more data.
[0057] The load level can be an indicator for comprehensively evaluating the current workload of the server node, usually calculated based on multiple factors such as the processor utilization rate and the data throughput. A high load level means that the server node is processing a large number of tasks and may not have enough resources to process new requests.
[0058] In some embodiments, the load level of the server node at the current moment is determined by the following formula (1):
[0059] LT i (t) = α · CPU i (t) + β · GPU i (t) + γ · IO i (t) (1);
[0060] Wherein, LT i (t) represents the load level of the i-th server node at time t, CPU i (t) represents the utilization rate of the CPU of the i-th server node at time t, GPU i (t) represents the utilization rate of the GPU of the i-th server node at time t, IO i (t) represents the data throughput of the i-th server node at time t, and α, β, and γ are weight coefficients, satisfying α + β + γ = 1. The weight coefficients can be adjusted according to the resource characteristics of the system and the business requirements. For example, the weight of α can be increased for CPU-intensive tasks, and the weight of β can be increased for GPU-intensive tasks.
[0061] Step S112: Determine the congestion level of the first server node at the current moment based on the network latency, packet loss rate, and bandwidth utilization rate of the first server node obtained at the current moment;
[0062] Here, network latency can refer to the time required for data to be transmitted in the network. Network latency may increase due to reasons such as network congestion, device failures, or excessive transmission distances. High network latency may lead to a slowdown in data transmission speed and affect the response speed of server nodes.
[0063] Packet loss rate can refer to the proportion of data packets lost during data transmission. The packet loss rate may increase due to reasons such as network congestion, device failures, or signal interference. A high packet loss rate may result in incomplete or lost data, affecting the performance and reliability of server nodes.
[0064] Bandwidth utilization can refer to the ratio of the network bandwidth currently used by a server node to the total available bandwidth, which can reflect the usage of the server node in terms of data transmission. For example, if the bandwidth utilization has reached or is close to its limit, then the server node may not be able to handle more data requests.
[0065] The degree of congestion can be an indicator for evaluating the network condition of a server node, usually calculated based on multiple factors such as network latency, packet loss rate, and bandwidth utilization. A high degree of congestion means that the network of the server node may be experiencing serious congestion problems, which may lead to a slowdown in data transmission speed, data loss, or a decrease in server response speed.
[0066] In some embodiments, the degree of congestion of the server node at the current moment is determined by the following formula (2):
[0067] CT i (t) = δ·LAT i (t) + ε·LOSS i (t) + ζ·BW i (t) (2);
[0068] Wherein, CT i (t) represents the degree of congestion of the i-th server node at time t, LAT i (t) represents the network latency of the i-th server node at time t, LOSS i (t) represents the packet loss rate of the i-th server node at time t, BW i (t) represents the bandwidth utilization of the i-th server node at time t, and δ, ε, ζ are weight coefficients, satisfying δ + ε + ζ = 1. The weight coefficients can be adjusted according to the characteristics of the network and service requirements. For example, for tasks sensitive to latency, the weight of δ can be increased, and for tasks sensitive to packet loss, the weight of ε can be increased.
[0069] Step S113, based on the load degree and the degree of congestion, determine whether the first server node can receive the initial data.
[0070] In some embodiments, the system determines whether the first server node can receive the initial data based on the load level and congestion level of the first server node at the current moment. If both the load level and congestion level of the first server node at the current moment are within an acceptable range, it is determined that the first server node can receive the initial data. On the contrary, if the load level or congestion level of the first server node at the current moment is too high, it is determined that the first server node cannot receive the initial data.
[0071] In the embodiments of the present application, on the one hand, by obtaining the utilization rate of the first processor, the utilization rate of the second processor, and the data throughput of the server node at the current moment, the load level of the server node can be accurately evaluated, so as to understand its current task processing ability; on the other hand, by obtaining the network latency, packet loss rate, and bandwidth utilization rate of the server node at the current moment, the network congestion level of the server node can be accurately evaluated, and then its data transmission ability can be judged; on the other hand, according to the load level and the congestion level of the server node at the current moment, the system can optimize the allocation of server resources. For example, in the case of light load and good network conditions, the initial data can be preferentially sent to this server node. On the contrary, in the case of heavy load or network congestion, the system can choose to send the initial data to other available server nodes.
[0072] In some embodiments, the implementation of step S113, "determine whether the first server node can receive the initial data based on the load level and the congestion level", may include the following steps S121 and S122, where:
[0073] Step S121, when both the load level and the congestion level are less than or equal to the first threshold, determine that the first server node can receive the initial data;
[0074] Here, the first threshold may be a preset value used to determine whether the load level and congestion level of the server node are within an acceptable range. This threshold is usually set according to factors such as the performance requirements of the system, historical data, and business needs. If both the load level and congestion level are less than or equal to the first threshold, it can be considered that the server node is in a good working state and has the ability to receive new tasks or data.
[0075] In some embodiments, the first threshold is determined by the following formula (3):
[0076]
[0077] where DPT1 i (t) represents the first threshold at time t, DPT1 i(t - 1) represents the first threshold at time t - 1, Δ low represents the basic step size for adjusting the first threshold, LT low represents the minimum value of the load level, CT low represents the minimum value of the congestion level, is used to calculate the relative proportion below the lower bound threshold, thereby dynamically adjusting the reduction amplitude of the threshold.
[0078] In some embodiments, if both the load level and the congestion level of the first server node at the current time are less than or equal to the first threshold, then the system considers that the first server node is currently in a good working state and can receive initial data.
[0079] Step S122, in the case where the load level is greater than the second threshold or the congestion level is greater than the second threshold, it is determined that the first server node cannot receive the initial data.
[0080] Here, the second threshold can be a preset value, which is used to determine whether the load level and the congestion level of the server node are within an acceptable range. This threshold is usually set according to factors such as the performance requirements of the system, historical data, and business requirements. If the load level or the congestion level is greater than the second threshold, then it can be considered that the server node is in an overloaded or network - congested state and has no ability to receive new tasks or data.
[0081] In some embodiments, the second threshold is determined by the following formula (4):
[0082]
[0083] where DPT2 i (t) represents the second threshold at time t, DPT2 i (t - 1) represents the second threshold at time t - 1, Δ high represents the basic step size for adjusting the second threshold, LT high represents the maximum value of the load level, CT high represents the maximum value of the congestion level, is used to calculate the relative proportion exceeding the upper bound threshold, thereby dynamically adjusting the increase amplitude of the threshold.
[0084] In some embodiments, if the load level of the first server node at the current time is greater than the first threshold or the congestion level is greater than the second threshold, then the system considers that the first server node may be in an overloaded or network - congested state and cannot receive initial data.
[0085] In the embodiments of the present application, by setting thresholds for the load level and congestion level, the system can take measures before the server node is about to be overloaded, avoiding performance degradation or crashes of the server due to processing too many tasks or data.
[0086] In some embodiments, the implementation of step S102, "in the case where the first server node cannot receive the initial data, based on the busy levels of other server nodes other than the first server node at the current moment, determine a target server node from the other server nodes", may include the following steps S131 and S132, where:
[0087] Step S131, in the case where the first server node cannot receive the initial data, respectively determine the load level and congestion level of the other server nodes at the current moment;
[0088] In some embodiments, first, the system attempts to send the initial data to the first server node, but since the load level or congestion level of the first server node exceeds a second threshold, it is determined that the first server node cannot receive the initial data. Next, the system needs to evaluate the load level and congestion level of all other server nodes in the server node cluster other than the first server node at the current moment.
[0089] Step S132, determine the server nodes whose load level and congestion level are both less than or equal to the first threshold as the target server nodes.
[0090] In some embodiments, the system traverses all other server nodes to check whether the load level and congestion level of the server nodes are both less than or equal to the first threshold. If both of these two metrics of a server node meet the conditions, then the server node is considered available and is selected as the target server node to receive the initial data.
[0091] In the embodiments of the present application, on the one hand, when the first server node cannot receive data, the system can automatically transfer the data to other available server nodes, thus avoiding service interruption caused by single point of failure; on the other hand, by dynamically selecting server nodes with lower load level and congestion level as the server nodes, the system can maintain the overall stability of operation.
[0092] In some embodiments, the implementation of step S103, "determine a target transmission path based on the quality of the transmission path between the first server node and the target server node", may include the following steps S141 and S142, where:
[0093] Step S141: Determine the quality of each transmission path at the current moment based on the network latency, bandwidth utilization, and stability of the transmission path between the obtained first server node and the target server node at the current moment.
[0094] Here, stability can refer to the reliability of the network path during data transmission. A stable network path can reduce errors and packet loss rates during data transmission, thus ensuring data integrity and consistency.
[0095] The quality of the transmission path can be comprehensively evaluated based on multiple performance indicators such as network latency, bandwidth utilization, and stability. A high-quality transmission path should have low network latency, high bandwidth utilization, and good stability.
[0096] In some embodiments, when transmitting data from the first server node to the target server node, there may be multiple different transmission paths to choose from. To determine which transmission path is optimal, the system evaluates all available transmission paths. The evaluation is based on performance indicators such as the network latency, bandwidth utilization, and stability of each path at the current moment. After obtaining these performance indicators, the system calculates the quality of each path according to a certain algorithm or rule.
[0097] In some embodiments, the quality of the transmission path at the current moment is determined by the following formula (5):
[0098] Q j (t) = φ · LAT j (t) + ψ · BW j (t) + ω · STAB j (t) (5);
[0099] Where, Q j (t) represents the quality of path j at time t, LAT j (t) represents the network latency of path j at time t, BW j (t) represents the bandwidth utilization of path j at time t, STAB j (t) represents the stability of path j at time t. φ, ψ, and ω are weight coefficients, satisfying φ + ψ + ω = 1. The weight coefficients can be adjusted according to business requirements and network characteristics. For example, for latency-sensitive services, the weight of φ can be increased; for bandwidth-sensitive services, the weight of ψ can be increased; and for services with high requirements for connection stability, the weight of ω can be increased.
[0100] Step S142: Determine the transmission path with the highest quality as the target transmission path.
[0101] In some embodiments, the system selects the transmission path with the highest quality as the target transmission path. This path is considered to be optimal at the current moment because it performs relatively well in terms of latency, bandwidth, and stability, ensuring the efficiency and reliability of data transmission.
[0102] In the embodiments of the present application, the system determines the target transmission path based on the network latency, bandwidth utilization, and stability of the transmission path between the first server node and the target server node at the current moment, which can significantly improve the efficiency and reliability of data transmission.
[0103] In some embodiments, the implementation of step S104, "transmitting the initial data of the first server node to the target server node through the target transmission path in the order of priority", may include the following steps S151 to S153, where:
[0104] Step S151, preprocessing, feature extraction, and quantity selection are sequentially performed on the initial data to obtain a target feature vector;
[0105] In some embodiments, before transmitting the initial data, it is first necessary to preprocess the initial data on the first server node. The preprocessing may include operations such as data cleaning (removing redundant or invalid data), data conversion (converting the data into a format suitable for subsequent processing), and data normalization (scaling the data to a unified scale), etc., to ensure the quality and consistency of the data.
[0106] In some embodiments, the preprocessed data then needs to undergo feature extraction. Feature extraction is to extract key or useful information from the data, which helps to reduce the complexity of the data while retaining sufficient information for subsequent analysis or processing.
[0107] In some embodiments, after feature extraction, quantity selection needs to be performed on the extracted features, which can reduce the complexity of the data while retaining the most useful information for subsequent processing.
[0108] Step S152, compressing and encapsulating the target feature vector in sequence to obtain a data packet of the initial data;
[0109] In some embodiments, after the initial data sequentially undergoes preprocessing, feature extraction, and quantity selection, a target feature vector is obtained. Before transmitting the target feature vector to the target server node, it is usually necessary to compress the target feature vector. The purpose of compression is to reduce the dimension of the target feature vector, thereby reducing the transmission time and bandwidth occupancy. The compressed feature vector needs to be encapsulated into a data packet. The data packet usually contains the data itself, metadata (such as the source, format, size, etc. of the data), and possible check information. The purpose of encapsulation is to ensure the integrity and consistency of the data during transmission.
[0110] In some embodiments, the compressed feature vector is determined by the following formula (6):
[0111] V c = PCA(V′, d) (6);
[0112] where V c represents the compressed feature vector, PCA(·) represents the PCA (Principal Component Analysis) function, V′ represents the target feature vector, and d represents the target dimension.
[0113] In some embodiments, the data packet is determined by the following formula (7):
[0114] P i = Packet(V c , i, s) (7);
[0115] where P i represents the data packet, Packet(·) represents the data encapsulation function, i represents the packet number, and s represents the packet size.
[0116] Step S153, transmit the data packet to the target server node through the target transmission path in the order of priority.
[0117] In some embodiments, the data packet needs to be transmitted to the target server node through the target transmission path in the order of priority. The order of priority may be determined based on the urgency, importance of the data, or other business logics. By preferentially transmitting high-priority data packets, it can be ensured that critical or urgent data can reach the target server node faster.
[0118] In the embodiments of the present application, on the one hand, by extracting key or useful information from the initial data, it helps to reduce the complexity of the data while retaining sufficient information for subsequent analysis or processing; on the other hand, by compressing the target feature vector, it can reduce the data storage space and lower the transmission cost.
[0119] In some embodiments, the implementation of step S151, "preprocessing, feature extraction, and quantity selection are sequentially performed on the initial data to obtain a target feature vector", may include the following steps S161 to S164, where:
[0120] Step S161, preprocess the initial data to obtain preprocessed data;
[0121] In some embodiments, before the initial data is transmitted, it is first necessary to preprocess the initial data on the first server node. The preprocessing may include operations such as data cleaning (removing redundant or invalid data), data conversion (converting the data into a format suitable for subsequent processing), and data normalization (scaling the data to a unified scale), etc., to ensure the quality and consistency of the data.
[0122] Step S162, standardize the preprocessed data to obtain standardized data;
[0123] In some embodiments, by standardizing the preprocessed data, the mean of the standardized data is 0 and the variance is 1.
[0124] In some embodiments, the standardized data is determined by the following formula (8):
[0125] X″ = (X′ - μ) / σ (8);
[0126] where X″ represents the standardized data, X′ represents the preprocessed data, and μ and σ are the mean and standard deviation of the data, respectively.
[0127] Step S163, extract features from the standardized data to obtain an initial feature vector;
[0128] In some embodiments, a Convolutional Neural Network (CNN) or a deep learning (Transformer) model based on the self-attention mechanism is used to extract features from the standardized data to obtain an initial feature vector.
[0129] In some embodiments, the initial feature vector is determined by the following formula (9):
[0130] V = F(X″, W) (9);
[0131] where V represents the initial feature vector, F(·) represents the CNN / transformer model, and W represents the weight matrix of the CNN / transformer model.
[0132] Step S164: Perform a quantity selection on the initial feature vector to obtain the target feature vector.
[0133] In some embodiments, for the quantity selection of the initial feature vector, the k most discriminative feature vectors selected are used as the target feature vector.
[0134] In some embodiments, the target feature vector is determined by the following formula (10):
[0135] V′ = SelectKBest(V, k) (10);
[0136] where V′ represents the target feature vector, SelectKBest(·) represents the feature vector selection function, and k represents the number of feature vectors selected.
[0137] In the embodiments of the present application, by successively performing preprocessing, standardization, feature extraction, and quantity selection on the initial data, the data quality can be improved, the features can be optimized, which is convenient for subsequent analysis.
[0138] In some embodiments, the implementation of step S151, "using the model corresponding to the target server node to perform inference on the initial data to obtain an inference result", may include the following steps S171 to S176, where:
[0139] Step S171: Use the target server node to decrypt the data packet of the initial data to obtain the decrypted feature vector;
[0140] In some embodiments, when the initial data is transmitted to the target server node in the form of a data packet, it is first necessary to perform a decryption operation on the data packet. The purpose of decryption is to restore the data in the data packet to the original format to obtain the decrypted feature vector.
[0141] In some embodiments, the decrypted feature vector is determined by the following formula (11):
[0142] V c = Unpack(P i ) (11);
[0143] where V c represents the decrypted feature vector, and Unpack(·) represents the data packet decryption function.
[0144] Step S172: Determine the matching degree between the decrypted feature vector and the model corresponding to the target server node;
[0145] Here, the matching degree may refer to the similarity degree between the target feature vector and the model corresponding to the target server node.
[0146] In some embodiments, a cosine similarity function is used to calculate the similarity between the target feature vector and the model corresponding to the target server node.
[0147] In some embodiments, the matching degree between the target feature vector and the model corresponding to the target server node is determined by the following formula (12):
[0148] C(M i ,V c ) = CosSim(M i ,V c ) = M i ·V c / ‖M i ‖‖V c ‖ (12);
[0149] Wherein, C(M i ,V c ) represents the matching degree between the target feature vector and the model corresponding to the target server, M i represents the model corresponding to the target server node, CosSim(·) represents the cosine similarity function, and ‖·‖ represents the Euclidean norm of the vector.
[0150] Step S173, when the matching degree is greater than the matching degree threshold, it is determined that the unsealed feature vector matches the model corresponding to the target server node;
[0151] Here, the matching degree threshold can be a preset value used to determine whether the target feature vector matches the model corresponding to the target server node.
[0152] In some embodiments, if the matching degree is greater than the matching degree threshold, it is considered that the target feature vector matches the model corresponding to the target server node. If the matching degree is less than the matching degree threshold, it is considered that the target feature vector does not match the model corresponding to the target server node.
[0153] Step S174, determine the model corresponding to the target server node with the highest matching degree with the unsealed feature vector as the target model;
[0154] In some embodiments, among multiple target server nodes, the model corresponding to the target server node with the highest matching degree with the unsealed feature vector is selected as the target model finally used to process the initial data.
[0155] Step S175, use the unsealed feature vector to adjust the target model to obtain an adjusted model;
[0156] In some embodiments, after determining the target model, the system adjusts the target model using the target feature vector. Such adjustment may include updating the parameters of the model, optimizing the structure of the model, etc., to ensure that the model can better adapt to the characteristics of the target feature vector.
[0157] In some embodiments, the adjusted model is determined by the following formula (13):
[0158] M′ max =LoRA(M max ,V c ,α,r) (13);
[0159] Wherein, M′ max represents the adjusted model, M max represents the model corresponding to the target server with the highest matching degree with the unsealed feature vector, LoRA(·) is the LoRA fine-tuning function, α is the fine-tuning learning rate, and r is the rank parameter.
[0160] Step S176, perform inference on the initial data using the adjusted model to obtain an inference result.
[0161] In some embodiments, assume that the user uploads a text "This mobile phone is really great, fast speed and good battery life", and hopes that the system can analyze the sentiment tendency (positive, negative or neutral) of this text. In this way, the model will perform inference on the sentiment tendency of the text through a series of calculation and inference steps (such as feature extraction, classifier judgment, etc.). After the model inference is completed, an inference result will be output. The inference result may be a probability distribution [0.9, 0.05, 0.05], indicating that the text has a 90% probability of positive sentiment, a 5% probability of negative sentiment, and a 5% probability of neutral sentiment.
[0162] In the embodiments of the present application, on the one hand, by determining the matching degree between the unsealed feature vector and the model corresponding to the target server node, a model most compatible with the data can be selected for inference, which helps to improve the accuracy and reliability of the inference; on the other hand, after determining the target model, the unsealed feature vector is used to adjust the target model to further improve the performance of the model, enabling the system to continuously learn and improve, and better adapt to the changing data and task environment.
[0163] In some embodiments, the inference method of the data may further include the following steps S181 and step S182, where:
[0164] Step S181, during the transmission of the initial data, detect the quality of the current transmission path;
[0165] In some embodiments, during the initial data transmission, the system will detect the quality of the currently used transmission path in real time. This quality may include multiple aspects, such as transmission speed, packet loss rate, latency, jitter, etc.
[0166] Step S182, when the quality of the current transmission path is less than the quality threshold, switch to other transmission paths to transmit the initial data.
[0167] Here, the quality threshold can be a preset value used to evaluate whether the quality of the transmission path meets the requirements. The quality threshold can be set according to the needs of the application scenario and is usually related to factors such as data transmission reliability, speed, latency, etc.
[0168] In some embodiments, when the system detects that the quality of the current transmission path is lower than the quality threshold, it is considered that the current transmission path is no longer suitable for subsequent data transmission, and the system will switch the subsequent data transmission to other available transmission paths.
[0169] In the embodiments of the present application, by detecting the quality of the current transmission path in real time, the system can timely discover and avoid potential problem paths. When the quality of the current transmission path decreases, the system can quickly switch to other transmission paths with better quality, thereby ensuring the continuous transmission of data and reducing the risk of data loss.
[0170] In the related art, distributed inference and fine-tuning systems usually adopt static resource allocation strategies and cannot be dynamically adjusted according to the real-time server node load and network conditions, resulting in low utilization of computing resources and situations of resource idling and waste. At the same time, the system has high costs in terms of hardware procurement and energy consumption. The embodiments of the present application realize the dynamic optimal configuration of resources through dynamic threshold adjustment technology, significantly improve the resource utilization rate, and reduce the operation cost.
[0171] The data transmission methods in the related art often rely on a single path and are prone to transmission interruption or failure in case of network congestion or server node failure, affecting the availability and stability of the system. The embodiments of the present application adopt a multi-path transmission strategy and an adaptive path selection mechanism, and through redundant transmission and dynamic path switching, greatly enhance the reliability of data transmission, improve the fault tolerance ability of the system, and ensure the continuity of critical services.
[0172] The inference and fine-tuning systems in the related art usually process tasks in a simple First In First Out (FIFO) manner and cannot give priority to critical tasks, resulting in long response times and difficulty in meeting real-time requirements. The embodiments of the present application introduce a priority scheduling mechanism that can dynamically adjust the processing order according to the urgency of tasks to ensure that critical tasks are processed first, significantly improving the user experience.
[0173] Inference and fine-tuning may have various modalities of input, such as text, video, voice, etc. The direct transmission of raw data will occupy a large amount of network bandwidth, bringing huge overhead to transmission and storage. The data compression methods in related technologies usually target specific fields and lack generality and adaptability. The embodiments of the present application adopt advanced feature extraction and data compression technologies, achieving a high compression ratio while being general, which can greatly save network bandwidth and storage space.
[0174] The embodiments of the present application propose a distributed inference and fine-tuning method based on dynamic threshold adjustment, multi-path transmission, and priority scheduling. By real-time monitoring the server node load and network congestion status, the data processing threshold is dynamically adjusted to optimize the resource utilization efficiency. At the same time, a multi-path transmission strategy is adopted to adaptively select the optimal transmission path according to the path quality and network status, improving the reliability and efficiency of data transmission. In addition, the system introduces a priority scheduling mechanism to perform hierarchical processing according to the importance and timeliness of data, ensuring the timely transmission and processing of critical data.
[0175] A distributed inference and fine-tuning method proposed by the embodiments of the present application, such as Figure 2 shown, the distributed inference and fine-tuning method may include the following steps S201 to step S209, where:
[0176] Step S201, the user inputs data;
[0177] Here, the data input by the user reaches a random server node and enters step S202.
[0178] Step S202, dynamic path selection;
[0179] In some embodiments, a multi-path transmission strategy is adopted to adaptively select the optimal transmission path according to the path quality and network status.
[0180] Step S203, priority scheduling;
[0181] In some embodiments, hierarchical processing is performed according to the importance and timeliness of data.
[0182] Step S204, multi-path transmission;
[0183] Here, if the data is transmitted to server node 1 in the server cluster, it enters step S205.
[0184] Step S205, the DPU preprocesses, extracts features, and compresses the data, and transmits the compressed data packet to other server nodes;
[0185] In some embodiments, the data is first preprocessed and feature extracted, and key features are extracted through existing machine learning models and principal component analysis techniques to obtain feature vectors; the feature vectors are compressed, and the compressed data packets are transmitted to other server nodes (e.g., server node 2, server node N).
[0186] Step S206, the GPU performs feature matching on the local model according to the data packet;
[0187] Here, the compressed feature vector is encapsulated into a data packet and transmitted to some server nodes through a dynamic path selection mechanism. The server node decapsulates the received data packet, extracts the feature vector, and calculates its matching degree with the local model in the model library. The data is dynamically allocated to the server node with the highest matching degree to achieve load balancing and resource optimization.
[0188] Step S207, fine-tuning the local model;
[0189] In some embodiments, the server node with the highest matching degree fine-tunes the local model using Low-Rank Adaptation (LoRA) technology of the large language model.
[0190] Step S208, the fine-tuned model performs the reasoning task;
[0191] In some implementations, the fine-tuned model performs an inference task to generate a final inference result.
[0192] Step S209: output the inference result and return it to the user.
[0193] In the embodiments of this application, detailed algorithm design and optimization are carried out in terms of multi-path transmission and priority scheduling, and mechanisms such as path selection, queue management, and scheduling strategies are proposed, fully considering the system's adaptability, reliability, real-time requirements, etc. At the same time, the system adopts a modular design concept, decoupling and encapsulating functions such as data preprocessing, feature extraction, data compression, data transmission, model fine-tuning, and reasoning tasks, thereby improving the scalability and maintainability of the system.
[0194] In short, in the embodiments of this application, through key technologies such as multipath transmission and priority scheduling, the system's resource utilization efficiency, data transmission reliability, and task processing real-time performance are effectively improved, while reducing data transmission and storage overhead, and has broad application prospects. The system can play an important role in scenarios such as intelligent manufacturing, autonomous driving, and smart cities, supporting large-scale, real-time, and efficient distributed reasoning and fine-tuning tasks.
[0195] The present application embodiment provides a task distribution method, such as Figure 3As shown, the task distribution method may include the following steps S301 to S317, where:
[0196] Step S301, when data arrives at the current server node, calculate the load level (LT) and congestion level (CT) of the current server node;
[0197] In some embodiments, define LT i (t) represents the load level of server node i at time t, and CT i (t) is the congestion level of server node i at time t.
[0198] Determine the load level of server node i at the current time t through the following formula (1):
[0199] LT i (t) = α·CPU i (t) + β·GPU i (t) + γ·IO i (t) (1);
[0200] Determine the congestion level of server node i at the current time t through the following formula (2):
[0201] CT i (t) = δ·LAT i (t) + ε·LOSS i (t) + ζ·BW i (t) (2);
[0202] Step S302, determine whether LT or CT is greater than the threshold;
[0203] Here, if so, that is, LT or CT is greater than the threshold, go to step S303; otherwise, that is, LT and CT are less than or equal to the threshold, go to step S304.
[0204] In some embodiments, determine the threshold through the following formula (3) or (4):
[0205]
[0206] Step S303, increase the data reception and processing threshold;
[0207] Step S304, lower the data reception and processing threshold;
[0208] Step S305, determine whether the data processing condition is met;
[0209] Here, if so, that is, the data processing condition is met, go to step S306.
[0210] In some embodiments, if LTi (t) ≤ DPT1 i (t) and CT i (t) ≤ DPT1 i (t), then the data processing condition is satisfied; if LT i (t) > DPT2 i (t) or CT i (t) > DPT2 i (t), then the data processing condition is not satisfied.
[0211] It should be noted that, in order to avoid infinite redirection of data, a maximum forwarding count K can be set for each data packet. When the forwarding count of the data packet reaches K, the current server node is forced to receive and process the data to ensure the final processing of the data.
[0212] Step S306, determine the data transmission path between the current service node and the target server node;
[0213] Step S307, select the optimal path according to the network congestion situation;
[0214] In some embodiments, the quality of transmission path j at the current time t is determined by the following formula (5):
[0215] Q j (t) = φ · LAT j (t) + ψ · BW j (t) + ω · STAB j (t) (5);
[0216] In some embodiments, the system will select the path with the highest quality as the target transmission path.
[0217] In some embodiments, by periodically sending probe packets, the quality of each path is calculated. The sending frequency of the probe packets can be adaptively adjusted according to the speed of network dynamic changes, reducing the sending frequency when the network is relatively stable and increasing the sending frequency when the network changes violently, so as to balance the probing overhead and the timely update of the path quality. In order to obtain more accurate and comprehensive path quality information, multiple probing methods can be used, such as a combination of end-to-end active probing and passive measurement of network intermediate nodes, so as to cover performance indicators at different levels.
[0218] It should be noted that in multi-path transmission, it is necessary to coordinate and synchronize data transmission on different paths to ensure data integrity and consistency. A data fragmentation and recombination mechanism based on sequence numbers can be adopted. The data is fragmented and numbered at the sending end, and the data is recombined at the receiving end according to the sequence numbers to achieve the reliability of multi-path transmission. In order to make full use of the bandwidth resources of multiple paths, load balancing needs to be performed between different paths. The data allocation ratio on different paths can be dynamically adjusted according to the path quality and current load conditions, and more data can be allocated to the path with better quality and lighter load to improve transmission efficiency and resource utilization. When performing multi-path load balancing, it is also necessary to consider the delay difference and bandwidth difference between different paths to avoid data out-of-order and congestion caused by path imbalance. A load balancing mechanism based on congestion control can be adopted to dynamically adjust the data allocation ratio according to the congestion state of the path and allocate the data to the path with a lower degree of congestion to achieve load balancing and congestion avoidance.
[0219] Step S308, determine whether the priority queue is non-empty;
[0220] Here, when it is time, that is, the priority queue is non-empty, go to step S309; otherwise, that is, the priority queue is empty, go to step S301.
[0221] Step S309, determine whether the high-priority queue is non-empty;
[0222] Here, when it is time, that is, the high-priority queue is non-empty, go to step S310; otherwise, that is, the high-priority queue is empty, go to step S311.
[0223] In some embodiments, the priority levels are defined as: Priority ∈ {high, medium, low}, where high represents high priority, medium represents medium priority, and low represents low priority. By introducing finer-grained priority levels, the priority requirements of different data can be met. According to the importance, timeliness, and business requirements of the data, a priority is assigned to each data packet. A rule-based priority assignment strategy can be adopted to automatically assign priorities to data packets according to predefined rules; or a machine learning-based priority assignment strategy can be adopted to dynamically predict the priorities of data packets by training a priority assignment model. When assigning priorities, it is also necessary to consider the size and processing time of the data packets to avoid large data packets or data packets with long processing times occupying the high-priority queue for a long time and affecting the transmission of other high-priority data. The priority can be dynamically adjusted according to the size and processing time of the data packets. For example, for data packets that exceed a certain size threshold or processing time threshold, their priorities are reduced.
[0224] In some embodiments, multiple priority queues are defined: Queuehigh (High - priority queue), Queue medium (Medium - priority queue), Queue low (Low - priority queue), each queue corresponds to a priority level and is used to store data packets with different priorities. According to the priority Priority of the data packet i , it is inserted into the corresponding priority queue. To avoid a certain priority queue not being processed for a long time, a maximum waiting time T can be set for each queue max , when the waiting time of the data packet in the queue exceeds T max , its priority is automatically increased and it is moved to a queue with a higher priority.
[0225] It should be noted that in the management of priority queues, dynamic adjustment of queues and overload protection also need to be considered. The service rate and scheduling strategy of the queue can be dynamically adjusted according to the length of the queue and the average waiting time of the data packet to adapt to the change of data traffic. When the length of a certain queue exceeds the predefined threshold, the overload protection mechanism is triggered to temporarily block the data reception of the queue until the queue length drops to the normal range.
[0226] In some embodiments, the queue scheduling order is defined as: Queue high →Queue medium →Queue low , and the data packets in each queue are processed in order from high to low priority. When sending data, the data packets in the high - priority queue are processed first. Only when the high - priority queue is empty, the data packets are taken from the second - highest - priority queue for sending.
[0227] It should be noted that a priority feedback mechanism is introduced to dynamically adjust the priority of the data packet according to the transmission situation and quality of service of the data packet. The priority increase and decrease rules can be defined. For example, for data packets with consecutive timeouts or packet losses, their priorities are increased; for data packets with consecutive successful transmissions and delays lower than the threshold, their priorities are decreased. Priority feedback can be based on end - to - end application - layer feedback or network - layer congestion control feedback. By comprehensively considering the feedback information of the application layer and the network layer, the priority of the data packet is dynamically adjusted to adapt to the change of network conditions and service requirements. In dynamic priority adjustment, the frequency and amplitude of priority adjustment also need to be considered to avoid overly frequent or drastic priority changes, which affect the stability of scheduling. The time window and adjustment step size of priority adjustment can be set to smooth and limit the priority adjustment.
[0228] Step S310, obtain data from the high - priority queue;
[0229] Step S311, determine whether the medium - priority queue is non - empty;
[0230] Here, when the medium-priority queue is not empty, step S312 is entered; otherwise, when the medium-priority queue is empty, step S313 is entered.
[0231] Step S312: Retrieve data from the medium-priority queue;
[0232] Step S313: Retrieve data from the low-priority queue;
[0233] Step S314: The current server node transmits data to the target server node through the optimal path;
[0234] Step S315: Determine whether the quality of the current path has deteriorated;
[0235] Here, when the quality of the current path has deteriorated, step S316 is entered; otherwise, when the quality of the current path has not deteriorated, step S317 is entered.
[0236] In some embodiments, a path switching threshold Q is defined switch = μ·min 1≤j≤N Q j (t)+(1 - μ)·Q current (t) where is the adjustment factor of the path switching threshold, and its value range is [0, 1]. The value of μ can be adaptively adjusted according to the severity of network changes, taking a larger value when the network changes gently and a smaller value when the network changes violently to balance the sensitivity and stability of path switching.
[0237] Step S316: Trigger path switching and select a path with better quality;
[0238] In some embodiments, when the quality Q current (t) of the current path < Q switch , path switching is triggered and the path with the highest quality is selected as the new transmission path. To avoid frequent path switching, a minimum duration T min of path switching can be introduced, that is, after switching to a new path, it is allowed to switch again only after maintaining for at least T min time.
[0239] Step S317: Continue to use the current path.
[0240] The embodiment of the present application provides a task processing method. As Figure 4 shown, the task distribution method may include the following steps S401 to S411, where:
[0241] Step S401: Send the preprocessed data X' to the GPU;
[0242] Step S402: The GPU extracts features from X' using a pre-trained machine learning model to obtain an initial feature vector V;
[0243] In some embodiments, before performing feature extraction on the preprocessed data X', the preprocessed data X' is normalized using the following formula (8) to obtain normalized data:
[0244] X″ = (X′ - μ) / σ (8);
[0245] Feature extraction is performed on the normalized data using the following formula (9) to obtain an initial feature vector:
[0246] V = F(X″, W) (9);
[0247] In some embodiments, a certain number of feature vectors are selected from the initial feature vector as target feature vectors using the following formula (10):
[0248] V′ = SelectKBest(V, k) (10);
[0249] Step S403: Determine whether the dimension of the target feature vector V’ is too high;
[0250] Here, if so, that is, the dimension of the target feature vector V’ is too high, go to step S404; otherwise, that is, the dimension of the target feature vector V’ is not high, go to step S405.
[0251] Step S404: Use PCA to compress the high-dimensional feature vector V’ into a low-dimensional feature vector Vc;
[0252] In some embodiments, the compressed feature vector is determined using the following formula (6):
[0253] V c = PCA(V′, d) (6);
[0254] Step S405: The GPU sends the feature vector V or Vc to the DPU through a high-speed channel;
[0255] Step S406: The DPU encapsulates Vc and transmits the encapsulated data packet to each server;
[0256] In some embodiments, the data packet is determined using the following formula (7):
[0257] P i = Packet(V c , i, s) (7);
[0258] Step S407: Each server receives and decrypts the data packet to obtain the feature vector Vc, and calculates the matching degree between Vc and its corresponding model M;
[0259] In some embodiments, the decrypted feature vector is determined by the following formula (11):
[0260] V c = Unpack(P i ) (11);
[0261] In some embodiments, the matching degree between the feature vector Vc and the target server's corresponding model is determined by the following formula (12):
[0262] C(M i ,V c ) = CosSim(M i ,V c ) = M i ·V c / ‖M i ‖‖V c ‖ (12);
[0263] Step S408: Traverse the matching degrees C(Mi, Vc) of all servers to find the server Nmax with the highest matching degree;
[0264] Step S409: Transmit the data to the server with the highest matching degree;
[0265] Step S410: The server Nmax uses Vc to fine-tune its model Mmax through deep learning;
[0266] In some embodiments, the adjusted model is determined by the following formula (13):
[0267] M′ max = LoRA(M max ,V c ,α,r) (13);
[0268] Step S411: Use the fine-tuned model to perform inference on the initial data to obtain the inference result Y.
[0269] In some embodiments, use the fine-tuned model M′ max to perform inference on the initial data X to obtain the inference result Y = M′ max (X).
[0270] In some embodiments, post-process the inference result Y, such as threshold filtering, softmax, etc., to obtain the final inference result where PostProcess(·) is the post-processing function and θ is the post-processing parameter.
[0271] The embodiments of the present application have the following advantages:
[0272] 1. When the distributed server node receives the data input by the user, optimize the data reception and processing strategies of the server node according to the load and network congestion status of the server node. And introduce path quality, dynamically select the optimal transmission path, and improve the reliability and efficiency of data transmission.
[0273] 2. Perform hierarchical processing and scheduling according to the importance and timeliness of the data to ensure the timely transmission and processing of critical data.
[0274] 3. According to the similarity between the feature vector and the local model, dynamically allocate the transmitted data through the above scheduling mechanism, fine-tune the model using the LoRA technology, and perform the inference task. Design reasonable evaluation indicators to dynamically optimize the model fine-tuning and inference process.
[0275] 4. Adopt a collaborative combination of multiple optimization mechanisms, including multi-path transmission, priority scheduling, etc., to provide a comprehensive and systematic optimization solution. This comprehensive optimization method is superior to the common local optimization or single-mechanism optimization in the related technologies, and can improve the performance and efficiency of the system from a global perspective.
[0276] 5. Adopt a multi-path transmission strategy and adaptive path selection, which can dynamically select the optimal transmission path according to the network status, effectively avoid network congestion and single-point failures, greatly improve the reliability and stability of data transmission, ensure the continuity of critical services to the greatest extent, and reduce the economic losses caused by service interruptions.
[0277] The embodiments of the present application provide an inference device for data, as Figure 5 shown. The inference device 500 for data includes:
[0278] The first determination module 510 is configured to determine whether the first server node can receive the initial data based on the busyness degree of the first server node at the current moment when the initial data input by the user arrives at the first server node;
[0279] The second determination module 520 is configured to, when the first server node cannot receive the initial data, determine a target server node from the other server nodes based on the busyness degree of the other server nodes except the first server node at the current moment;
[0280] The third determination module 530 is configured to determine a target transmission path based on the quality of the transmission path between the first server node and the target server node;
[0281] A transmission module 540, configured to transmit the initial data of the first server node to the target server node through the target transmission path in the order of priority;
[0282] An inference module 550, configured to perform inference on the initial data by using the model corresponding to the target server node to obtain an inference result.
[0283] In some embodiments, the busyness degree includes a load degree and / or a congestion degree, and the first determination module 510 includes: a first determination unit, configured to determine the load degree of the first server node at the current moment based on the utilization rate of the first processor, the utilization rate of the second processor, and the data throughput of the first server node obtained; a second determination unit, configured to determine the congestion degree of the first server node at the current moment based on the network latency, packet loss rate, and bandwidth utilization rate of the first server node obtained; a third determination unit, configured to determine whether the first server node can receive the initial data based on the load degree and the congestion degree.
[0284] In some embodiments, the third determination unit includes: a first determination subunit, configured to determine that the first server node can receive the initial data when both the load degree and the congestion degree are less than or equal to a first threshold; a second determination subunit, configured to determine that the first server node cannot receive the initial data when the load degree is greater than a second threshold or the congestion degree is greater than the second threshold.
[0285] In some embodiments, the second determination module 520 includes: a fourth determination unit, configured to respectively determine the load degree and the congestion degree of the other server nodes at the current moment when the first server node cannot receive the initial data; a fifth determination unit, configured to determine the server node corresponding to both the load degree and the congestion degree being less than or equal to the first threshold as the target server node.
[0286] In some embodiments, the third determination module 530 includes: a sixth determination unit, configured to determine the quality of each transmission path at the current moment based on the network latency, bandwidth utilization rate, and stability of the transmission path between the first server node and the target server node obtained; a seventh determination unit, configured to determine the transmission path with the highest quality as the target transmission path.
[0287] In some embodiments, the transmission module 540 includes: a first processing unit configured to perform preprocessing, feature extraction, and quantity selection on the initial data in sequence to obtain a target feature vector; a second processing unit configured to perform compression and encapsulation on the target feature vector in sequence to obtain a data packet of the initial data; and a transmission unit configured to transmit the data packet to the target server node through the target transmission path in accordance with a priority order.
[0288] In some embodiments, the first processing unit includes: a preprocessing subunit configured to perform preprocessing on the initial data to obtain preprocessed data; a normalization subunit configured to normalize the preprocessed data to obtain normalized data; a feature extraction subunit configured to perform feature extraction on the normalized data to obtain an initial feature vector; and a selection subunit configured to perform quantity selection on the initial feature vector to obtain the target feature vector.
[0289] In some embodiments, the inference module 550 includes: a de-encapsulation unit configured to de-encapsulate the data packet of the initial data by using the target server node to obtain a de-encapsulated feature vector; an eighth determination unit configured to determine a matching degree between the de-encapsulated feature vector and a corresponding model of the target server node; a ninth determination unit configured to determine that the de-encapsulated feature vector matches the corresponding model of the target server node when the matching degree is greater than a matching degree threshold; a tenth determination unit configured to determine a model corresponding to the target server node with the highest matching degree with the de-encapsulated feature vector as a target model; an adjustment unit configured to adjust the target model by using the de-encapsulated feature vector to obtain an adjusted model; and an inference unit configured to perform inference on the initial data by using the adjusted model to obtain an inference result.
[0290] In some embodiments, the inference device 500 for data further includes: a detection module configured to detect a quality of a current transmission path during transmission of the initial data; and a switching module configured to switch to another transmission path to transmit the initial data when the quality of the current transmission path is less than a quality threshold.
[0291] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the related technology, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific hardware, software, or firmware, or any combination among hardware, software, and firmware.
[0292] The embodiments of the present application further provide an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.
[0293] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, it implements some or all of the steps in the above method. The computer-readable storage medium can be transient or non-transient.
[0294] The embodiments of the present application further provide a computer program, including computer-readable code. When the computer-readable code runs in a computing device, the processor in the computing device executes to implement some or all of the steps in the above method.
[0295] The embodiments of the present application further provide a computer program product. The computer program product includes a non-transient computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above method. The computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.
[0296] It should be noted here that: the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities can be referred to each other. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.
[0297] An embodiment of the present application provides a hardware entity of an electronic device, such as Figure 6 As shown, the hardware entity of the electronic device 600 includes: The processor 601 generally controls the overall operation of the electronic device 600. The communication interface 602 enables the electronic device to communicate with other terminals or servers through a network. The memory 603 is configured to store instructions and applications executable by the processor 601, and can also cache data to be processed or already processed by the processor 601 and each module in the electronic device 600 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM). Data transmission can be performed between the processor 601, the communication interface 602, and the memory 603 through the bus 604.
[0298] It should be understood that the term "one embodiment" or "an embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above steps / processes do not mean the order of execution, and the order of execution of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0299] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0300] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the couplings, direct couplings, or communication connections between the various components shown or discussed may be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms.
[0301] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0302] In addition, in each embodiment of the present application, the various functional units can all be integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit. The above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0303] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage medium includes various media that can store program codes, such as removable storage devices, read-only memories, magnetic disks, or optical discs.
[0304] Alternatively, if the above-mentioned integrated units of the present application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the related technology can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The foregoing storage medium includes various media that can store program codes, such as removable storage devices, ROMs, magnetic disks, or optical discs.
[0305] As described above, it is only the implementation mode of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application.
Claims
1. A method for reasoning about data, characterized in that: The method comprises: When the initial data input by the user arrives at the first server node, determining whether the first server node can receive the initial data based on the busyness of the first server node at the current moment; In the case that the first server node cannot receive the initial data, determining a target server node from the other server nodes based on the busyness of the other server nodes except the first server node at the current moment; determining a target transmission path based on the quality of the transmission path between the first server node and the target server node; Transmitting the initial data of the first server node to the target server node through the target transmission path in order of priority; The initial data is inferred using the model corresponding to the target server node to obtain an inference result.
2. The inference method according to claim 1, characterized in that: The busyness level includes load level and / or congestion level. Determining whether the first server node can receive the initial data based on the busyness of the first server node at the current moment includes: Determine the load level of the first server node at the current moment based on the acquired utilization rate of the first processor, the utilization rate of the second processor and the data throughput of the first server node at the current moment; Determine the congestion level of the first server node at the current moment based on the acquired network delay, packet loss rate, and bandwidth utilization rate of the first server node at the current moment; Based on the load level and the congestion level, it is determined whether the first server node can receive the initial data.
3. The inference method according to claim 2, characterized in that: Determining whether the first server node can receive the initial data based on the load level and the congestion level includes: In a case where both the load degree and the congestion degree are less than or equal to a first threshold, determining that the first server node is capable of receiving the initial data; When the load level is greater than a second threshold or the congestion level is greater than the second threshold, it is determined that the first server node cannot receive the initial data.
4. The inference method according to claim 1, characterized in that: In the case that the first server node cannot receive the initial data, based on the busyness of other server nodes except the first server node at the current moment, determining a target server node from the other server nodes includes: In the case that the first server node cannot receive the initial data, respectively determining the load degree and congestion degree of the other server nodes at the current moment; A server node corresponding to which both the load degree and the congestion degree are less than or equal to a first threshold is determined as the target server node.
5. The inference method according to claim 1, characterized in that: Determining a target transmission path based on the quality of the transmission path between the first server node and the target server node includes: Determine the quality of each transmission path at the current moment based on the acquired network delay, bandwidth utilization, and stability of the transmission path between the first server node and the target server node at the current moment; The transmission path with the highest quality is determined as the target transmission path.
6. The inference method according to claim 1, characterized in that: Transmitting the initial data of the first server node to the target server node through the target transmission path according to the priority order, comprising: Preprocessing, feature extraction and quantity selection are performed on the initial data in sequence to obtain a target feature vector; Compressing and encapsulating the target feature vector in sequence to obtain a data packet of the initial data; The data packets are transmitted to the target server node through the target transmission path in order of priority.
7. The inference method according to claim 6, characterized in that: The initial data is sequentially preprocessed, feature extracted and quantity selected to obtain a target feature vector, including: Preprocessing the initial data to obtain preprocessed data; Standardizing the preprocessed data to obtain standardized data; Performing feature extraction on the standardized data to obtain an initial feature vector; The initial feature vector is quantitatively selected to obtain the target feature vector.
8. The inference method according to claim 1, characterized in that: The initial data is inferred using the model corresponding to the target server node to obtain an inference result, including: Decapsulating the data packet of the initial data using the target server node to obtain a decapsulated feature vector; Determining the matching degree between the unsealed feature vector and the corresponding model of the target server node; When the matching degree is greater than a matching degree threshold, determining that the unsealed feature vector matches the target server node corresponding model; Determine the model corresponding to the target server node with the highest matching degree of the unsealed feature vector as the target model; Using the unsealed feature vector to adjust the target model to obtain an adjusted model; The adjusted model is used to infer the initial data to obtain an inference result.
9. The inference method according to any one of claims 1 to 8, characterized in that: The method further comprises: During the initial data transmission, detecting the quality of the current transmission path; When the quality of the current transmission path is less than a quality threshold, switching to another transmission path to transmit the initial data.
10. A data reasoning device, characterized in that: The device comprises: A first determination module, configured to determine, when the initial data input by the user arrives at the first server node, whether the first server node can receive the initial data based on the busyness of the first server node at a current moment; A second determination module is used to determine a target server node from other server nodes except the first server node based on the busyness of the other server nodes at the current moment when the first server node cannot receive the initial data; A third determination module, configured to determine a target transmission path based on the quality of the transmission path between the first server node and the target server node; a transmission module, configured to transmit the initial data of the first server node to the target server node through the target transmission path in order of priority; The inference module is used to use the model corresponding to the target server node to infer the initial data to obtain an inference result.
11. An electronic device comprising a processor and a memory, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the computer program, the steps in the method according to any one of claims 1 to 9 are implemented.
12. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 9 are implemented.
13. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps in the method according to any one of claims 1 to 9 are implemented.