Request processing method and device, storage medium and computer equipment

By real-time monitoring of node load performance indicators and dynamically allocating requests to nodes with lighter loads or better performance, the node overload problem caused by request balancing in distributed systems is solved, achieving efficient and stable request processing.

CN120812059APending Publication Date: 2025-10-17BEIJING CENTURY TAL EDUCATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511081752.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In a distributed system, evenly distributed request processing results in low-capacity nodes handling high-resource consumption requests, which may cause failures, affect stability, increase maintenance costs, and prolong processing time.

Method used

By obtaining the load performance indicator parameters of multiple request processing nodes in real time, determining the load score and scheduling weight, selecting nodes with lighter load or better performance to process requests, building a hash ring for node mapping, and optimizing scheduling in combination with cache relevance.

Benefits of technology

It improves the efficiency and stability of request processing, avoids system failures and maintenance costs caused by node overload, and achieves flexible node allocation and improved resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120812059A_ABST
    Figure CN120812059A_ABST
Patent Text Reader

Abstract

The invention discloses a request processing method and device, a storage medium and computer equipment, relates to the technical field of intelligent question answering, and mainly aims to improve the request processing efficiency, guarantee the system stability and save the system maintenance cost. The method comprises the following steps: acquiring load performance index parameters of a plurality of request processing nodes in real time in response to a processing signal of a to-be-processed request; determining a load score of each request processing node based on the load performance index parameter, and determining a load scheduling weight of the corresponding request processing node based on the load score of each request processing node; and based on the load scheduling weight, selecting a target request processing node in each request processing node, and scheduling the target request processing node to process the to-be-processed request. The method and the device are suitable for scenes for processing requests.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent question answering, and in particular to a request processing method and device, a storage medium and computer equipment. BACKGROUND

[0002] With the rapid development of Internet technology and the wide popularity of various application services, the number of network requests is showing explosive growth, such as in the field of intelligent question answering in educational technology, the number of requests for questions from students is increasing. In a complex distributed system architecture, how to efficiently process a large number of requests has become a key problem for improving the overall performance and user experience of the system.

[0003] At present, an equal number of requests is usually allocated to each request processing node for processing. However, the processing capacity of different nodes is different, and allocating requests equally will cause some low-capacity nodes to process high-resource-consumption requests, and the overload of nodes may cause faults or crashes, thereby affecting the stability and reliability of the nodes and increasing the maintenance cost, while prolonging the processing time of the requests. SUMMARY

[0004] The present application provides a request processing method and device, a storage medium and computer equipment, which can improve the processing efficiency of requests, ensure system stability and save system maintenance costs.

[0005] According to a first aspect of the present application, a request processing method is provided, comprising:

[0006] In response to a processing signal of a to-be-processed request, real-time acquisition of load performance index parameters of a plurality of request processing nodes is performed;

[0007] Based on the load performance index parameters, a load score of each request processing node is determined, and a load scheduling weight of the corresponding request processing node is determined based on the load score of each request processing node;

[0008] Based on the load scheduling weight, a target request processing node is selected from each request processing node, and the target request processing node is scheduled to process the to-be-processed request.

[0009] Optionally, based on the load scheduling weight, a target request processing node is selected from each request processing node, comprising:

[0010] Real-time acquisition of cache information of a plurality of request processing nodes is performed;

[0011] The cache association degree of the to-be-processed request and the cache information of each request processing node is determined respectively, and based on the cache association degree, a cache association weight of each request processing node is determined;

[0012] Select a target request processing node in each of the request processing nodes based on the load scheduling weight and the cache association weight.

[0013] Optionally, the target request processing node is scheduled to process the to-be-processed request, including:

[0014] In a case where the target request processing node is multiple, node attribute information of each target request processing node is acquired, wherein the node attribute information includes at least one of a node IP address and a host name;

[0015] Each target request processing node is converted into a node hash value and the to-be-processed request is converted into a request hash value by using a preset hash function respectively, a hash ring is constructed based on a node hash value range, and each target request processing node is mapped to different positions on the hash ring based on the node hash value;

[0016] The to-be-processed request is mapped to a target position on the hash ring based on the request hash value, a first encountered target request processing node is found on the hash ring clockwise from the target position, and the first encountered target request processing node is scheduled to process the to-be-processed request.

[0017] Optionally, a cache association degree of the to-be-processed request and cache information of each request processing node is determined respectively, including:

[0018] A radix tree is constructed for cache information corresponding to each request processing node respectively;

[0019] A plurality of hash list items are constructed based on the to-be-processed request, a request hash list is formed by the plurality of hash list items, and a list length of the request hash list is determined;

[0020] The request hash list is matched with each radix tree respectively, and a matching depth of the request hash list in each radix tree is determined based on a matching result;

[0021] A cache association degree of the to-be-processed request and cache information of each request processing node is determined based on the matching depth and the list length.

[0022] Optionally, a plurality of hash list items are constructed based on the to-be-processed request, including:

[0023] A system prompt word and a user prompt word in the to-be-processed request are determined, and the user prompt word is split into a plurality of groups of sub-user prompt words based on a preset number of characters;

[0024] determine a system hash value corresponding to the system prompt word and a user hash value corresponding to each group of sub-user prompt words respectively, and take the system hash value and each group of user hash values as corresponding hash list items.

[0025] Optionally, based on the cache association degree, a cache association weight of each request processing node is determined, including:

[0026] Obtain a node concurrency of each request processing node and a node concurrency average corresponding to each node concurrency.

[0027] Based on the node concurrency, the node concurrency average and the cache association degree, a cache association weight of each request processing node is determined.

[0028] Optionally, based on the load scheduling weight and the cache association weight, a target request processing node is selected from each request processing node, including:

[0029] Determine an importance coefficient corresponding to the load scheduling weight and the cache association weight respectively, add the load scheduling weight and the cache association weight based on the importance coefficient to obtain a node selection weight, and select a target request processing node from each request processing node based on the node selection weight.

[0030] Optionally, the load performance index parameter includes at least one of a queuing request number, a current running request number, a cache utilization rate and a scheduling request number in a previous window period corresponding to the request processing node.

[0031] Based on the load performance index parameter, a load score of each request processing node is determined, including:

[0032] Determine a weight coefficient corresponding to each load performance index parameter corresponding to each request processing node respectively, and based on the weight coefficient, perform weighted summation on each load performance index parameter to obtain a load score of each request processing node.

[0033] Optionally, the target request processing node is scheduled to process the to-be-processed request, including:

[0034] Determine whether there is first response information matching the to-be-processed request in cache information corresponding to the target request processing node.

[0035] If there is first response information matching the to-be-processed request in the cache information, the first response information is fed back to an initiation end of the to-be-processed request.

[0036] If the first response information matching the to-be-processed request does not exist in the cache information, the target request processing node is scheduled to generate second response information corresponding to the to-be-processed request, and the second response information is fed back to the initiator of the to-be-processed request, and the second response information is stored in association with the to-be-processed request in the cache corresponding to the target request processing node.

[0037] Optionally, after the first response information matching the to-be-processed request exists in the cache information, the method further comprises:

[0038] determining the storage medium type of the first response information, if the storage medium type is a display memory, directly calling the first response information in the display memory to feed back to the initiator of the to-be-processed request;

[0039] if the storage medium type is a block storage, moving the first response information in the block storage to the display memory, and calling the first response information in the display memory after moving to feed back to the initiator of the to-be-processed request.

[0040] Optionally, the method further comprises:

[0041] acquiring the access times, access time intervals and information amounts of each cache information in the display memory;

[0042] determining the activity of each cache information based on the access times, the access time intervals and the information amounts, and removing the cache information with an activity less than a preset threshold from the display memory to obtain the display memory after releasing space.

[0043] Optionally, a radix tree is constructed for each cache information corresponding to the request processing node, comprising:

[0044] determining first target cache information with a common prefix and second target cache information without a common prefix in the cache information;

[0045] merging the common prefix in the first target cache information as a first intermediate node to be constructed in the radix tree, and taking each non-common prefix of the second target cache information as a second intermediate node to be constructed in the radix tree;

[0046] taking the first target cache information excluding the common prefix as a first leaf node corresponding to the first intermediate node, and taking the second target cache information excluding the non-common prefix as a second leaf node corresponding to the second intermediate node;

[0047] constructing the radix tree corresponding to the cache information based on the first intermediate node, the second intermediate node, the first leaf node and the second leaf node.

[0048] Optionally, determining a load score of each request processing node based on the load performance indicator parameter includes:

[0049] The load performance indicator parameters corresponding to each of the request processing nodes are respectively input into a preset score prediction model for score prediction to obtain a load score for each of the request processing nodes, wherein the preset score prediction model is pre-constructed based on a sample indicator parameter data set with score labels.

[0050] According to a second aspect of the present invention, there is provided a request processing device, comprising:

[0051] an acquisition unit, configured to acquire load performance indicator parameters of a plurality of request processing nodes in real time in response to a processing signal of a request to be processed;

[0052] a determining unit, configured to determine a load score of each of the request processing nodes based on the load performance indicator parameter, and determine a load scheduling weight of the corresponding request processing node based on the load score of each of the request processing nodes;

[0053] The request processing unit is configured to select a target request processing node from each of the request processing nodes based on the load scheduling weight, and schedule the target request processing node to process the pending request.

[0054] According to a third aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, which implements the above-requested processing method when executed by a processor.

[0055] According to a fourth aspect of the present invention, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the processing method requested above when executing the program.

[0056] According to the request processing method, device, storage medium and computer equipment provided by the application, compared with the current method of processing equal number of requests for each request processing node, the application scores the load of each node by the load performance index parameter of each request processing node, determines the scheduling weight of each node according to the load score, and finally selects the node to process the request according to the scheduling weight, so that the request is allocated to the node with lighter load or better performance by monitoring the load performance index of the node in real time, the resource waste of the node is avoided, the fast and stable processing of the request is ensured, the system failure and maintenance cost caused by the single point overload of the node are avoided, and the dynamic allocation of the scheduling weight is realized by monitoring the load performance index parameter in real time, so that the flexible allocation of the node is realized. BRIEF DESCRIPTION OF DRAWINGS

[0057] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and serve to explain the principles of the application, and do not limit the application. In the drawings:

[0058] Figure 1 A flow chart of a request processing method provided by an embodiment of the application is shown;

[0059] Figure 2 A flow chart of another request processing method provided by an embodiment of the application is shown;

[0060] Figure 3 A schematic diagram of a weight calculation process provided by an embodiment of the application is shown;

[0061] Figure 4 An example diagram of a radix tree provided by an embodiment of the application is shown;

[0062] Figure 5 A structural schematic diagram of a request processing device provided by an embodiment of the application is shown;

[0063] Figure 6 A structural schematic diagram of another request processing device provided by an embodiment of the application is shown;

[0064] Figure 7 An entity structural schematic diagram of a computer device provided by an embodiment of the application is shown. DETAILED DESCRIPTION

[0065] Hereinafter, the application will be described in detail with reference to the accompanying drawings and in combination with the embodiments. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.

[0066] Currently, the way of processing the same number of requests for each request processing node, the equal distribution of requests will make some low-capability nodes process high-resource consumption requests, and the node may fail or crash due to overload, thereby affecting its stability and reliability, increasing maintenance costs, and prolonging the processing time of the request.

[0067] To solve the above problems, the embodiment of the application provides a request processing method, as shown in the figure, the method comprises: Figure 1

[0068] 101, in response to the processing signal of the request to be processed, the load performance index parameters of the plurality of request processing nodes are acquired in real time.

[0069] For the embodiment of the application, the state of the node is monitored by the cluster controller, the IP address of the node is reported to the gateway, the configuration of the gateway is updated, the performance index parameters of each node are collected through the index interface exposed by the inference framework (such as vllm), and the collected index parameters are written into the index warehouse (such as Prometheus) for persistent storage. The inference gateway will dynamically update the scheduling weight of each backend request processing node according to the collected load performance index parameters, so that the traffic is more reasonably distributed, and the resource utilization is improved.

[0070] 102, determine the load score of each request processing node based on the load performance index parameters, and determine the load scheduling weight of the corresponding request processing node based on the load score of each request processing node.

[0071] Among them, the load performance index parameters include but are not limited to the number of queued requests of the processing node, the number of current running requests, the key value cache utilization rate, the number of scheduled requests in the last window period, the GPU utilization rate, etc. The number of queued requests refers to the number of requests waiting for the request processing node to process. The number of running requests refers to the number of requests being processed by the node. The key value cache utilization rate refers to the actual usage rate of the key value (Key-Value, KV) cache. The number of scheduled requests in the last window period refers to the number of scheduled requests of the node in the last time period. Since the weight calculation and index collection of the gateway are periodically executed, the period of weight calculation is relatively short (such as 1s), and the period of index collection is relatively long (such as 10s). Before each weight update, requests will be preferentially scheduled to the node with the lowest load, but the index warehouse needs a relatively long period of time to update, and the number of scheduled requests during the period is not considered, which may cause uneven load, thereby the application introduces the number of scheduled requests in the last window period for weight calculation, which can improve the load balancing of the node.

[0072] ​For the embodiment of the present application, in order to determine the load scheduling weight of the request processing node, it is necessary to first determine the load score of each request processing node, and based on this, step 102 specifically comprises: determining the weight coefficient corresponding to each load performance index parameter of each request processing node respectively, and based on the weight coefficient, performing weighted summation on each load performance index parameter to obtain the load score of each request processing node.

[0073] Specifically, the weight coefficient is set for each load performance index parameter according to actual needs, and then the load score of each request processing node is calculated according to the following formula:

[0074] a i = -(λ1R1+λ2R2+λ3R3+λ4R4+λ5R5)

[0075] wherein a i represents the load score of the i-th request processing node, λ1, λ2, λ3, λ4, λ5 respectively represent the weight coefficients corresponding to the queued request number, the current running request number, the key value cache utilization rate, the scheduled request number in the last window period, and the GPU utilization rate, and R1, R2, R3, R4, R5 respectively represent the queued request number, the current running request number, the key value cache utilization rate, the scheduled request number in the last window period, and the GPU utilization rate. Thus, the load score of each request processing node can be calculated in the above manner. In another embodiment of the present application, if the main factors affecting the node scheduling weight are the queued request number and the scheduled request number in the window period, and the node has queued requests, the scheduling weight will be greatly reduced (therefore, a higher weight can be set for the queued request number and the scheduled request number in the last window period, such as 100); if a certain instance schedules a request in the window period, its scheduling weight will also be greatly reduced in the next window period (therefore, the weight coefficient can be set to 100); if the queued request number + the scheduled request number in the window period of each node is close, the node with less running request number, lower key value cache utilization rate and lower GPU utilization rate can be preferentially scheduled (therefore, the weight coefficients of the running request number, the key value cache utilization rate and the GPU utilization rate can be set to 10, 1 and 1 respectively). It should be noted that the above examples are only illustrative and do not specifically limit the embodiment of the present application. Further, after determining the load score of each processing node, the load scheduling weight of the node can be determined according to the load score, such as the higher the load score, the greater the load scheduling weight; the lower the load score, the smaller the load scheduling weight. The load scheduling weight can be reasonably set according to the size of the load score. The embodiment of the present application can improve the rationality of scheduling weight setting by performing load scoring for each node and setting the scheduling weight according to the load score, thereby ensuring the node resource utilization rate.

[0076] 103. Based on the load scheduling weight, a target request processing node is selected in each request processing node, and the target request processing node is scheduled to process the request to be processed.

[0077] For the embodiment of the application, a node with a load scheduling weight greater than a preset threshold (the preset threshold is set according to actual needs) is selected as a target request processing node in each request processing node, and finally the target request processing node is used for request processing. Thus, by monitoring the load performance indicators of the nodes in real time, the request is allocated to the node with lighter load or better performance, avoiding the waste of resources of the node, ensuring the rapid and stable processing of the request, and avoiding the system failure and maintenance cost caused by the single-point overload caused by allocating the request to the node with heavier load. Meanwhile, the application realizes the dynamic allocation of the scheduling weight by monitoring the load performance indicator parameters in real time, and can realize the flexible allocation of the node.

[0078] In another embodiment of the application, if the target request processing node is multiple, in order to further improve the request processing effect, the optimal node can also be selected in the multiple target request processing nodes. Based on this, the method comprises: acquiring node attribute information of each target request processing node in the case that the target request processing node is multiple, wherein the node attribute information comprises at least one of node IP address and host name; converting the node attribute information of each target request processing node into a node hash value and converting the request to be processed into a request hash value by using a preset hash function respectively, constructing a hash ring based on the node hash value range, and mapping each target request processing node to different positions of the hash ring based on the node hash value; mapping the request to be processed to a target position of the hash ring based on the request hash value, finding the first encountered target request processing node on the hash ring clockwise from the target position, and scheduling the first encountered target request processing node to process the request to be processed.

[0079] Specifically, the same hash function is used to hash each node's node attribute information and the unique identifier of the pending request (such as the request ID, user ID, and request content summary) to obtain a hash value. The output range of the hash function is then connected end to end to form a ring space, namely a hash ring. Based on the calculated node hash value, each target request processing node is mapped to the corresponding position on the hash ring. For example, the hash value 0x1234abcd of node A corresponds to position 305419897 on the ring, and the hash value 0x5678ef12 of node B corresponds to position 1450743314 on the ring. The nodes are arranged in ascending order of hash value (if the hash values ​​conflict, they can be distinguished by additional information such as the port number). The hash value of the pending request is mapped to the target position of the hash ring. For example, the hash value of the pending request, 0x3a7b9c2f, corresponds to position 981234543 on the ring. Starting from the position corresponding to the request hash value, the hash ring is searched clockwise. The first target request processing node encountered is the node responsible for processing the request. For example, if the node positions on the hash ring are 305419897 (node ​​A) and 1450743314 (node ​​B), the first node encountered clockwise from the request position 981234543 is node B. The request is dispatched to the found node (such as node B) for processing. The embodiment of the present invention selects the optimal node for processing the request by constructing a hash ring. The hash function evenly maps nodes and requests to the ring. Combined with virtual node technology, it can further avoid skew caused by physical node performance differences, thereby improving the accuracy of node selection. At the same time, when adding a new node, it only needs to insert the hash value of the new node into the ring, without the need for global reconstruction, supporting seamless expansion.

[0080] According to a request processing method provided by the present invention, compared with the current method of evenly distributing an equal number of requests to each request processing node for processing, the present invention uses the load performance index parameters of each request processing node to load score each node, and determines the scheduling weight of each node based on the load score, and finally selects a node to process the request based on the scheduling weight. Thus, by real-time monitoring of the load performance index of the node, the request is allocated to the node with lighter current load or better performance, thereby avoiding waste of node resources, ensuring fast and stable processing of requests, and avoiding system failures and maintenance costs caused by single-point overload caused by allocating requests to nodes with heavier loads. At the same time, the present invention realizes dynamic allocation of scheduling weights by real-time monitoring of load performance index parameters, which can achieve flexible allocation of nodes.

[0081] Furthermore, in order to better illustrate the above-mentioned request processing process, as a refinement and extension of the above-mentioned embodiment, the embodiment of the present invention provides another request processing method, such as Figure 2 As shown, the method includes:

[0082] 201、real-time acquisition of load performance index parameters and cache information of multiple request processing nodes in response to a processing signal of a request to be processed.

[0083] The cache information refers to various information stored in the cache corresponding to each node. For example, if each node is a node in an intelligent question and answer system, the cache information corresponding to the node can be question and answer pairs stored in the cache. Figure 3 As shown in the figure, in a cloud-native environment, the inference gateway (load balancer) monitors the available backend request processing nodes (Endpoints) through the cluster controller (Controller) and reports the IP addresses (Pod IP) in its cluster to automatically update the routing configuration. For each request processing node, collect the load performance index parameters (metrics) and store them in the index warehouse (metrics collect), and collect the cache information of each node. The inference gateway will dynamically update the weight of each backend request processing node based on the collected index parameters and cache information, so that the traffic is more reasonably distributed.

[0084] 202、determine the load score of each request processing node based on the load performance index parameters, and determine the load scheduling weight of the corresponding request processing node based on the load score of each request processing node.

[0085] For the embodiment of the application, in order to determine the load scheduling weight of the node, the load score of the node needs to be determined first. Based on this, step 202 specifically includes: inputting each load performance index parameter corresponding to each request processing node into a preset score prediction model for score prediction to obtain the load score of each request processing node, wherein the preset score prediction model is constructed based on a sample index parameter data set with score labels in advance.

[0086] Specifically, in order to improve the prediction accuracy of the preset score prediction model, the preset score prediction model needs to be trained and constructed first. Based on this, the method includes: constructing a preset initial score prediction model; obtaining a sample data set, wherein the sample data set includes sample load performance index parameters of multiple backend processing nodes in a sample processing system with load score labels; dividing the sample data set into a training set and a test set, training the preset initial score prediction model using the training set, and testing the trained preset initial score prediction model using the test set, and finally using the trained preset initial score prediction model that meets the test condition as the preset score prediction model.

[0087] Specifically, in the model training process, first, a preset initial score prediction model is constructed, and then a sample data set is obtained. Ensure that the data set contains all necessary files, including various load performance indicator parameters of multiple request processing nodes (queued request number, running request number, key value cache utilization, scheduled request number in the last window period, GPU utilization, etc.). Convert the data into a format that the preset initial score prediction model can understand, and finally train and test the model. Specifically, the data set can be divided first: use random or specific strategies (such as stratified sampling) to divide the sample data set into a training set and a test set. Then use the training set to train the model, and use the test set to test the trained model to evaluate its performance on unseen data. Calculate and record the mCP, accuracy, recall rate and other indicators on the test set. If the model performance does not meet the requirements, it can return to the training stage for more iterations or adjustments. A preset score prediction model that meets the requirements is obtained. Finally, input the various load performance indicator parameters corresponding to the processing nodes into the preset score prediction model, and the preset score prediction model can directly output the load score of the corresponding processing node. The embodiment of the present application predicts the load score through the model, which can improve the prediction accuracy and efficiency of the load score. Further, according to the load score, set a scheduling weight for each node, such as each score interval corresponding to a scheduling weight, determine the target score interval to which the load score belongs, and take the weight corresponding to the target score interval as the load scheduling weight of the corresponding node.

[0088] 203、respectively determine the cache association degree of the cache information of the to-be-processed request and each request processing node, and determine the cache association weight of each request processing node based on the cache association degree.

[0089] For the embodiment of the present application, in order to determine the cache association weight, it is necessary to first determine the cache association degree, based on which, step 203 specifically comprises: constructing a radix tree for the cache information corresponding to each of the request processing nodes; constructing a plurality of hash list items based on the to-be-processed request, forming a request hash list from the plurality of hash list items, and determining the list length of the request hash list; respectively matching the request hash list with each of the radix trees, and determining the matching depth of the request hash list in each of the radix trees based on the matching result; and determining the cache association degree of the to-be-processed request and the cache information of each of the request processing nodes based on the matching depth and the list length. The method of constructing a radix tree comprises: determining first target cache information with a common prefix and second target cache information without a common prefix in the cache information; merging the common prefix in the first target cache information as a first intermediate node of the to-be-constructed radix tree, and respectively taking each non-common prefix of the second target cache information as a second intermediate node of the to-be-constructed radix tree; taking the first target cache information excluding the common prefix as a first leaf node corresponding to the first intermediate node, and respectively taking the second target cache information excluding the non-common prefix as a second leaf node corresponding to the second intermediate node; and constructing the radix tree corresponding to the cache information based on the first intermediate node, the second intermediate node, the first leaf node, and the second leaf node. The method of constructing a plurality of hash list items comprises: determining the system prompt word and the user prompt word in the to-be-processed request, and splitting the user prompt word into a plurality of groups of sub-user prompt words based on a preset number of characters; respectively determining the system hash value corresponding to the system prompt word and the user hash value corresponding to each group of sub-user prompt words, and taking the system hash value and each group of user hash values as corresponding hash list items.

[0090] Specifically, for efficient storage and retrieval of key-value pairs, cache information (keys) needs to be organized into a radix tree. Taking a request processing node as an example, from the cache information corresponding to the request processing node, identify keys with common prefixes (first target cache information) and keys without common prefixes (second target cache information), merge common prefixes into intermediate nodes of the radix tree, and non-common prefixes as independent intermediate nodes, and the remaining parts as leaf nodes mounted to the corresponding intermediate nodes, to build a complete radix tree. It should be noted that in the field of intelligent Q&A, the key represents the question and the value represents the answer. For example, Q1: What is the content of Newton's first law? -> A1: An object at rest stays at rest or in uniform motion in a straight line when subjected to no external force; Q2: What is the formula of Newton's second law? -> A2: F = ma (force equals mass times acceleration); Q3: What is the content of Ohm's law? -> A3: The current in a conductor is proportional to the voltage and inversely proportional to the resistance. The common prefix of Q1 and Q2 is "Newton", and Q3 has no common prefix. The common prefix "Newton" of Q1 and Q2 is merged into the first intermediate node N1, and Q3 has no common prefix, so its complete keyword "What is the content of Ohm's law?" is directly split into an independent node, the second intermediate node N2: stores "Ohm"; the subsequent nodes are in turn "law", "content", "is", "what", starting from N1, split the non-common parts of Q1 and Q2: path 1: N1 -> "first" -> "law" -> "content" -> "is" -> "what" -> leaf node L1 (stores A1); path 2: N1 -> "second" -> "law" -> "formula" -> "is" -> "what" -> leaf node L2 (stores A2). Starting from N2, split the complete keyword of Q3: path: N2 -> "law" -> "content" -> "is" -> "what" -> leaf node L3 (stores A3), and the final constructed radix tree is as shown in Figure 4 .

[0091] Further, the hash list item needs to be determined according to the to-be-processed request. For example, if the system prompt word is "user asks a question", the user prompt word is "how to learn artificial intelligence", the user prompt word is divided into two groups of sub-user prompt words "how to learn" and "artificial intelligence" according to a preset number of characters (the preset number of characters is set according to actual needs). Then, the hash values of the system prompt word and each group of sub-user prompt words are calculated by using a hash function. For example, the hash value of the system prompt word is "0x7f5c3b2a1d9e8f6a", the hash values of the two groups of sub-user prompt words are "0x3e4d5c6b7a8f9e0d" and "0x1a2b3c4d5e6f7a8b" respectively. "0x7f5c3b2a1d9e8f6a", "0x3e4d5c6b7a8f9e0d" and "0x1a2b3c4d5e6f7a8b" are taken as the hash list items, and the final hash list is [0x7f5c3b2a1d9e8f6a, 0x3e4d5c6b7a8f9e0d, 0x1a2b3c4d5e6f7a8b]. The length of the hash list is 3.

[0092] Further, the information of each node in the radix tree can be converted into a hash value for storage. Then, the request hash list is matched with each node of the radix tree layer by layer. For example, if the first item hash value in the request hash list is matched with the radix tree first, if the first item hash value is H_sys1, and the starting point of path A in the radix tree corresponding to a certain node is H_sys1, the matching is successful, the matching depth is recorded as 1, then the second item hash value H_userC1 in the request hash list is matched with the next node H_userA1 of path A, the result is not matched, and the matching is terminated, and finally the matching depth is determined as 1. The ratio of the matching depth to the list length of the request hash list is taken as the cache association degree of the to-be-processed request and the cache information of the corresponding request processing node. In this way, the cache association degree of the to-be-processed request and the cache information of each request processing node can be determined in the above manner.

[0093] Further, the cache association weight needs to be determined based on the cache association degree. Based on this, the method comprises: acquiring the node concurrency of each request processing node and the node concurrency average corresponding to each node concurrency; determining the cache association weight of each request processing node based on the node concurrency, the node concurrency average and the cache association degree.

[0094] Specifically, the cache association weight is determined according to the following formula:

[0095]

[0096] wherein ω idenotes the cache association weight corresponding to the ith request processing node, f i denotes the cache association degree of the ith request processing node, b i denotes the node concurrency of the ith request processing node, j denotes the request processing node, n denotes the total number of nodes of the request processing node, b j denotes the node concurrency of the jth request processing node.

[0097] 204. Select a target request processing node in each request processing node based on the load scheduling weight and the cache association weight, and schedule the target request processing node to process the to-be-processed request.

[0098] For the embodiment of the application, after determining the load scheduling weight and the cache association weight of each request processing node, a target request processing node needs to be selected according to the above weights. Based on this, step 204 specifically includes: respectively determining the importance coefficients corresponding to the load scheduling weight and the cache association weight, adding the load scheduling weight and the cache association weight based on the importance coefficients to obtain a node selection weight, and selecting a target request processing node in each request processing node based on the node selection weight.

[0099] Specifically, the importance coefficients of the load scheduling weight and the cache association weight are set according to actual needs, then the load scheduling weight and the cache association weight are weighted and summed according to the importance coefficients to obtain a selection weight, and finally the request processing node corresponding to the maximum selection weight is selected as the target request processing node, or the node with a selection weight greater than a preset weight threshold (the preset weight threshold can be set according to actual needs) is selected as the target request processing node. The embodiment of the application selects the node for processing the request through the load scheduling weight and the cache association weight, which can not only ensure the load balancing of the node, but also improve the cache hit rate, and can maximize the resource utilization rate and the response efficiency.

[0100] Further, after determining the target request processing node, the target request processing node needs to be scheduled to process the to-be-processed request. The process of scheduling the target request processing node to process the to-be-processed request includes: judging whether there is first response information matching the to-be-processed request in the cache information corresponding to the target request processing node; if there is first response information matching the to-be-processed request in the cache information, feeding back the first response information to the initiation end of the to-be-processed request; if there is no first response information matching the to-be-processed request in the cache information, scheduling the target request processing node to generate second response information corresponding to the to-be-processed request, feeding back the second response information to the initiation end of the to-be-processed request, and storing the second response information and the to-be-processed request in association in the cache corresponding to the target request processing node.

[0101] Specifically, the to-be-processed request is matched with the cache information corresponding to the target request processing node, for example, in the field of intelligent question answering, the to-be-processed request is matched with the question in the question and answer pair in the cache information, if there is a question with high similarity to the to-be-processed request, the answer corresponding to the question is taken as the first response information, and then it is needed to determine the storage medium based on the first response information to select the information feedback mode. Based on this, the method comprises: determining the storage medium type of the first response information, if the storage medium type is the video memory, the first response information in the video memory is directly called to feed back to the initiator of the to-be-processed request; if the storage medium type is the block storage, the first response information in the block storage is moved to the video memory, and the first response information after moving in the video memory is called to feed back to the initiator of the to-be-processed request.

[0102] Wherein, the storage medium type includes but is not limited to video memory and block storage. Specifically, if the first response information is stored in the video memory, the first response information is directly fed back to the request initiating user from the video memory, if the first response information is stored in the block storage, the response information needs to be moved to the video memory, such as using GDS (Global Distribution System) to move the first response information. Then the first response information is fed back to the request initiating user from the video memory. Further, if there is no first response information matching the to-be-processed request in the cache information, a new video memory is opened to generate the second response information corresponding to the to-be-processed request and feed back to the request initiating user. The embodiment of the application first queries the response information from the cache information, which can improve the information acquisition efficiency and reduce the computing resources. The embodiment of the application can solve the cache size limitation problem through the video memory exchange mechanism.

[0103] Further, in order to release the storage space of the video memory corresponding to each request processing node, so as to increase the information query efficiency, the method further comprises: acquiring the access times, access time interval and information amount of each cache information in the video memory; based on the access times, the access time interval and the information amount, determining the activity degree of each cache information, and removing the cache information with an activity degree less than a preset threshold from the video memory to obtain the video memory after the space is released. Specifically, the access times, the access time interval and the information amount are weighted and summed, and the activity degree of the video memory is determined according to the weighted sum result, for example, the more the access times, the shorter the access time interval and the larger the information amount, the higher the corresponding activity degree. Finally, the cache information with an activity degree less than a preset threshold (the preset threshold is set according to actual needs) is removed, so as to release the storage space of the video memory, facilitating the fast query of the subsequent response information.

[0104] According to another request processing method provided by the present invention, compared with the current method of evenly distributing an equal number of requests to each request processing node for processing, the present invention uses the load performance index parameters of each request processing node to load score each node, and determines the scheduling weight of each node based on the load score, and finally selects the node to process the request based on the scheduling weight. Thus, by real-time monitoring of the load performance index of the node, the request is allocated to the node with lighter current load or better performance, thereby avoiding waste of node resources, ensuring fast and stable processing of requests, and avoiding system failures and maintenance costs caused by single-point overload caused by allocating requests to nodes with heavier loads. At the same time, the present invention realizes dynamic allocation of scheduling weights by real-time monitoring of load performance index parameters, which can achieve flexible allocation of nodes.

[0105] Further, as Figure 1 The embodiment of the present invention provides a request processing device, such as Figure 5 As shown, the device includes: an acquisition unit 31, a determination unit 32, and a request processing unit 33.

[0106] The acquisition unit 31 may be configured to acquire load performance indicator parameters of multiple request processing nodes in real time in response to a processing signal of a request to be processed.

[0107] The determining unit 32 may be configured to determine a load score of each of the request processing nodes based on the load performance indicator parameter, and determine a load scheduling weight of the corresponding request processing node based on the load score of each of the request processing nodes.

[0108] The request processing unit 33 may be configured to select a target request processing node from each of the request processing nodes based on the load scheduling weight, and schedule the target request processing node to process the pending request.

[0109] In a specific application scenario, in order to select a target request processing node in each request processing node, such as Figure 6 As shown, the request processing unit 33 includes an acquisition module 331 , a determination module 332 , and a node selection module 333 .

[0110] The acquisition module 331 can be used to acquire cache information of multiple request processing nodes in real time.

[0111] The determination module 332 may be configured to determine cache relevance between the pending request and cache information of each request processing node, and determine a cache relevance weight of each request processing node based on the cache relevance.

[0112] The node selection module 333 may be configured to select a target request processing node from each of the request processing nodes based on the load scheduling weight and the cache association weight.

[0113] In a specific application scenario, in order to schedule the target request processing node to process the pending request, the request processing unit 33 further includes a construction module 334 and a search module 335 .

[0114] The acquisition module 331 may also be configured to acquire node attribute information of each target request processing node when there are multiple target request processing nodes, wherein the node attribute information includes at least one of a node IP address and a host name.

[0115] The construction module 334 can be used to convert the node attribute information of each target request processing node into a node hash value and convert the pending request into a request hash value using a preset hash function, construct a hash ring based on the node hash value range, and map each target request processing node to a different position on the hash ring based on the node hash value.

[0116] The search module 335 can be used to map the pending request to the target position of the hash ring based on the request hash value, search for the first encountered target request processing node clockwise on the hash ring from the target position, and schedule the first encountered target request processing node to process the pending request.

[0117] In a specific application scenario, in order to determine the cache association, the determination module 332 can be specifically used to construct a radix tree for the cache information corresponding to each of the request processing nodes; construct multiple hash list items based on the pending request, and the request hash list is composed of multiple hash list items, and the list length of the request hash list is determined; the request hash list is matched with each of the radix trees respectively, and the matching depth of the request hash list in each of the radix trees is determined based on the matching result; based on the matching depth and the list length, the cache association of the pending request and the cache information of each of the request processing nodes is determined.

[0118] In a specific application scenario, in order to construct multiple hash list items, the determining module 332 may be specifically configured to determine the system prompt word and the user prompt word in the request to be processed, and split the user prompt word into multiple groups of sub-user prompt words based on a preset number of characters;

[0119] A system hash value corresponding to the system prompt word and a user hash value corresponding to each group of sub-user prompt words are determined respectively, and the system hash value and each group of user hash values ​​are used as corresponding hash list items.

[0120] In a specific application scenario, in order to determine the cache association weight, the determination module 332 can be specifically configured to obtain the node concurrency of each request processing node and the node concurrency average corresponding to each node concurrency; determine the cache association weight of each request processing node based on the node concurrency, the node concurrency average and the cache association degree.

[0121] In a specific application scenario, in order to select the target request processing node, the node selection module 333 can be specifically configured to determine the importance coefficient corresponding to the load scheduling weight and the cache association weight respectively, add the load scheduling weight and the cache association weight based on the importance coefficient to obtain a node selection weight, and select the target request processing node in each request processing node based on the node selection weight.

[0122] In a specific application scenario, the load performance index parameter includes at least one of the queuing request number, the current running request number, the cache utilization rate and the scheduling request number in the last window period corresponding to the request processing node; in order to determine the load score of each request processing node, the determination unit 32 can be specifically configured to determine the weight coefficient corresponding to each load performance index parameter corresponding to each request processing node respectively, and based on the weight coefficient, the load performance index parameters are weighted and summed to obtain the load score of each request processing node.

[0123] In a specific application scenario, in order to schedule the target request processing node to process the to-be-processed request, the request processing unit 33 further includes a judgment module 336, a feedback module 337 and a generation module 338.

[0124] The judgment module 336 can be configured to judge whether there is first response information matched with the to-be-processed request in the cache information corresponding to the target request processing node.

[0125] The feedback module 337 can be configured to feed back the first response information to the initiation end of the to-be-processed request if there is first response information matched with the to-be-processed request in the cache information.

[0126] The generation module 338 can be configured to generate second response information corresponding to the to-be-processed request if there is no first response information matched with the to-be-processed request in the cache information, feed back the second response information to the initiation end of the to-be-processed request, and store the second response information and the to-be-processed request in the cache corresponding to the target request processing node.

[0127] In a specific application scenario, in order to feed back the response information, the feedback module 337 can also be configured to determine the storage medium type of the first response information, if the storage medium type is a display memory, directly call the first response information in the display memory to feed back to the initiator of the request to be processed, and if the storage medium type is a block storage, move the first response information in the block storage to the display memory, and call the first response information after moving in the display memory to feed back to the initiator of the request to be processed.

[0128] In a specific application scenario, in order to clean up the information in the display memory, the feedback module 337 can also be configured to obtain the access times, access time intervals and information amounts of each cache information in the display memory, determine the activity of each cache information based on the access times, access time intervals and information amounts, remove the cache information with an activity less than a preset threshold from the display memory, and obtain the display memory after releasing the space.

[0129] In a specific application scenario, in order to construct a radix tree, the determination module 332 can specifically be configured to determine, in the cache information, first target cache information with a common prefix and second target cache information without a common prefix, combine the common prefix in the first target cache information as a first intermediate node of a radix tree to be constructed, and respectively combine each non-common prefix of the second target cache information as a second intermediate node of a radix tree to be constructed, take the first target cache information after excluding the common prefix as a first leaf node corresponding to the first intermediate node, and respectively take the second target cache information after excluding the non-common prefix as a second leaf node corresponding to the second intermediate node, and construct a radix tree corresponding to the cache information based on the first intermediate node, the second intermediate node, the first leaf node and the second leaf node.

[0130] In a specific application scenario, in order to determine the load score of each request processing node, the determination unit 32 can specifically be configured to input each load performance index parameter corresponding to each request processing node into a preset score prediction model for score prediction to obtain the load score of each request processing node, wherein the preset score prediction model is constructed in advance based on a sample index parameter data set with a score label.

[0131] It should be noted that other corresponding descriptions of the functions of the request processing device provided by the embodiments of the present application can be referred to the corresponding descriptions of the method shown in FIG. 13, and will not be described here in detail. Figure 1

[0132] Based on the above, the method shown in FIG. 13 can be implemented by the request processing device shown in FIG. 12. Figure 1 ​The method shown, accordingly, an embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, which implements the following steps when executed by a processor: in response to a processing signal of a pending request, obtaining load performance indicator parameters of multiple request processing nodes in real time; determining a load score of each of the request processing nodes based on the load performance indicator parameters, and determining a load scheduling weight of the corresponding request processing node based on the load score of each of the request processing nodes; based on the load scheduling weight, selecting a target request processing node in each of the request processing nodes, and scheduling the target request processing node to process the pending request.

[0133] Based on the above Figure 1 The method shown and Figure 5 The embodiment of the device shown in the figure, the embodiment of the present invention also provides a physical structure diagram of a computer device, such as Figure 7 As shown, the computer device includes: a processor 41, a memory 42, and a computer program stored in the memory 42 and executable on the processor, wherein the memory 42 and the processor 41 are both arranged on a bus 43 and implement the following steps when the processor 41 executes the program: in response to a processing signal of a request to be processed, obtaining load performance index parameters of a plurality of request processing nodes in real time; determining a load score of each of the request processing nodes based on the load performance index parameters, and determining a load scheduling weight of the corresponding request processing node based on the load score of each of the request processing nodes; selecting a target request processing node in each of the request processing nodes based on the load scheduling weight, and scheduling the target request processing node to process the request to be processed.

[0134] Through the technical solution of the present invention, the present invention uses the load performance index parameters of each request processing node to load score each node, and determines the scheduling weight of each node based on the load score, and finally selects the node to process the request based on the scheduling weight. Therefore, by real-time monitoring of the load performance index of the node, the request is allocated to the node with lighter current load or better performance, avoiding waste of node resources, ensuring fast and stable processing of requests, and avoiding system failures and maintenance costs caused by single-point overload caused by allocating requests to nodes with heavier loads. At the same time, the present invention realizes dynamic allocation of scheduling weights by real-time monitoring of load performance index parameters, which can achieve flexible allocation of nodes.

[0135] It should be apparent to those skilled in the art that the modules or steps of the application described above can be implemented with a general purpose computing device, which can be centralized on a single computing device or distributed over a network of multiple computing devices, and optionally implemented with program code executable by a computing device, which can be stored in a storage device and executed by a computing device, and in some cases, the steps shown or described can be performed in a different order than shown, or made into individual integrated circuit modules, or multiple modules or steps made into a single integrated circuit module. Thus, the application is not limited to any particular combination of hardware and software.

[0136] The preferred embodiments of the application described above are intended to be merely exemplary and those skilled in the art will recognize that changes can be made to the above-described embodiments without departing from the spirit and scope of the application. What is desired to be protected by letters patent is set forth in the appended claims.

Claims

1. A method for processing a request, characterized in that: include: Responding to a processing signal of a pending request, obtaining load performance indicator parameters of a plurality of request processing nodes in real time; Determining a load score for each of the request processing nodes based on the load performance indicator parameter, and determining a load scheduling weight for the corresponding request processing node based on the load score for each of the request processing nodes; Based on the load scheduling weight, a target request processing node is selected from each of the request processing nodes, and the target request processing node is scheduled to process the pending request.

2. The method according to claim 1, characterized in that Selecting a target request processing node in each of the request processing nodes based on the load scheduling weight includes: Get cache information of multiple request processing nodes in real time; Determining cache relevance of the pending request and cache information of each request processing node respectively, and determining a cache relevance weight of each request processing node based on the cache relevance; A target request processing node is selected in each of the request processing nodes based on the load scheduling weight and the cache association weight.

3. The method according to any one of claims 1 to 2, characterized in that Scheduling the target request processing node to process the pending request includes: In the case where there are multiple target request processing nodes, obtaining node attribute information of each target request processing node, wherein the node attribute information includes at least one of a node IP address and a host name; Using a preset hash function, the node attribute information of each target request processing node is converted into a node hash value and the request to be processed is converted into a request hash value, a hash ring is constructed based on the node hash value range, and each target request processing node is mapped to a different position of the hash ring based on the node hash value; Based on the request hash value, the pending request is mapped to the target position of the hash ring, and the first encountered target request processing node is searched clockwise on the hash ring from the target position, and the first encountered target request processing node is scheduled to process the pending request.

4. The method according to claim 2, characterized in that Determining cache relevance of the pending request and cache information of each request processing node respectively includes: Construct a radix tree for the cache information corresponding to each request processing node respectively; Constructing a plurality of hash list items based on the pending request, forming a request hash list from the plurality of hash list items, and determining a list length of the request hash list; Matching the requested hash list with each of the radix trees respectively, and determining a matching depth of the requested hash list in each of the radix trees based on a matching result; Based on the matching depth and the list length, a cache association between the request to be processed and the cache information of each request processing node is determined.

5. The method according to claim 4, characterized in that A plurality of hash list items are constructed based on the pending request, including: determining system prompt words and user prompt words in the request to be processed, and dividing the user prompt words into multiple groups of sub-user prompt words based on a preset number of characters; A system hash value corresponding to the system prompt word and a user hash value corresponding to each group of sub-user prompt words are determined respectively, and the system hash value and each group of user hash values ​​are used as corresponding hash list items.

6. The method according to claim 2, characterized in that Determining a cache relevance weight of each of the request processing nodes based on the cache relevance includes: Obtaining the node concurrency number of each request processing node and the node concurrency average corresponding to each node concurrency number; A cache relevance weight of each of the request processing nodes is determined based on the node concurrency number, the node concurrency average, and the cache relevance.

7. The method according to claim 2, characterized in that Selecting a target request processing node in each of the request processing nodes based on the load scheduling weight and the cache association weight includes: Determine the importance coefficients corresponding to the load scheduling weight and the cache association weight respectively, add the load scheduling weight and the cache association weight based on the importance coefficients to obtain a node selection weight, and select a target request processing node in each of the request processing nodes based on the node selection weight.

8. The method according to claim 1, characterized in that The load performance indicator parameter includes at least one of the number of queued requests corresponding to the request processing node, the number of currently running requests, cache utilization, and the number of scheduled requests in the previous window period; Determining a load score of each of the request processing nodes based on the load performance indicator parameter includes: The weight coefficients corresponding to the load performance index parameters corresponding to each of the request processing nodes are determined respectively, and based on the weight coefficients, the load performance index parameters are weighted and summed to obtain the load score of each of the request processing nodes.

9. The method according to claim 1, characterized in that Scheduling the target request processing node to process the pending request includes: Determine whether there is first response information matching the request to be processed in the cache information corresponding to the target request processing node; If there is first response information matching the pending request in the cache information, feeding back the first response information to the initiator of the pending request; If there is no first response information matching the pending request in the cache information, the target request processing node is scheduled to generate second response information corresponding to the pending request, and the second response information is fed back to the initiator of the pending request, and the second response information is associated with the pending request and stored in the cache corresponding to the target request processing node.

10. The method according to claim 9, characterized in that If the cache information contains first response information that matches the pending request, the method further includes: Determine the storage medium type of the first response information, and if the storage medium type is video memory, directly call the first response information in the video memory and feed it back to the initiator of the pending request; If the storage medium type is block storage, the first response information in the block storage is moved to the video memory, and the first response information moved to the video memory is called and fed back to the initiator of the request to be processed.

11. The method according to claim 10, characterized in that The method further comprises: Obtaining the number of accesses, access time interval, and information volume of each cache information in the video memory; Based on the number of accesses, the access time interval, and the amount of information, the activity of each cached information is determined, and cached information with an activity less than a preset threshold is removed from the video memory to obtain the video memory after freeing up space.

12. The method according to claim 4, characterized in that Constructing a radix tree for the cache information corresponding to each request processing node respectively, including: Determining, in the cache information, first target cache information having a common prefix and second target cache information not having a common prefix; Merge the common prefixes in the first target cache information as the first intermediate node of the radix tree to be constructed, and use each non-common prefix of the second target cache information as the second intermediate node of the radix tree to be constructed; The first target cache information after excluding the common prefix is ​​used as the first leaf node corresponding to the first intermediate node, and the second target cache information after excluding the non-common prefix is ​​used as the second leaf node corresponding to the second intermediate node; Based on the first intermediate node, the second intermediate node, the first leaf node, and the second leaf node, a radix tree corresponding to the cache information is constructed.

13. The method according to claim 1, wherein Determining a load score of each of the request processing nodes based on the load performance indicator parameter includes: The load performance indicator parameters corresponding to each of the request processing nodes are respectively input into a preset score prediction model for score prediction to obtain a load score for each of the request processing nodes, wherein the preset score prediction model is pre-constructed based on a sample indicator parameter data set with score labels.

14. A request processing device, characterized in that: include: an acquisition unit, configured to acquire load performance indicator parameters of a plurality of request processing nodes in real time in response to a processing signal of a request to be processed; a determining unit, configured to determine a load score of each of the request processing nodes based on the load performance indicator parameter, and determine a load scheduling weight of the corresponding request processing node based on the load score of each of the request processing nodes; The request processing unit is configured to select a target request processing node from each of the request processing nodes based on the load scheduling weight, and schedule the target request processing node to process the pending request.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.

16. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.