Request processing method and device, electronic equipment and storage medium

By calculating parallel probabilities in a large model system and dynamically adjusting resource allocation strategies, the problem of low hardware resource utilization during traffic troughs is solved, achieving more efficient resource utilization and faster response capabilities.

CN120614360APending Publication Date: 2025-09-09BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510854810.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

In large-scale model systems, hardware resources have low utilization during low-traffic periods, and the existing system cannot respond to traffic fluctuations in real time, resulting in idle resources and high overall time consumption.

Method used

By determining the idle load and current inlet traffic of the target subsystem, calculating the parallel probability, and dynamically adjusting the system resource allocation strategy, the upstream subsystem and the target subsystem are allowed to process user requests in parallel or serially.

Benefits of technology

It improves resource utilization during low traffic periods, reduces waiting time, improves the system's service quality and resource utilization efficiency, and avoids resource idleness or overload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120614360A_ABST
    Figure CN120614360A_ABST
Patent Text Reader

Abstract

The invention provides a request processing method and device, electronic equipment and a storage medium, relates to the technical field of computers, in particular to the fields of artificial intelligence, data processing, deep learning and the like, and can be applied to application scenes such as large model dialogue and the like. According to the specific implementation scheme, the method comprises the steps of determining an idle load of a target subsystem in response to a received user request; determining a parallel probability according to the current inlet flow and idle load of the processing system; the processing system comprises a target subsystem and an upstream subsystem of the target subsystem. And generating a request processing result of the user request by using the upstream subsystem and the target subsystem based on the parallel probability. According to the scheme, the resource utilization rate in the flow trough period can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, in particular to the fields of artificial intelligence, data processing, deep learning, etc., and can be used in application scenarios such as large-scale model dialogues, and specifically relates to a request processing method, device, electronic device, and storage medium. Background Art

[0002] In large-scale model systems, the hardware resources of each subsystem are generally configured based on the capacity of traffic during peak periods. However, resource utilization is extremely low during low traffic periods, leaving a large amount of hardware resources idle. The existing system consumes a lot of time and cannot respond to traffic fluctuations in real time, making it difficult to improve resource utilization in a fixed resource deployment scenario. Summary of the Invention

[0003] The present disclosure provides a request processing method, device, electronic device, and storage medium.

[0004] According to a first aspect of the present disclosure, a request processing method is provided, comprising: determining an idle load of a target subsystem in response to a received user request; determining a parallel probability based on the current inlet traffic and idle load of the processing system; the processing system comprising a target subsystem and an upstream subsystem of the target subsystem; and generating a request processing result of the user request using the upstream subsystem and the target subsystem based on the parallel probability.

[0005] According to a second aspect of the present disclosure, a request processing device is provided, comprising: a load determination module for determining the idle load of a target subsystem in response to a received user request; a probability calculation module for determining a parallel probability based on the current inlet traffic and idle load of the processing system; the processing system comprising a target subsystem and an upstream subsystem of the target subsystem; and a request processing module for generating a request processing result of a user request using the upstream subsystem and the target subsystem based on the parallel probability.

[0006] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.

[0007] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0008] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements any method according to the embodiments of the present disclosure when executed by a processor.

[0009] The solution disclosed in this disclosure can improve resource utilization during periods of low traffic.

[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0012] Figure 1 is a flowchart of a request processing method according to an embodiment of the present disclosure;

[0013] Figure 2 is a schematic diagram of load changes of a subsystem according to an embodiment of the present disclosure;

[0014] Figure 3 is a flow chart of generating request processing results based on parallel probability according to an embodiment of the present disclosure;

[0015] Figure 4 is a schematic diagram of a processing flow of a serial call according to an embodiment of the present disclosure;

[0016] Figure 5 is a timing diagram of serial calls according to an embodiment of the present disclosure;

[0017] Figure 6 is a schematic diagram of a processing flow of parallel calls according to an embodiment of the present disclosure;

[0018] Figure 7 is a timing diagram of parallel calls according to an embodiment of the present disclosure;

[0019] Figure 8 is a schematic structural diagram of a request processing device according to an embodiment of the present disclosure;

[0020] Figure 9 is a scenario diagram of a request processing method according to an embodiment of the present disclosure;

[0021] Figure 10 4 is a structural diagram of an electronic device used to implement the request processing method of an embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0023] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this article refer to multiple similar technical terms and distinguish them, and do not mean to limit the order or to limit to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature can be one or more, and the second feature can also be one or more.

[0024] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0025] Before introducing the technical solutions of the embodiments of the present disclosure, the following technical terms that may be used in the present disclosure are further explained:

[0026] A large-scale AI system is composed of multiple subsystems, typically small models or modular components. Large-scale systems are typically designed to handle complex tasks by breaking them down into multiple subtasks and assigning each subtask to a different subsystem. Overall, a large-scale system can be viewed as a distributed AI architecture working in concert.

[0027] Large-scale model systems usually require multiple subsystems to work together when processing a single event, and each subsystem is responsible for handling different specialized functions. Due to the dependencies between subsystems, the processing results of the upstream subsystem will determine the type and calling method of the downstream subsystem. The traffic of an upstream subsystem is equal to the sum of the traffic of all its child nodes, and the total time consumed for event processing is the sum of the time consumed by all subsystems. In actual applications, the hardware resources of each subsystem are usually configured with capacity according to peak traffic. Existing technologies can organize subsystems into pipelines in the form of call chains and arrange subsystems as nodes according to dependencies. A serial execution strategy is adopted within the link, while a parallel execution strategy is adopted between links. However, the existing technology solutions have the defect of being unable to adjust the concurrency in real time according to traffic, which limits the system's responsiveness in dynamic traffic scenarios.

[0028] In order to at least partially solve one or more of the above problems and other potential problems, the present disclosure proposes a request processing method that can improve resource utilization during low traffic periods.

[0029] The present disclosure provides a request processing method. Figure 1 It is a flow chart of a request processing method according to an embodiment of the present disclosure, and the request processing method can be applied to a request processing device. The request processing device is located in an electronic device. The electronic device includes but is not limited to fixed devices and / or mobile devices. For example, fixed devices include but are not limited to servers, and servers can be cloud servers or ordinary servers. For example, mobile devices include but are not limited to large model dialogue devices, and large model dialogue devices can be mobile phones, tablet computers, vehicle-mounted terminals, etc. In some possible implementations, the request processing method can also be implemented by a processor calling computer-readable instructions stored in a memory. For example Figure 1 As shown, the request processing method includes:

[0030] S101 : In response to a received user request, determine the idle load of a target subsystem.

[0031] S102: Determine the parallel probability according to the current inlet traffic and idle load of the processing system.

[0032] S103 : Based on the parallel probability, the upstream subsystem and the target subsystem are used to generate a request processing result of the user request.

[0033] The user request may be an instruction or demand sent by the user, such as a search query, voice input, a recommendation service request, etc. In the embodiment of the present disclosure, the user request may be a task that the system needs to process, which usually triggers a call chain of the system to complete the request processing.

[0034] A subsystem is a model or component within a larger model system that handles a specific function or task. A target subsystem is the specific subsystem that needs to be called by the user request to complete a specific subtask.

[0035] The idle load refers to the computing resources or processing capacity of the target subsystem that are not currently being used. In the embodiment of the present disclosure, the idle load represents the amount of resources available to the subsystem at the current moment, and is used to measure the load status of the system.

[0036] In the embodiment of the present disclosure, a user request may be received first. Exemplarily, the user request may be received through an interface such as an Application Programming Interface (API), front-end interaction, or internal triggering of the system, wherein the user request generally includes specific parameters, such as request type, related data, and the like. Subsequently, the user request may be mapped to a specific subsystem based on the user request, with the subsystem being used as the target subsystem. Exemplarily, the request mapping process may be implemented through predefined routing rules. Finally, a reporting mechanism may be pre-set so that the subsystem can report its own resource usage to the central scheduling module or monitoring system in real time, thereby obtaining the current idle load of the target subsystem by querying the monitoring system. The above is merely an exemplary explanation and is not intended to limit all possible situations for determining the idle load, but is not exhaustive here.

[0037] The processing system may refer to a large model system, which is a complete computing system or architecture that can be used to receive, process and respond to user requests. In the embodiment of the present disclosure, the processing system includes a target subsystem and an upstream subsystem of the target subsystem.

[0038] The upstream subsystem refers to the subsystem in the request processing chain that is responsible for the previous step of data processing relative to the target subsystem.

[0039] The current inlet traffic represents the number of user requests received by the system at the current moment and can be expressed in queries per second (QPS). In the disclosed embodiments, the current inlet traffic can indicate that the system is in a certain traffic state, such as a traffic peak or a traffic trough. The current traffic state can affect the resource allocation strategy.

[0040] The parallel probability is a probability value used to determine whether to assign a request processing task to multiple subsystems for parallel processing. In the embodiment of the present disclosure, the parallel probability can be calculated based on traffic and idle load and used to dynamically adjust the system's concurrency to optimize processing efficiency.

[0041] In the disclosed embodiment, the traffic monitoring module can first count the number of requests received by the system in real time, using this number as the current ingress traffic. Subsequently, a probability value can be calculated based on the current ingress traffic and the idle load of the target subsystem to guide whether to perform parallel processing. This probability value is used as the parallel probability. The above is merely an example and does not limit all possible scenarios for determining the parallel probability. This is simply not an exhaustive list.

[0042] The request processing result refers to the output of the user request processing, which is the result generated after the system completes the task. In the embodiment of the present disclosure, the request processing result can be the final output of the user request processing process, or it can be an intermediate output of the user request processing process.

[0043] In the disclosed embodiments, a request can first be split into multiple subtasks based on the parallel probability and distributed to the target subsystem and upstream subsystems for parallel processing; or it can be processed step by step according to the established call chain without splitting the tasks. Subsequently, the final request processing result is generated based on the processing results of the target subsystem and upstream subsystem. The above is only an example and does not limit all possible situations for generating request processing results. It is just not an exhaustive list here.

[0044] The technical solution of the disclosed embodiments can dynamically adjust the parallel probability based on traffic and load, making more efficient use of system resources and avoiding idle or overloaded resources. This allows for a dynamic adjustment strategy that combines traffic monitoring and load conditions, avoiding system crashes or instability caused by insufficient or overallocated resources. By calculating the parallel probability, during peak traffic times, serial processing can conserve resources and reduce hardware and energy costs. During low traffic times, increased concurrency can be used to quickly respond to user requests, reducing wait times while increasing the utilization of idle resources and improving the system's service quality.

[0045] In some embodiments, the request processing method further includes: obtaining architecture information of the processing system; determining a set of independent subsystems based on the architecture information; and determining a target subsystem from the set of independent subsystems.

[0046] Architectural information refers to descriptive data or meta-information that describes the structure of the entire processing system and how its components collaborate. In the disclosed embodiments, architectural information can be used to analyze and determine the relationships between subsystems, particularly dependencies, and thereby determine which subsystems can execute independently or need to wait for other subsystems to complete before starting. For example, architectural information may include information such as the organization of each subsystem, its dependencies, input and output interface definitions, and the distribution of functional modules.

[0047] In the disclosed embodiments, architecture information is typically defined and recorded during the system design phase and stored in a configuration file or database. Therefore, architecture information can be first obtained by invoking a system management module or configuration management tool. The above description is merely illustrative and does not limit all possible scenarios for obtaining architecture information. This is not intended to be exhaustive.

[0048] An independent subsystem is one whose execution does not depend on the results of other upstream subsystems and can operate independently. In the disclosed embodiments, an independent subsystem does not need to wait for the completion of an upstream subsystem and can be started in advance, executing simultaneously with the upstream subsystem. Furthermore, the results of the upstream subsystem only determine whether the independent subsystem is executed, and are not an input prerequisite for the independent subsystem. In other words, the startup of an independent subsystem is not constrained by other subsystems.

[0049] In the disclosed embodiment, the architecture information can first be parsed to extract the dependencies, call chain structure, and execution logic between subsystems. The subsystem call chains in the architecture information are then analyzed to determine whether each subsystem depends on the execution results of upstream subsystems. Finally, all subsystems that meet the requirements are collected to form a set of independent subsystems. The above is merely an example and does not limit all possible scenarios for determining a set of independent subsystems. This is simply not an exhaustive list.

[0050] In the disclosed embodiment, based on the specific task requested by the user, a subsystem suitable for processing the current request can be selected from the set of independent subsystems and the corresponding subsystem can be selected as the target subsystem. The above description is merely illustrative and does not limit all possible scenarios for determining the target subsystem. This is not intended to be exhaustive.

[0051] By identifying independent subsystems and allowing them to execute ahead of time, the overall time required for request processing can be reduced. By analyzing architectural information, we can identify subsystems that can execute independently, optimizing the system's parallel capabilities. In low-traffic scenarios, independent subsystems can run simultaneously with upstream subsystems, improving concurrency, rapidly responding to user requests, reducing wait times, and increasing the utilization of idle resources and throughput. By analyzing architectural information, we can clarify the dependencies between subsystems, simplifying system design and optimization. Preemptively identifying independent subsystems can also reduce the complexity of the task processing chain.

[0052] In some embodiments, determining the idle load of the target subsystem includes: obtaining the total capacity of the target subsystem and the occupied load of the target subsystem; and determining the idle load according to the total capacity and the occupied load.

[0053] The total capacity refers to the maximum processing capability of the subsystem in terms of hardware resources and software resources. In the embodiment of the present disclosure, the total capacity is the upper limit of the load that the target subsystem can handle under ideal circumstances.

[0054] The occupied load is the amount of resources currently actually used or required by the subsystem. In the disclosed embodiment, the occupied load can be the hardware and software resources occupied by the task being processed by the target subsystem; it can also be the load that the current task needs to occupy when it is executed serially, that is, assuming the resources required when the task is executed serially, which can be used to estimate the minimum resource requirements for task execution; it can also be the sum of the current actual occupied load and the required occupied load, that is, the sum of the occupied load of the currently running task and the resources required by the new task, which can be used to evaluate the total load of the target subsystem after the new task is added.

[0055] In the disclosed embodiments, the total capacity of a subsystem is typically defined during system initialization or configuration. Thus, the total capacity of the target subsystem can be read from the configuration file. Specifically, if the target subsystem supports dynamic expansion, the total capacity can be calculated in real time, and the calculated total capacity can be used as the total capacity of the target subsystem. Subsequently, the occupied load can be defined based on actual conditions, and the occupied load can be determined based on the definition of occupied load. For example, when the occupied load refers to the hardware and software resources occupied by the tasks currently being processed by the target subsystem, the subsystem's monitoring module can be used to query the number of tasks currently being processed and resource usage in real time to obtain the current actual occupied load, and use the actual occupied load as the occupied load. For example, when the occupied load refers to the load required for the current task to be executed serially, the current task can be assumed to be executed serially, and the resources required to be occupied can be calculated, with the calculated required resources being used as the occupied load. For example, when the occupied load refers to the sum of the current actual occupied load and the required load, after obtaining the current actual occupied load, the total load after the new task is added can be evaluated, and the evaluated total load can be used as the occupied load. The above is only an example and is not intended to limit all possible situations for obtaining total capacity and occupied load. This is just not an exhaustive list.

[0056] In the embodiment of the present disclosure, the idle resources of the subsystem can be calculated based on the total capacity and the occupied load, and the idle resources can be used as the idle load. For example, the difference between the total capacity and the occupied load can be used as the idle load. In particular, if the idle load is zero, it means that the subsystem is running at full load. Similarly, if the calculation result is a negative value, that is, the occupied load exceeds the total capacity, it means that the subsystem resources are insufficient and resource scheduling may be required. For example, resource scheduling can be performed by means of capacity expansion or task delay. The above is only an illustrative description and is not intended to limit all possible situations for determining the idle load, but it is not exhaustive here.

[0057] In this way, by monitoring the total capacity and occupied load of the target subsystem in real time, the resource status of the subsystem can be understood at any time, which helps the system to respond quickly. By calculating the idle load, the system can reasonably allocate resources to avoid idle or overloaded resources. By calculating the occupied load and idle load, the system can determine the degree of task parallelism. The higher the idle load, the more tasks the system can arrange for parallel processing, thereby improving overall throughput. When scheduling tasks, giving priority to subsystems with higher idle loads can reduce task waiting time and improve task processing efficiency. Especially in low-load scenarios, calculating the idle load helps to fully utilize computing resources and accelerate task completion.

[0058] In some implementations, the idle load of the target subsystem may be calculated using the following formula:

[0059] Q idle =CQ sub

[0060] Where C represents the total capacity of the target subsystem; Q sub Indicates the occupied load of the target subsystem; Q idle Indicates the idle load of the target subsystem.

[0061] Figure 2 Figure 2 shows a schematic diagram of the load changes of the subsystem. Figure 2 As shown, when the total capacity of the subsystem is C, at any moment, the occupied load of the subsystem is Q sub When the idle load of the subsystem is Q idle .

[0062] In some embodiments, determining the parallel probability based on the current inlet traffic and idle load of the processing system includes: determining the idle load rate of the target subsystem based on the current inlet traffic and idle load; and determining the parallel probability based on the idle load rate.

[0063] The idle load ratio is the percentage of the target subsystem's currently available computing resources to its total resource capacity, reflecting the remaining resource situation of the target subsystem. In the disclosed embodiments, the idle load ratio is a dynamic indicator that can be used to measure the system's load status and guide task scheduling strategies.

[0064] In the embodiment of the present disclosure, the current number of user requests can first be counted in real time through the traffic monitoring module, and the statistical results can be used as the current inlet traffic. In particular, the current inlet traffic represents the source of the current task load of the system and is a dynamically changing indicator. Subsequently, the idle load rate of the target subsystem can be calculated based on the idle load and the current inlet traffic. For example, the ratio between the displayed load and the current inlet traffic can be used as the idle load rate. The above is only an exemplary explanation and is not intended to limit all possible situations for determining the idle load rate. It is just that this is not an exhaustive list.

[0065] In the disclosed embodiment, the parallel probability is dynamically calculated based on the idle load rate and predefined rules. For example, a linear function or a nonlinear function can be used for calculation, and interval probability values ​​can be defined in combination with a configuration file to obtain the parallel probability. In particular, the value range of the parallel probability is generally 0%-100%, indicating the possibility that the target subsystem and the upstream subsystem will execute in parallel. The above is only an example and is not intended to limit all possible situations for determining the parallel probability, but it is not intended to be exhaustive.

[0066] By calculating the idle load ratio and parallel probability, resource allocation strategies can be dynamically adjusted to ensure efficient utilization of computing resources. By calculating the idle load ratio, the system can predict resource availability in advance, preventing subsystem overload caused by task allocation. It also limits parallel processing when resources are insufficient, improving system stability. When ingress traffic fluctuates, the parallel probability can be adjusted based on the real-time idle load ratio, dynamically optimizing task scheduling strategies.

[0067] In some implementations, the idle load rate of the target subsystem may be calculated using the following formula:

[0068]

[0069] Where, P idle Indicates the idle load rate of the target subsystem; Q idle Indicates the idle load of the subsystem; Q all Indicates the current ingress traffic.

[0070] Specifically, P idle It can represent the ratio between the current idle load of the target subsystem and the current inlet flow. idle When it is greater than 1, it means that the idle load of the target subsystem can fully process all requests of the current inlet traffic; when P idle When it is greater than 0 but less than 1, it means that the target subsystem has idle load at this time, but the idle load at this time cannot fully process all requests of the current inlet traffic and can only process part of the requests; when P idle When it is less than 0, it means that the target subsystem has insufficient resources and no idle load, and parallel processing needs to be limited.

[0071] Furthermore, the parallel probability can be calculated by the following formula:

[0072] P d =min(P idle , 1)×100%

[0073] Where, P d represents the parallel probability.

[0074] Specifically, when P idle When it is greater than 1, take P idle The minimum value between 1 and 1 is 1, and the parallel probability is 100%. That is, when the idle load of the target subsystem can fully process all requests of the current inlet traffic, parallel scheduling must be adopted; when P idle When it is greater than 0 but less than 1, take P idle The minimum value between P and 1 is idle , that is, when the target subsystem has idle load, but the idle load at this time cannot fully process all requests of the current inlet traffic, P idle ×100% as the probability value, choose to adopt parallel scheduling or serial scheduling; when P idle When it is less than 0, take P idle The minimum value between P and 1 is idle , that is, when the target subsystem resources are insufficient and there is not enough idle load, parallel processing needs to be limited, and the probability of parallel processing is less than or equal to 0%, that is, serial processing must be used. d The additional load introduced to the target subsystem is not included in Q sub .

[0075] In some embodiments, based on the parallel probability, the upstream subsystem and the target subsystem are used to generate the request processing result of the user request, including: determining the calling method according to the parallel probability; based on the calling method, using the upstream subsystem and the target subsystem to generate the request processing result of the user request.

[0076] The calling method refers to the strategy by which the system decides how to call the upstream subsystem and the target subsystem to process the request when processing a user request, including the order of calls, parallel and serial processing methods, and how to coordinate the execution between multiple subsystems. In the embodiment of the present disclosure, the calling method includes serial calls and parallel calls. Specifically, serial calls refer to calling the upstream subsystem and the target subsystem in sequence. Parallel calls refer to the simultaneous execution of the upstream subsystem and the target subsystem without blocking each other.

[0077] In the embodiment of the present disclosure, a suitable calling method can be selected according to the parallel probability. For example, when the parallel probability is 100%, parallel calling is selected; when the parallel probability is less than or equal to 0%, serial calling is selected. Furthermore, when the parallel probability is greater than v% and less than 100%, a judgment can be made according to the parallel probability threshold, that is, when the parallel probability is greater than or equal to the parallel probability threshold, parallel calling is selected; when the parallel probability is less than the parallel probability threshold, serial calling is selected. Similarly, random number judgment can also be made according to the parallel probability, that is, a random number generator is used to generate a random number, and the interval of the random number is divided according to the parallel probability. When the random number falls within the parallel probability interval, parallel calling is selected, and when the random number falls outside the parallel probability interval, serial calling is selected. The above is only an exemplary explanation and is not intended to limit all possible situations for determining the calling method. It is just that an exhaustive list is not given here.

[0078] In the disclosed embodiments, the execution process of the upstream subsystem and the target subsystem can be initiated according to the calling method, and the processing results of the upstream subsystem and the target subsystem can be aggregated to generate a complete user request processing result. The above is only an example and does not limit all possible situations for generating request processing results. However, this is not an exhaustive list.

[0079] By dynamically selecting the calling method, the system maximizes resource utilization and reduces overall task processing time. When the target subsystem is independent and has a high probability of parallelization, it can fully utilize the subsystem's idle resources for parallel processing, avoiding resource waste. By determining the calling method based on the probability of parallelization, tasks can be distributed across multiple subsystems when the probability of parallelization is high, preventing overloading of a single subsystem. When the probability of parallelization is low, the concurrency of tasks is controlled, prioritizing sequential processing of tasks and reducing system pressure.

[0080] In some embodiments, based on the calling method, according to the user request, the upstream subsystem and the target subsystem are used to generate the request processing result, including: according to the calling method, the upstream subsystem and the target subsystem are used to obtain the upstream processing result generated by the upstream subsystem and the target processing result generated by the target subsystem according to the calling method; based on the upstream processing result and the target processing result, the request processing result is generated.

[0081] Among them, the upstream processing result is an intermediate processing result generated by the upstream subsystem. In the embodiment of the present disclosure, the upstream processing result is the execution basis of the target subsystem, but it is not used as input data of the target subsystem. Specifically, the upstream processing result can be a binary judgment result generated by the upstream subsystem after preliminary processing of the user request. Exemplarily, the upstream processing result can be calling the target subsystem or not calling the target subsystem. In particular, the upstream processing result can be represented by a Boolean value; when the upstream processing result is calling the target subsystem, True is output; conversely, when the upstream processing result is not calling the target subsystem, False is output.

[0082] The target processing result is a processing result generated by the target subsystem. In the embodiment of the present disclosure, the target processing result can be a single result or a composite result.

[0083] In the disclosed embodiments, an upstream subsystem can be called to process a task based on the content of a user request, generating an upstream processing result. Furthermore, depending on the calling method, a target subsystem can be used to generate a target processing result simultaneously with the upstream subsystem being called or based on the upstream processing result. The above description is merely illustrative and does not limit all possible scenarios for generating upstream processing results and target processing results, but this is not intended to be exhaustive.

[0084] In the disclosed embodiments, the upstream processing results and the target processing results can be integrated to generate the final request processing result. In particular, in some scenarios, the request processing result may require further optimization, such as sorting, filtering, etc. In this case, the integrated and optimized result can be output as the final request processing result. The above is only an example and does not limit all possible situations for generating request processing results. This is just not an exhaustive list.

[0085] By dynamically calling upstream and target subsystems, the system can rationally allocate tasks based on the calling method, improving task processing efficiency. Different calling methods can flexibly adapt to different scenarios and task requirements. By integrating upstream and target processing results, the resulting request processing results are more complete and targeted.

[0086] In some embodiments, according to the calling method, the upstream subsystem and the target subsystem are used to generate upstream processing results and target processing results, including: when the calling method is a serial call, the upstream subsystem is called to generate the upstream processing results; according to the upstream processing results, the target subsystem is used to generate the target processing results.

[0087] The upstream processing result may be a binary judgment result generated by the upstream subsystem after preliminary processing of the user request. For example, the upstream processing result may include the need to call the target subsystem.

[0088] In the disclosed embodiments, an upstream subsystem may be first invoked for processing, and the processing result generated by the upstream subsystem may be used as the upstream processing result. In particular, the tasks of the upstream subsystem may include data preprocessing, analysis, and / or generation of intermediate results. The above description is merely illustrative and does not limit all possible scenarios for generating upstream processing results, but this is not intended to be exhaustive.

[0089] In the embodiment of the present disclosure, the target subsystem can be used to generate the target processing result according to the content of the upstream processing result. The above is only an example and does not limit all possible situations for generating the target processing result. It is just not exhaustive here.

[0090] In this way, the serial calling method can ensure the execution order of tasks, reduce the resource pressure of the subsystem when the subsystem is idle and underloaded, and avoid resource conflicts within the subsystem.

[0091] In some embodiments, the target processing result is generated by using the target subsystem according to the upstream processing result, including: when the upstream processing result requires calling the target subsystem, calling the target subsystem to generate the target processing result.

[0092] In the disclosed embodiments, if the determination result indicates that the target subsystem needs to be called, the execution process of the target subsystem can be initiated. For example, when the determination result indicates that the target subsystem needs to be called, the target subsystem can be called to process the original information or global configuration information requested by the user, thereby generating the target processing result. The above is merely an example and does not limit all possible scenarios for generating the target processing result, but this is not intended to be an exhaustive list.

[0093] In this way, the upstream subsystem makes a preliminary judgment to decide whether to call the target subsystem. The target subsystem, as an independent module, can independently generate the target processing results, thereby effectively improving the efficiency and resource utilization of the system, while ensuring the quality of the processing results and the flexibility of the system.

[0094] In some embodiments, the target processing result is generated using the target subsystem based on the upstream processing result, including: when the processing result of the upstream subsystem does not require calling the target subsystem, the target processing result is set to a default value using the target subsystem.

[0095] In the disclosed embodiments, if the upstream processing result indicates that the target subsystem does not need to be called, the complex logic of the target subsystem can be directly skipped and the target processing result can be set to a default value. In particular, the default value is typically a predefined static value or an empty value to handle scenarios where no result is obtained or the task is not triggered. The above is merely an example and does not limit all possible scenarios for generating the target processing result, but it is not intended to be exhaustive.

[0096] In this way, the upstream subsystem generates processing results to determine whether the target subsystem needs to be called. If the target subsystem does not need to be called, the system directly sets the target processing result to the default value, effectively avoiding resource waste, improving processing efficiency, and ensuring system stability and flexibility. By setting default values, user requests can be quickly responded to, improving the user experience, and simplifying the logical design of the target subsystem.

[0097] In some embodiments, generating a request processing result based on the upstream processing result and the target processing result includes: merging the upstream processing result and the target processing result to generate the request processing result.

[0098] In the disclosed embodiments, merging rules can be defined based on business needs to integrate upstream processing results and target processing results into a unified request processing result. For example, a direct combination method can be used to directly combine upstream processing results and target processing results into a complete result. The above is merely an example and does not limit all possible situations for generating request processing results. This is just not an exhaustive list.

[0099] By merging the upstream processing results with the target processing results to generate the request processing result, we can significantly improve the quality of the processing results and the user experience, while optimizing resource utilization and system design logic. Furthermore, the merging logic can be dynamically adjusted to meet diverse business needs while ensuring system stability and scalability.

[0100] In some embodiments, according to the calling method, the upstream subsystem and the target subsystem are used to generate upstream processing results and target processing results, including: when the calling method is a parallel call, the upstream subsystem and the target subsystem are called at the same time to obtain the upstream processing results generated by the upstream subsystem and the target processing results generated by the target subsystem.

[0101] In the disclosed embodiments, when the call mode is parallel, the upstream subsystem and the target subsystem can be started simultaneously, and both independently perform their respective tasks. That is, the upstream subsystem and the target subsystem run simultaneously and each completes its task. Subsequently, the processing results can be obtained from the upstream subsystem and the target subsystem respectively, resulting in the upstream processing result and the target processing result. The above is only an example and does not limit all possible situations in which the upstream processing result and the target processing result can be generated. However, this is not an exhaustive list.

[0102] In this way, by simultaneously starting the upstream subsystem and the target subsystem, the system processing efficiency is significantly improved, resource utilization is optimized, and the system coupling is reduced, thereby enabling a quick response to user requests while maintaining flexibility and scalability. This is an efficient and reliable task processing strategy.

[0103] In some embodiments, a request processing result is generated based on the upstream processing result and the target processing result, including: when the upstream processing result requires calling the target subsystem, the upstream processing result and the target processing result are merged to generate the request processing result.

[0104] In the embodiments of the present disclosure, when the upstream processing result requires calling the target subsystem, a merge rule can be defined based on business needs to integrate the upstream processing result and the target processing result into a unified request processing result. For example, a direct combination method can be used to directly combine the upstream processing result and the target processing result into a complete result. The above is only an example and does not limit all possible situations for generating request processing results. It is just not an exhaustive list here.

[0105] This approach leverages information from both upstream and target subsystems by determining whether to call the target subsystem and, if necessary, merging upstream and target processing results. This improves processing efficiency and result quality, while optimizing resource utilization and enhancing the user experience. Through layered design and dynamic merging logic, the system enhances overall scalability and flexibility, making it suitable for a variety of application scenarios.

[0106] In some embodiments, a request processing result is generated based on the upstream processing result and the target processing result, including: when the processing result of the upstream subsystem does not require calling the target subsystem, a request processing result is generated based on the upstream processing result.

[0107] In the disclosed embodiments, when the upstream processing result indicates that the target subsystem does not need to be called, the complex logic of the target subsystem can be skipped, and the request processing result can be directly generated using only the information generated by the upstream subsystem. The above is only an example and does not limit all possible situations for generating request processing results. This is just not an exhaustive list.

[0108] In this way, the need to call the target subsystem is determined by the upstream processing results. When the upstream processing result shows that the target subsystem does not need to be called, the complex logic of the target subsystem is directly skipped, and the request processing result is generated directly according to the upstream processing result, thereby avoiding unnecessary waste of resources, improving system processing efficiency and stability, and improving user experience.

[0109] In some embodiments, the large model system may include a determination subsystem and a rewriting subsystem. For the natural language of a user request, the determination subsystem is used to determine whether the user request requires an online search. If an online search is required, the rewriting subsystem is called to rewrite the natural language to obtain the rewritten content. The results of the determination subsystem are then integrated with the rewritten content to generate prompt words that the large language model can understand. Conversely, if an online search is not required, the rewriting subsystem does not need to be called, and the prompt words can be directly generated using the results of the determination subsystem. Specifically, the determination subsystem can be used as the upstream subsystem, and the rewriting subsystem can be used as the target subsystem, and they can be called based on the parallel probability.

[0110] For example, when the call method is serial, the judgment subsystem is first called to generate a judgment result, which is used as the upstream processing result. Subsequently, if the upstream processing result is that an online search is required, the rewriting subsystem is called to generate rewritten content, which is used as the target processing result. The online search and rewriting content are merged, and the merged result is "Internet connection required + rewritten content". A prompt word is generated based on the merged result, and the prompt word is used as the request processing result. Conversely, if the upstream processing result is that an online search is not required, the processing of the rewriting subsystem is skipped, the default null value is used as the target processing result, the no-network-need and null values ​​are merged, and the merged result is "no-network-needed". A prompt word is generated based on the merged result, and the prompt word is used as the request processing result.

[0111] For example, when the call method is parallel, the judgment subsystem is called to generate a judgment result, which is used as the upstream processing result. Simultaneously, the rewriting subsystem is called to generate rewritten content, which is used as the upstream processing result. Subsequently, if the upstream processing result indicates that an online search is required, the online search requirement and the rewritten content are merged. The merged result is "Internet required + rewritten content." A prompt word is generated based on the merged result, and the prompt word is used as the request processing result. Conversely, if the upstream processing result indicates that an online search is not required, the target processing result is discarded, and a prompt word is directly generated based on the upstream processing result, and the prompt word is used as the request processing result.

[0112] Figure 3 A schematic diagram of a process for generating request processing results based on parallel probability is shown, as shown in FIG. Figure 3 Shown, including:

[0113] S301, receiving a user request;

[0114] S302a, calling the upstream processing subsystem to generate upstream processing results;

[0115] S302b, according to the parallel probability P d Determine the calling method; there is P d% probability call mode is parallel call, then go to S303; there is (100%-P d %) is a serial call, and then go to S307;

[0116] S303, calling the target processing subsystem to generate a target processing result;

[0117] S304: Determine whether the target processing result is needed based on the upstream processing result; if so, go to S305; if not, go to S306;

[0118] S305: Merge the upstream processing result and the target processing result, and go to S311;

[0119] S306, discard the target processing result and go to S311;

[0120] S307: Determine whether the target processing subsystem needs to be called based on the upstream processing result; if so, go to S308; if not, go to S310;

[0121] S308, call the target processing subsystem, generate the target processing result, and go to S311;

[0122] S309: Merge the upstream processing result and the target processing result, and go to S311;

[0123] S310, skip the target subsystem and go to S311;

[0124] S311. Generate request processing results.

[0125] Figure 4 A schematic diagram of the processing flow of serial calls is shown, as Figure 4 Shown, including:

[0126] S401, receiving a user request;

[0127] S402: Call the upstream subsystem to generate upstream processing results;

[0128] S403, judging whether the target subsystem needs to be called according to the upstream processing result; if the target subsystem needs to be called, go to S404; if the target subsystem does not need to be called, go to S405;

[0129] S404, calling the target subsystem to generate the target processing result;

[0130] S405: Generate request processing result.

[0131] Figure 5 A timing diagram of serial calls is shown, as Figure 5As shown, the overall time consumption of the processing system during serial calls can be expressed as follows:

[0132] C S =C1+p%×C2

[0133] Where C S Indicates the overall time consumption of serial calls; C1 indicates the processing time consumption of the upstream subsystem; p% indicates the probability of needing to call the target subsystem; C2 indicates the processing time consumption of the target subsystem.

[0134] Specifically, in a serial call, the upstream subsystem is first called, and then the target subsystem is called based on the upstream processing result. Therefore, it is necessary to first wait for the upstream subsystem to generate the upstream processing result after the C1 time has passed. Subsequently, there is a p% probability that it will be necessary to continue waiting for the target subsystem to generate the target processing result after the C2 time has passed before combining the upstream processing result and the target processing result to generate the request processing result. At the same time, there is a (100% - p%) probability that it is not necessary to wait for the target subsystem. In this case, after waiting for the C1 time, the request processing result can be directly generated based on the upstream processing result.

[0135] Figure 6 A schematic diagram of the processing flow of serial calls is shown, as Figure 6 Shown, including:

[0136] S601, receiving a user request;

[0137] S602a, calling the upstream subsystem to generate upstream processing results;

[0138] S602b, calling the target subsystem to generate the target processing result;

[0139] S603: Determine whether to use the target processing result based on the upstream processing result;

[0140] S604: Generate request processing results.

[0141] Figure 7 A timing diagram of parallel calls is shown, as Figure 7 As shown, the overall time consumption of the processing system during parallel calls can be expressed as follows:

[0142] C P =p%×[max(C1,C2)]+(1-p%)×C1

[0143] Where C P Indicates the overall time consumption of parallel calls; C1 indicates the processing time consumption of the upstream subsystem; p% indicates the probability of needing to call the target subsystem; C2 indicates the processing time consumption of the target subsystem; max(C1, C2) indicates the maximum value between C1 and C2.

[0144] Specifically, because parallel calls involve calling both the upstream and target subsystems simultaneously, there's a p% probability that the call will need to wait for both the upstream and target subsystems to complete processing, meaning it will need to wait for the maximum of C1 and C2. At the same time, there's a (1-p%) probability that the call will not need to wait for the target subsystem to complete processing, but only needs to wait for C2, ensuring that the upstream subsystem completes processing.

[0145] Based on the overall time consumption during serial calls and the overall time consumption during parallel calls, we can see that the average time consumption reduced by parallel execution compared to serial execution can be expressed as follows:

[0146] ΔC=C S -C P

[0147] Where ΔC represents the time reduction of parallel execution compared to serial execution.

[0148] Furthermore, C S and C P After substituting and sorting, we can get:

[0149] ΔC = p% × [min(C1, C2)]

[0150] Wherein, min(C1, C2) represents the minimum value between C1 and C2.

[0151] It can be seen from this that the technical solution of the embodiment of the present disclosure can improve concurrency, shorten processing time, quickly respond to user requests when traffic is low, and at the same time improve the utilization rate of idle resources and enhance the service quality of the system.

[0152] It should be understood that Figures 2 to 7 The schematic diagram shown is only exemplary and not restrictive, and it is scalable, and those skilled in the art can Figures 2 to 7 Various obvious changes and / or substitutions can be made to the examples, and the resulting technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0153] The embodiment of the present disclosure provides a request processing device, such as Figure 8 As shown, the device may include: a load determination module 801, which is used to determine the idle load of the target subsystem in response to a received user request; a probability calculation module 802, which is used to determine the parallel probability based on the current inlet traffic and idle load of the processing system; the processing system includes a target subsystem and an upstream subsystem of the target subsystem; and a request processing module 803, which is used to generate a request processing result of the user request using the upstream subsystem and the target subsystem based on the parallel probability.

[0154] In some embodiments, the request processing device further includes: an architecture information acquisition module 804 ( Figure 8 ), used to obtain the architecture information of the processing system; independent system determination module 805 ( Figure 8 ), for determining a set of independent subsystems according to the architecture information; a target subsystem determination module 806 ( Figure 8 (not shown) is used to determine the target subsystem from the set of independent subsystems.

[0155] In some embodiments, the load determination module 801 includes: a subsystem information acquisition submodule, used to obtain the total capacity of the target subsystem and the occupied load of the target subsystem; and an idle load determination submodule, used to determine the idle load based on the total capacity and the occupied load.

[0156] In some embodiments, the probability calculation module 802 includes: an idle load calculation submodule for determining the idle load rate of the target subsystem based on the current inlet traffic and idle load; and a parallel probability calculation submodule for determining the parallel probability based on the idle load rate.

[0157] In some embodiments, the request processing module 803 includes: a calling method determination submodule, which is used to determine the calling method based on the parallel probability; the calling method includes serial calling and parallel calling; and a processing result generation submodule, which is used to generate the request processing result of the user request based on the calling method using the upstream subsystem and the target subsystem.

[0158] In some embodiments, the processing result generation submodule is used to: use the upstream subsystem and the target subsystem according to the calling method to obtain the upstream processing result generated by the upstream subsystem and the target processing result generated by the target subsystem; generate the request processing result based on the upstream processing result and the target processing result.

[0159] In some embodiments, the processing result generation submodule is used to: when the calling mode is a serial call, call the upstream subsystem to generate the upstream processing result; and generate the target processing result by using the target subsystem according to the upstream processing result.

[0160] In some embodiments, the processing result generation submodule is used to: when the upstream processing result requires calling the target subsystem, call the target subsystem to generate the target processing result.

[0161] In some embodiments, the processing result generating submodule is used to: when the processing result of the upstream subsystem does not require calling the target subsystem, use the target subsystem to set the target processing result to a default value.

[0162] In some embodiments, the processing result generation submodule is used to merge the upstream processing result and the target processing result to generate a request processing result.

[0163] In some embodiments, the processing result generation submodule is used to: when the calling mode is parallel calling, simultaneously call the upstream subsystem and the target subsystem to obtain the upstream processing result generated by the upstream subsystem and the target processing result generated by the target subsystem.

[0164] In some embodiments, the processing result generation submodule is used to: when the upstream processing result requires calling the target subsystem, merge the upstream processing result and the target processing result to generate a request processing result.

[0165] In some embodiments, the processing result generating submodule is used to: when the processing result of the upstream subsystem is that there is no need to call the target subsystem, generate a request processing result according to the upstream processing result.

[0166] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0167] The request processing device in the disclosed embodiment can dynamically adjust the parallel probability based on traffic and load, making more efficient use of system resources and avoiding resource idleness or overload. This allows for a dynamic adjustment strategy that combines traffic monitoring and load conditions, avoiding system crashes or instability caused by insufficient or over-allocated resources. By calculating the parallel probability, during peak traffic times, serial processing can conserve resources and reduce hardware and energy costs. During low traffic times, increased concurrency can be used to quickly respond to user requests, reducing waiting times while increasing the utilization of idle resources and improving the system's service quality.

[0168] The embodiment of the present disclosure provides a scenario diagram of a request processing method, such as Figure 9 shown.

[0169] As mentioned above, the request processing method provided by the embodiment of the present disclosure is applied to electronic devices. The electronic devices are intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers.

[0170] Specifically, the electronic device can perform the following operations:

[0171] In response to a received user request, an idle load of a target subsystem is determined; a parallel probability is determined based on current inlet traffic and idle load of the processing system; the processing system includes a target subsystem and an upstream subsystem of the target subsystem; based on the parallel probability, a request processing result of the user request is generated using the upstream subsystem and the target subsystem.

[0172] It should be understood that Figure 9 The scene diagram shown is only illustrative and not restrictive. Those skilled in the art can Figure 9 Various obvious changes and / or substitutions can be made to the examples, and the resulting technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0173] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0174] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0175] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0176] like Figure 10 As shown, the device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0177] Various components in device 1000 are connected to I / O interface 1005, including an input unit 1006, such as a keyboard, mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, optical disk, etc.; and a communication unit 1009, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0178] The computing unit 1001 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as the request processing method. For example, in some embodiments, the request processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the request processing method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to execute the request processing method in any other appropriate manner (eg, by means of firmware).

[0179] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0180] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0181] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0182] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0183] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0184] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0185] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0186] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A request processing method, comprising: In response to the received user request, determining an idle load of the target subsystem; Determining a parallel probability according to a current inlet flow of a processing system and the idle load; the processing system includes the target subsystem and an upstream subsystem of the target subsystem; Based on the parallel probability, a request processing result of the user request is generated using the upstream subsystem and the target subsystem.

2. The method according to claim 1, wherein The method further comprises: Obtaining architectural information of the processing system; Determining a set of independent subsystems based on the architecture information; The target subsystem is determined from the set of independent subsystems.

3. The method according to claim 1, wherein Determining the idle load of the target subsystem includes: Obtaining the total capacity of the target subsystem and the occupied load of the target subsystem; The idle load is determined according to the total capacity and the occupied load.

4. The method according to claim 1, wherein The determining of the parallel probability according to the current inlet flow of the processing system and the idle load includes: determining an idle load rate of the target subsystem according to the current inlet flow and the idle load; The parallel probability is determined according to the idle load rate.

5. The method according to claim 1, wherein The generating, based on the parallel probability, a request processing result of the user request by using the upstream subsystem and the target subsystem includes: Determine a calling mode according to the parallel probability; the calling mode includes serial calling and parallel calling; Based on the calling mode, the upstream subsystem and the target subsystem are used to generate a request processing result of the user request.

6. The method according to claim 5, wherein: The generating, based on the calling mode, a request processing result of the user request by using the upstream subsystem and the target subsystem includes: According to the calling mode, using the upstream subsystem and the target subsystem, obtaining the upstream processing result generated by the upstream subsystem and the target processing result generated by the target subsystem; The request processing result is generated according to the upstream processing result and the target processing result.

7. The method according to claim 6, wherein: The method of obtaining, by using the upstream subsystem and the target subsystem according to the calling mode, an upstream processing result generated by the upstream subsystem and a target processing result generated by the target subsystem includes: When the calling mode is a serial call, calling the upstream subsystem to generate an upstream processing result; According to the upstream processing result, the target subsystem is utilized to generate a target processing result.

8. The method according to claim 7, wherein: Generating a target processing result by using the target subsystem according to the upstream processing result includes: When the upstream processing result requires calling a target subsystem, the target subsystem is called to generate a target processing result.

9. The method according to claim 7, wherein: Generating a target processing result by using the target subsystem according to the upstream processing result includes: When the processing result of the upstream subsystem does not require calling the target subsystem, the target processing result is set as a default value using the target subsystem.

10. The method according to claim 8 or 9, wherein: Generating the request processing result according to the upstream processing result and the target processing result includes: The upstream processing result and the target processing result are merged to generate the request processing result.

11. The method according to claim 6, wherein: The method of obtaining, by using the upstream subsystem and the target subsystem according to the calling mode, an upstream processing result generated by the upstream subsystem and a target processing result generated by the target subsystem includes: When the calling mode is parallel calling, the upstream subsystem and the target subsystem are called simultaneously to obtain the upstream processing result generated by the upstream subsystem and the target processing result generated by the target subsystem.

12. The method according to claim 11, wherein Generating the request processing result according to the upstream processing result and the target processing result includes: When the upstream processing result requires calling a target subsystem, the upstream processing result and the target processing result are merged to generate the request processing result.

13. The method according to claim 11, wherein Generating the request processing result according to the upstream processing result and the target processing result includes: When the processing result of the upstream subsystem is that the target subsystem does not need to be called, the request processing result is generated according to the upstream processing result.

14. A request processing device, comprising: a load determination module, configured to determine an idle load of a target subsystem in response to a received user request; a probability calculation module, configured to determine a parallel probability based on a current inlet flow of a processing system and the idle load; the processing system comprising the target subsystem and an upstream subsystem of the target subsystem; A request processing module is used to generate a request processing result of the user request by using the upstream subsystem and the target subsystem based on the parallel probability.

15. An electronic device comprising: at least one processor; as well as a memory communicatively connected to at least one processor; wherein, The memory stores instructions that can be executed by at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 13.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are for causing a computer to perform a method according to any one of claims 1-13.

17. A computer program product comprising a computer program stored on a storage medium, the computer program implementing the method according to any one of claims 1 to 13 when executed by a processor.