Service request scheduling method and device, electronic equipment and storage medium
By comprehensively considering computing resources and network resources, selecting candidate service instances from the server cluster and calculating the optimal path, the problem of unbalanced resource utilization in the existing technology is solved, and the efficiency and user experience of service request scheduling are improved.
Patent Information
- Application Number
- CN202510706625.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-05
AI Technical Summary
Existing business request scheduling methods such as polling method and computing power priority method fail to effectively balance computing resources and network resources, resulting in unbalanced resource utilization and unstable service quality, affecting user experience.
By selecting candidate service instances from the server cluster, comprehensively considering computing resources and network resources, calculating the optimal path and forwarding service requests, we ensure that we have sufficient computing power and good network connections are selected.
It achieves the reduction of response time and processing delay while meeting business needs, and improves service quality and user experience.
Smart Images

Figure CN120434306A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a service request scheduling method, device, electronic device, and storage medium. Background Art
[0002] With the acceleration of digital transformation, enterprise business systems are facing an increasing volume of business requests. These requests not only come from a wide range of sources and cover multiple regions, but also exhibit varying computational complexity and response time requirements. Therefore, an efficient scheduling mechanism is required to accurately match the appropriate microservice for processing.
[0003] Current mainstream scheduling schemes, such as round-robin and compute-power priority, have significant limitations: the former uses a sequential allocation strategy, distributing traffic to candidate microservices in turn; the latter prioritizes traffic to microservices with abundant compute resources. Both approaches ignore the synergistic relationship between compute and network resources. In practice, these approaches have exposed problems such as unbalanced resource utilization and unstable service quality, which directly impact user experience. To effectively address the complex nature of business traffic, a more scientific business request scheduling method is urgently needed to ensure service quality and improve user satisfaction. Summary of the Invention
[0004] The present application provides a service request scheduling method, device, electronic device and storage medium, which can reduce response time and processing delay, and improve service quality and user experience satisfaction.
[0005] In a first aspect, the present application provides a service request scheduling method, the method comprising:
[0006] Selecting at least one candidate service instance for processing the service request from the server cluster based on the service request to be scheduled;
[0007] Obtain computing resources and network resources for each candidate service instance;
[0008] Selecting a target service instance from the candidate service instances based on the computing resources and the network resources;
[0009] Calculating an optimal path from the current service node to the target service instance based on the network resources;
[0010] The service request is forwarded to the target service instance based on the optimal path.
[0011] Furthermore, the selecting of at least one candidate service instance from the server cluster for processing the business request based on the business request to be scheduled includes: determining the service type corresponding to processing the business request based on the identification code of the business request; and selecting at least one service instance corresponding to the service type from the server cluster as a candidate service instance.
[0012] Furthermore, the selecting of a target service instance from the candidate service instances based on the computing power resources and the network resources includes: calculating the computing power performance score and resource utilization of each candidate service instance based on the computing power resources; calculating the load balancing score of each candidate service instance based on the resource utilization; calculating the score adjustment coefficient of each candidate service instance based on the network resources; calculating the score of each candidate service instance based on the computing power performance score, the load balancing score and the score adjustment coefficient; and selecting the candidate service instance whose score meets the preset criteria as the target service instance.
[0013] Furthermore, before calculating the computing power performance score and resource utilization of each candidate service instance based on the computing power resources, it also includes: if the number of the candidate service instances is less than the preset number, judging whether the computing power resources of each candidate service instance meet the preset computing power requirements and whether the network resources meet the preset network requirements; taking the candidate service instance that meets the preset computing power requirements and the preset network requirements as the alternative service instance; accordingly, calculating the computing power performance score and resource utilization of each candidate service instance based on the computing power resources includes: calculating the computing power performance score and resource utilization of each alternative service instance based on the computing power resources.
[0014] Furthermore, the types of computing power resources include at least central processing unit CPU, memory, machine bandwidth and machine disk; the computing power performance score and resource utilization of each candidate service instance are calculated based on the computing power resources, including: for each candidate service instance, determining the weight, total amount and real-time usage of each computing power resource; calculating the first computing power performance score of each computing power resource based on the weight, the total amount and the real-time usage, and taking the sum of the first computing power performance scores as the computing power performance score of the candidate service instance; calculating the first resource utilization of each computing power resource based on the total amount and the real-time usage; and determining the resource utilization of the candidate service instance based on the first resource utilization.
[0015] Furthermore, forwarding the service request to the target service instance based on the optimal path includes: encapsulating the optimal path into an extended header of an Internet Protocol version 6 IPv6 data packet to obtain a segment routing IPv6 message header; adding the segment routing IPv6 message header to the data packet of the service request to obtain a target data packet; and forwarding the target data packet to the target service instance.
[0016] Furthermore, before selecting at least one candidate service instance from the server cluster for processing the business request based on the business request to be scheduled, it also includes: selecting the business request to be scheduled from multiple first business requests according to preset scheduling rules, and the preset scheduling rules include first-in-first-out, computing power demand priority or network demand priority.
[0017] In a second aspect, the present application provides a service request scheduling device, which includes:
[0018] A service initial selection module, configured to select at least one candidate service instance for processing a service request to be scheduled from a server cluster based on the service request to be scheduled;
[0019] A data acquisition module, configured to acquire computing resources and network resources of each candidate service instance;
[0020] A service determination module, configured to select a target service instance from the candidate service instances based on the computing resources and the network resources;
[0021] A path calculation module, configured to calculate an optimal path from a current service node to a target service instance based on the network resources;
[0022] A request forwarding module is used to forward the business request to the target service instance based on the optimal path.
[0023] In a third aspect, the present application provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the service request scheduling method described in any embodiment of the present application.
[0024] In a fourth aspect, the present application provides a computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a processor to implement the service request scheduling method described in any embodiment of the present application when executed.
[0025] In order to solve the defects of the prior art in the background technology, an embodiment of the present application provides a service request scheduling method, and the execution of this method can bring the following beneficial effects: after receiving a service request, the present application determines the candidate service instance corresponding to the service request according to the type of the service request, comprehensively considers the computing power resources and network resources of each candidate service instance, determines the target service instance through heuristic calculation, and forwards the service request to the target service instance based on the forwarding rule. When scheduling service requests, the present application considers the computing power factors and network factors of traffic processing at the same time, so that the processing delay and transmission delay of the service request reach a certain balance, thereby selecting a target service instance that has both sufficient computing power and good network connection. This balanced resource utilization method avoids the resource waste or overload problems caused by single-factor scheduling. Since the present invention fully considers the dual constraints of computing and network when scheduling, it can minimize response time and processing delay while meeting business needs, thereby significantly improving service quality and user experience.
[0026] It should be noted that the above-mentioned computer instructions may be stored in whole or in part on a computer-readable storage medium. The computer-readable storage medium may be packaged together with the processor of the service request scheduling device or may be packaged separately from the processor of the service request scheduling device, and this application does not limit this.
[0027] The description of the second, third and fourth aspects in this application can refer to the detailed description of the first aspect; and the beneficial effects of the description of the second, third and fourth aspects can refer to the analysis of the beneficial effects of the first aspect, which will not be repeated here.
[0028] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description.
[0029] It is understandable that before using the technical solutions disclosed in the embodiments of this application, the type, scope of use, and usage scenarios of the personal information involved in this application should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0031] Figure 1A flowchart of a service request scheduling method provided in an embodiment of the present application;
[0032] Figure 2 A flowchart of a process for determining a target service instance provided in an embodiment of the present application;
[0033] Figure 3 A schematic diagram of the structure of a service request scheduling device provided in an embodiment of the present application;
[0034] Figure 4 This is a block diagram of an electronic device used to implement a service request scheduling method according to an embodiment of the present application. DETAILED DESCRIPTION
[0035] To make the purpose, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0036] It should be noted that the terms "first," "second," "target," and "original" in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or apparatus.
[0037] Figure 1 This is a flow chart of a service request scheduling method provided in an embodiment of the present application. This embodiment is applicable to scenarios where diverse service requests with varying computational complexity and network requirements are matched to appropriate microservice instances for processing. The service request scheduling method provided in this embodiment can be executed by a service request scheduling device provided in an embodiment of the present application. This device can be implemented in software and / or hardware and integrated into an electronic device that executes this method.
[0038] See also Figure 1 The method of this embodiment includes but is not limited to the following steps:
[0039] S110 . Select at least one candidate service instance for processing the service request from a server cluster based on the service request to be scheduled.
[0040] The type of business request depends on the type of business system. When the business system is a financial system, business requests include but are not limited to account inquiries, transaction processing, credit approval, and payment settlement. A server cluster is a logical collection of machines (such as physical machines, virtual machines, or containers) that work together and are interconnected through a network to provide unified service capabilities to the outside world. The server cluster is the operating carrier of the service instance. A service instance is a specific operating instance of a microservice and is the actual operating unit of a microservice on computing resources (such as machines or containers). A service instance is also a microservice instance.
[0041] Specifically, based on the business request to be scheduled, at least one candidate service instance is selected from the server cluster for processing the business request, including: determining the service type corresponding to processing the business request based on the identification code of the business request; and selecting at least one service instance corresponding to the service type from the server cluster as a candidate service instance.
[0042] In an embodiment of the present application, different business requests are usually processed by different modules. Therefore, it is necessary to use a routing rule engine or a service registration center to determine which type of service instance the business request needs to be processed by, that is, to determine the required service type based on the unique identification code of the business request (such as the domain name of the Domain Name System (DNS), or other preset identifier (ID) information or number, etc.).
[0043] When a service instance starts, it registers its information (health status, load metrics, etc.) with the registry and periodically renews its heartbeat. The scheduling system obtains a list of available instances for this service type from the registry in real time. After eliminating faulty or overloaded instances from the list, it selects at least one server instance as a candidate service instance.
[0044] Furthermore, before selecting at least one candidate service instance for processing the service request from the server cluster based on the service request to be scheduled, the method further includes: selecting the service request to be scheduled from the plurality of first service requests according to a preset scheduling rule;
[0045] During the selection phase of service traffic to be scheduled, considering that there are multiple service traffic flows (i.e., first service requests) to be scheduled simultaneously, these service traffic flows can be prioritized using preset scheduling rules to select the service traffic with the highest priority as the service request to be scheduled. Preset scheduling rules include first-in-first-out, computing power priority, or network demand priority.
[0046] S120: Obtain computing resources and network resources of each candidate service instance.
[0047] Among them, computing power resources can include dynamic resource information and static resource information. Dynamic resource information is used to identify the current usage of resources, which refers to the actual real-time usage of various types of computing power resources. Static resource information is used to indicate the basic performance of resources, which refers to the total amount or allocated amount of various types of computing power resources.
[0048] Specifically, computing resource information is shown in Table 1. Dynamic resource information includes at least the real-time usage of the central processing unit (CPU), memory, machine bandwidth, and machine disk resources. Static resource information includes at least the total and allocated CPU, memory, machine bandwidth, and machine disk resources. In this embodiment, only CPU, memory, machine disk, and machine bandwidth are used for illustration. Furthermore, heterogeneous computing resource information can be included, such as computing resource information related to the graphics processing unit (GPU).
[0049] Table 1 Computing power resource information
[0050]
[0051] Table 2 shows network resource information, which includes at least latency, bandwidth, packet loss rate, and connectivity. Latency refers to the time required for data to travel from the source Internet Protocol (IP) to the destination IP. Low latency is crucial for real-time applications (such as online games and real-time trading systems). Bandwidth refers to the maximum amount of data a network can transmit per unit time. High bandwidth means faster data transmission, which is crucial for large data volumes and high-speed computing tasks. Packet loss rate refers to the rate at which data packets are lost during network transmission. A high packet loss rate can result in incomplete data, impacting application performance and reliability. Connectivity refers to the connection status between two clusters, describing whether a network connection can be established between the two clusters.
[0052] Table 2 Network resource information
[0053]
[0054] In the embodiments of the present application, static resource information of computing resources can be pre-stored in a database and can be directly queried from the database when needed. For dynamic information of computing resources and network resource information, it is necessary to use a real-time perception system to obtain resource information, such as through monitoring systems such as Prometheus, or through in-band network telemetry technology, data packet Border Gateway Protocol (BGP) resource announcements, and other technical means to obtain it.
[0055] S130: Select a target service instance from the candidate service instances based on computing resources and network resources.
[0056] Specifically, a target service instance is selected from candidate service instances based on computing power resources and network resources, including: calculating the computing power performance score and resource utilization of each candidate service instance based on computing power resources; calculating the load balancing score of each candidate service instance based on resource utilization; calculating the score adjustment coefficient of each candidate service instance based on network resources; calculating the score of each candidate service instance based on the computing power performance score, load balancing score and score adjustment coefficient; and selecting the candidate service instance whose score meets the preset criteria as the target service instance.
[0057] In an embodiment of the present application, in order to more efficiently utilize the resources in the entire business system, meet user needs, and improve user experience, this embodiment proposes a resource scoring method that comprehensively considers computing resources and network resources. Computing resource information is mainly reflected in the total amount and idle rate of the computing resources of the service instance. The greater the total amount, the higher the score of the service instance; the higher the idle rate, the higher the score of the service instance. Network resource information is mainly reflected in the network latency from the user to the service instance. The lower the latency, the higher the score of the service instance.
[0058] like Figure 2 The figure shows a flowchart for determining the target service instance, calculating the computing power performance score and load balancing score of each candidate service instance, calculating the score (i.e., comprehensive score) of each candidate service instance based on the computing power performance score and load balancing score, sorting the scores of all candidate service instances, and selecting the candidate service instances whose scores meet the preset standards as the target service instances.
[0059] Furthermore, the types of computing power resources include at least CPU, memory, machine bandwidth and machine disk; the computing power performance score and resource utilization of each candidate service instance are calculated based on the computing power resources, including: for each candidate service instance, determining the weight, total amount and real-time usage of each computing power resource; calculating the first computing power performance score of each computing power resource based on the weight, total amount and real-time usage, and taking the sum of the first computing power performance scores as the computing power performance score of the candidate service instance; calculating the first resource utilization of each computing power resource based on the total amount and real-time usage; and determining the resource utilization of the candidate service instance based on the first resource utilization.
[0060] Specifically, the computing power performance score of the candidate service instance (denoted as S1) is calculated using the following formula (1):
[0061]
[0062] Where n represents the number of types of computing resources. This embodiment takes four types of computing resources as an example, such as CPU, memory, machine bandwidth, and machine disk. i represents the index number of the computing resource type. i Represents the weight of the i-th type of computing power resources. In this embodiment, it can be assumed that the weights of various types of computing power resources are the same, such as one-quarter. The weights can be adjusted according to actual conditions. represents the total amount of the i-th resource, Indicates the real-time usage of the i-th resource.
[0063] The real-time usage of each computing resource is divided by the total amount to represent the first resource utilization rate, and the first resource utilization rates are combined to represent the resource utilization rate of the candidate service instance (denoted as U R ), and is calculated using the following formula (2):
[0064]
[0065] Where, Indicates the real-time usage of the first resource. Indicates the total amount of the first resource, Indicates the real-time usage of the second resource. Indicates the total amount of the second resource, Indicates the real-time usage of the third resource. Indicates the total amount of the third resource, Indicates the real-time usage of the fourth resource. Indicates the total amount of the fourth resource.
[0066] The load balancing score of the candidate service instance (denoted as S2) is measured by the standard deviation of the resource utilization of each type of computing resources and is calculated using the following formula (3):
[0067] S2=[1-σ(U R )] (3)
[0068] Taking network resources into consideration, the longer the delay, the lower the overall score, and the shorter the delay, the higher the overall score. Therefore, it can be used as a score adjustment coefficient (denoted as k) and calculated using the following formula (4):
[0069]
[0070] Among them, τ represents the delay information, b τ is a constant that can take any reasonable value.
[0071] The computing power performance score, load balancing score and score adjustment coefficient of each candidate service instance are calculated to obtain the score of each candidate service instance (denoted as score), which is calculated using the following formula (5):
[0072]
[0073] Finally, the candidate service instances are sorted according to the scores, and the candidate service instances whose scores meet the preset standards are selected as the target service instances.
[0074] Furthermore, before calculating the computing power performance score and resource utilization of each candidate service instance based on the computing power resources, service instances that do not meet the requirements (i.e., service instance candidates) are filtered out based on the computing power resources and network resources of the candidate service instances to reduce the calculation scale during scoring.
[0075] The specific process of service instance selection is: if the number of candidate service instances is less than the preset number, determine whether the computing power resources of each candidate service instance meet the preset computing power requirements and whether the network resources meet the preset network requirements; use the candidate service instances that meet the preset computing power requirements and the preset network requirements as candidate service instances; accordingly, calculate the computing power performance score and resource utilization of each candidate service instance based on the computing power resources, including: calculating the computing power performance score and resource utilization of each candidate service instance based on the computing power resources.
[0076] If there are too many candidate service instances, service instance selection can be skipped. Specifically, when a certain percentage of candidate service instances meet the requirements, the service instance selection phase can be skipped, thereby speeding up scheduling. After the service instance selection process, a candidate list is generated. Each service instance in the candidate list can meet the computing and network resource requirements of the business traffic, but further evaluation of these candidate service instances is required to determine the optimal target service instance.
[0077] S140: Calculate the optimal path from the current service node to the target service instance based on network resources.
[0078] In the embodiments of the present application, resource information such as network topology and node load, as well as information such as segment identifiers related to Segment Routing over IPv6 (SRv6), is first collected. This information is then combined with SRv6 path calculation algorithms, such as those based on shortest path first or traffic engineering, to ultimately determine an optimal path from the current node to the target microservice instance. This path meets specific service requirements, such as minimal latency and sufficient bandwidth.
[0079] S150: Forward the service request to the target service instance based on the optimal path.
[0080] Specifically, forwarding the service request to the target service instance based on the optimal path includes: encapsulating the optimal path into an extended header of an Internet Protocol Version 6 (IPv6) data packet to obtain an SRv6 message header; adding the SRv6 message header to the data packet of the service request to obtain a target data packet; and forwarding the target data packet to the target service instance.
[0081] In an embodiment of the present application, the calculated optimal path information is first encapsulated into a special header (i.e., an SRv6 message header) according to the rules of SRv6. This header contains the key identification information of the path. This header is then added to the front of the service request data packet to be sent to form a complete target data packet. Finally, this data packet with path information is sent to the target service instance, and the data packet is forwarded according to the encapsulated path. Path-based forwarding is implemented to ensure that banking business traffic can be efficiently and accurately transmitted to the target microservice instance, ensuring service quality and improving user experience.
[0082] Furthermore, this embodiment includes a timed forwarding policy correction rule. Specifically, it captures server and global network resource usage in real time, calculates a relevant score in real time according to the calculation rules selected for the target service instance, compares this score with the data in the configuration table, and performs real-time corrections on the forwarding information table to ensure the global optimization of the forwarding rules. This rule ensures that the forwarding policy rules are globally optimized in real time, and maintains service quality even when server and network resources fluctuate.
[0083] The technical solution provided in this embodiment selects at least one candidate service instance for processing the business request from the server cluster based on the business request to be scheduled; obtains the computing power resources and network resources of each candidate service instance; selects the target service instance from the candidate service instances based on the computing power resources and network resources; calculates the optimal path from the current service node to the target service instance based on the network resources; and forwards the business request to the target service instance based on the optimal path. After receiving the business request, the present application determines the candidate service instance corresponding to the business request according to the type of business request, comprehensively considers the computing power resources and network resources of each candidate service instance, determines the target service instance through heuristic calculation, and forwards the business request to the target service instance based on the forwarding rules. When scheduling business requests, the present application considers the computing power factors and network factors of traffic processing at the same time, so that the processing delay and transmission delay of the business request reach a certain balance, thereby selecting the target service instance that has both sufficient computing power and good network connection. This balanced resource utilization method avoids the resource waste or overload problems caused by single-factor scheduling. Since the present invention fully considers the dual constraints of computing and network during scheduling, it can minimize response time and processing delay while meeting business needs, thereby significantly improving service quality and user experience.
[0084] Figure 3 A schematic diagram of the structure of a service request scheduling device provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the apparatus 300 may include:
[0085] A service initial selection module 310 is configured to select at least one candidate service instance for processing a service request from a server cluster based on the service request to be scheduled;
[0086] A data acquisition module 320 is configured to acquire computing resources and network resources of each candidate service instance;
[0087] A service determination module 330 is configured to select a target service instance from the candidate service instances based on the computing resources and the network resources;
[0088] A path calculation module 340 is configured to calculate an optimal path from a current service node to a target service instance based on the network resources;
[0089] The request forwarding module 350 is configured to forward the service request to the target service instance based on the optimal path.
[0090] Furthermore, the service preliminary selection module 310 may be specifically configured to: determine a service type corresponding to the service request based on the identification code of the service request; and select at least one service instance corresponding to the service type from the server cluster as a candidate service instance.
[0091] Furthermore, the above-mentioned service determination module 330 can be specifically used to: calculate the computing power performance score and resource utilization of each candidate service instance based on the computing power resources; calculate the load balancing score of each candidate service instance based on the resource utilization; calculate the score adjustment coefficient of each candidate service instance based on the network resources; calculate the score of each candidate service instance based on the computing power performance score, the load balancing score and the score adjustment coefficient; and select the candidate service instance whose score meets the preset criteria as the target service instance.
[0092] Furthermore, the above-mentioned service request scheduling device may further include: a service alternative module;
[0093] The service candidate module is configured to, before respectively calculating the computing power performance score and resource utilization rate of each candidate service instance based on the computing power resources, determine whether the computing power resources of each candidate service instance meet the preset computing power requirements and whether the network resources meet the preset network requirements if the number of candidate service instances is less than a preset number; and select the candidate service instance that meets the preset computing power requirements and the preset network requirements as a candidate service instance;
[0094] Accordingly, the service determination module 330 may be specifically configured to calculate the computing power performance score and resource utilization rate of each candidate service instance based on the computing power resources.
[0095] In one embodiment, the types of computing resources include at least a central processing unit (CPU), memory, machine bandwidth, and machine disk;
[0096] Furthermore, the above-mentioned service determination module 330 can be specifically used to: determine the weight, total amount and real-time usage of each computing power resource for each candidate service instance; calculate the first computing power performance score of each computing power resource based on the weight, total amount and real-time usage, and use the sum of the first computing power performance scores as the computing power performance score of the candidate service instance; calculate the first resource utilization of each computing power resource based on the total amount and real-time usage; and determine the resource utilization of the candidate service instance based on the first resource utilization.
[0097] Furthermore, the above-mentioned request forwarding module 350 can be specifically used to: encapsulate the optimal path into the extended header of the Internet Protocol version 6 IPv6 data packet to obtain a segment routing IPv6 message header; add the segment routing IPv6 message header to the data packet of the service request to obtain a target data packet; and forward the target data packet to the target service instance.
[0098] Furthermore, the service request scheduling device may further include: a request acquisition module;
[0099] The request acquisition module is used to select the business request to be scheduled from multiple first business requests according to preset scheduling rules before selecting at least one candidate service instance for processing the business request based on the business request to be scheduled from the server cluster. The preset scheduling rules include first-in-first-out, computing power demand priority or network demand priority.
[0100] The service request scheduling device provided in this embodiment can be applied to the service request scheduling method provided in any of the above embodiments, and has corresponding functions and beneficial effects.
[0101] Figure 4 It is a block diagram of an electronic device for implementing a business request scheduling method of an embodiment of the present application. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0102] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0103] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0104] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the service request scheduling method.
[0105] In some embodiments, the service request scheduling method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the service request scheduling method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the service request scheduling method in any other suitable manner (e.g., via firmware).
[0106] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0107] Computer programs for implementing the methods of the present application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0108] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0109] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0110] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0111] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0112] Note that the above are only preferred embodiments of the present application and the technical principles used. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. For example, those skilled in the art can use the various forms of processes shown above, reorder, add, or delete steps; and can perform the steps described in the present application in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present application can be achieved, and this document does not limit them here.
[0113] The above specific embodiments do not limit the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A service request scheduling method, characterized in that: The method comprises: Selecting at least one candidate service instance for processing the service request from the server cluster based on the service request to be scheduled; Obtain computing resources and network resources for each candidate service instance; Selecting a target service instance from the candidate service instances based on the computing resources and the network resources; Calculating an optimal path from the current service node to the target service instance based on the network resources; The service request is forwarded to the target service instance based on the optimal path.
2. The service request scheduling method according to claim 1, wherein: The selecting, based on the service request to be scheduled, at least one candidate service instance for processing the service request from the server cluster includes: Determining a service type corresponding to processing the service request based on the identification code of the service request; At least one service instance corresponding to the service type is selected from the server cluster as a candidate service instance.
3. The service request scheduling method according to claim 1, wherein: The selecting a target service instance from the candidate service instances based on the computing resources and the network resources includes: Calculate the computing power performance score and resource utilization of each candidate service instance based on the computing power resources; Calculating a load balancing score for each candidate service instance based on the resource utilization; Calculating a score adjustment coefficient for each candidate service instance based on the network resources; Calculating a score for each candidate service instance based on the computing power performance score, the load balancing score, and the score adjustment coefficient; The candidate service instance whose score meets a preset standard is selected as the target service instance.
4. The service request scheduling method according to claim 3, wherein: Before respectively calculating the computing power performance score and resource utilization rate of each candidate service instance based on the computing power resources, the method further includes: If the number of the candidate service instances is less than the preset number, determining whether the computing power resources of each candidate service instance meet the preset computing power requirements and whether the network resources meet the preset network requirements; The candidate service instance that meets the preset computing power requirement and the preset network requirement is used as a candidate service instance; Accordingly, the computing power performance score and resource utilization of each candidate service instance are calculated based on the computing power resources, including: The computing power performance score and resource utilization rate of each candidate service instance are calculated based on the computing power resources.
5. The service request scheduling method according to claim 3, characterized in that: The types of computing resources include at least a central processing unit (CPU), memory, machine bandwidth, and machine disk; and calculating the computing performance score and resource utilization of each candidate service instance based on the computing resources includes: For each candidate service instance, determine the weight, total amount, and real-time usage of each computing resource; Calculate a first computing power performance score for each computing power resource based on the weight, the total amount, and the real-time usage, and use the sum of the first computing power performance scores as the computing power performance score of the candidate service instance; Calculate a first resource utilization rate of each computing resource based on the total amount and the real-time usage; The resource utilization of the candidate service instance is determined based on the first resource utilization.
6. The service request scheduling method according to claim 1, wherein: The forwarding the service request to the target service instance based on the optimal path includes: Encapsulating the optimal path into an extended header of an Internet Protocol version 6 (IPv6) data packet to obtain a segment routing IPv6 message header; Adding the segment routing IPv6 message header to the data packet of the service request to obtain a target data packet; Forward the target data packet to the target service instance.
7. The service request scheduling method according to claim 1, wherein: Before selecting at least one candidate service instance for processing the service request from the server cluster based on the service request to be scheduled, the method further includes: The service request to be scheduled is selected from multiple first service requests according to a preset scheduling rule, where the preset scheduling rule includes first-in-first-out, computing power demand priority, or network demand priority.
8. A service request scheduling device, characterized in that: The device comprises: A service initial selection module, configured to select at least one candidate service instance from a server cluster for processing a service request to be scheduled based on the service request to be scheduled; A data acquisition module, configured to acquire computing resources and network resources of each candidate service instance; A service determination module, configured to select a target service instance from the candidate service instances based on the computing resources and the network resources; A path calculation module, configured to calculate an optimal path from a current service node to a target service instance based on the network resources; A request forwarding module is used to forward the business request to the target service instance based on the optimal path.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the service request scheduling method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the service request scheduling method according to any one of claims 1 to 7 when executed.
Citation Information
Cited By
Service instance scheduling method, intelligent gateway platform, equipment and storage medium
CN121098935A
Service node discovery method, device and equipment under edge computing power network and medium
CN121262270A
Request processing method and device based on large model, electronic equipment and storage medium
CN121328705A
Service path scheduling method and device, equipment, storage medium and product
CN121691461A