Service scheduling method, device, equipment and medium

By introducing the first load threshold and the second load threshold in the server, dynamically scheduling service requests is solved, and the problem of low hardware resource utilization in the prior art is achieved, and higher resource utilization and lower service crash risk is achieved.

CN120066717APending Publication Date: 2025-05-30太保科技有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510129412.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, in order to ensure server stability, a relatively conservative maximum load threshold is usually set, resulting in a low hardware resource utilization rate.

Method used

By introducing a first load threshold and a second load threshold, the service request is dynamically scheduled. When the hardware load information is lower than the first load threshold, the service request is satisfied; when the hardware load information is between the first and second load thresholds, the service request is added to the waiting queue; when the load information is higher than the second load threshold, the service is degraded or denied.

Benefits of technology

On the basis of ensuring the stability of the server, it makes full use of hardware resources, improves resource utilization, and reduces the risk of service crashes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066717A_ABST
    Figure CN120066717A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a service scheduling method and device, equipment and a medium, and relates to the technical field of data processing. By introducing the first load threshold value and the second load threshold value, hardware resources of the server are allowed to be more fully utilized on the basis of ensuring the stability of the server. Specifically, when the hardware load information of the server is lower than a first load threshold value, the service request is directly satisfied; when the hardware load information is between the first load threshold value and the second load threshold value, the service request is added into the waiting queue instead of being rejected immediately, so that idle resources can be utilized more effectively, and the resource utilization rate is improved; and when the hardware load information is higher than the second load threshold value, the service request can be subjected to degradation processing or rejection processing, so that the service quality can be adjusted according to the actual condition of the server, the risk of service crash is reduced, and the resource utilization rate is also improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a service scheduling method, apparatus, device, and medium. Background Art

[0002] With the continuous progress of artificial intelligence (AI) and deep learning algorithms, AI models have demonstrated excellent capabilities in tasks such as generating images and answering questions. These AI models are widely used in multiple fields such as customer service and content creation, significantly improving service efficiency and enhancing the user experience.

[0003] In the related art, when processing a large number of user service requests, in order to maintain service stability and response speed, the common practice is to evenly distribute all service requests triggered by users to multiple servers for processing. Prior to this, relevant technical personnel need to determine the maximum load threshold that the server can theoretically withstand in a stress test environment. Once the real-time load of the server approaches or reaches the maximum load threshold, the server will directly reject new service requests to prevent overload.

[0004] However, in actual operation, in order to ensure the absolute stability of the server, relevant technical personnel often set a relatively conservative maximum load threshold. Although this approach effectively reduces the risk of server overload, it also means that the hardware resources of the server may not be utilized most effectively, resulting in a problem of low resource utilization rate. Summary of the Invention

[0005] Based on the above problems, this application provides a service scheduling method, apparatus, device, and medium, which can improve the resource utilization rate of the server.

[0006] The embodiments of this application disclose the following technical solutions:

[0007] In a first aspect, this application discloses a service scheduling method, and the method includes:

[0008] When a service request is received at a first moment, determine the hardware load information of the server within a first time window, where the first time window is the time window in which the first moment is located when the time axis is divided into multiple non-overlapping time windows;

[0009] When the hardware load information within the first time window is lower than a first load threshold, satisfy the service request;

[0010] When the hardware load information within the first time window is higher than the first load threshold and lower than a second load threshold, add the service request to a waiting queue;

[0011] When the hardware load information within the first time window is higher than the second load threshold, perform degradation service on the service request, or reject the service request.

[0012] Optionally, the method further includes:

[0013] When the real-time hardware load information of the server is lower than the first load threshold, satisfy the service requests in the waiting queue until the real-time hardware load information reaches the second load threshold.

[0014] Optionally, satisfying the service requests in the waiting queue includes:

[0015] Satisfy the service requests in the waiting queue according to the service priority or reception time of the service requests in the waiting queue, where the reception time is the time when the service request is added to the waiting queue.

[0016] Optionally, the hardware load information of the server within the first time window is obtained periodically, and the acquisition frequency of the periodic acquisition is positively correlated with the hardware load information of the server within the historical time window.

[0017] Optionally, the length of the time window is negatively correlated with the hardware load information of the server within the historical time window.

[0018] In a second aspect, the present application discloses a service scheduling device, which includes: an information determination module, a first scheduling module, a second scheduling module, and a third scheduling module;

[0019] The information determination module is configured to determine the hardware load information of the server within the first time window when a service request is received at the first moment, where the first time window is the time window in which the first moment is located in the case of dividing the time axis into multiple non-overlapping time windows;

[0020] The first scheduling module is configured to satisfy the service request when the hardware load information within the first time window is lower than the first load threshold;

[0021] The second scheduling module is configured to add the service request to the waiting queue when the hardware load information within the first time window is higher than the first load threshold and lower than the second load threshold;

[0022] The third scheduling module is configured to perform degradation service on the service request, or reject the service request, when the hardware load information within the first time window is higher than the second load threshold.

[0023] Optionally, the device further includes: a fourth scheduling module;

[0024] The fourth scheduling module is specifically used to: when the real-time hardware load information of the server is lower than the first load threshold, satisfy the service requests in the waiting queue until the real-time hardware load information reaches the second load threshold.

[0025] Optionally, the fourth scheduling module is specifically used to: meet the service requests in the waiting queue according to the service priority or receiving time of the service requests in the waiting queue, wherein the receiving time is the time when the service request is added to the waiting queue.

[0026] Optionally, the hardware load information of the server in the first time window is acquired periodically, and an acquisition frequency of the periodic acquisition is positively correlated with the hardware load information of the server in the historical time window.

[0027] Optionally, the length of the time window is negatively correlated with hardware load information of the server within the historical time window.

[0028] In a third aspect, the present application discloses a service scheduling device, the device comprising: a memory and a processor;

[0029] The memory is used to store programs;

[0030] The processor is used to execute the program to implement the various steps of the service scheduling method as described in the first aspect.

[0031] In a fourth aspect, the present application discloses a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, the various steps of the service scheduling method described in the first aspect are implemented.

[0032] Compared with the prior art, this application has the following beneficial effects:

[0033] The embodiments of the present application provide a service scheduling method, device, equipment and medium. The present application allows the server's hardware resources to be more fully utilized on the basis of ensuring the stability of the server by introducing a first load threshold and a second load threshold. Specifically, when the server's hardware load information is lower than the first load threshold, the service request will be directly satisfied; when the hardware load information is between the first load threshold and the second load threshold, the service request will be added to the waiting queue instead of being rejected immediately, so that idle resources can be used more effectively and resource utilization can be improved; when the hardware load information is higher than the second load threshold, the service request can be downgraded or rejected, so that the service quality can be adjusted according to the actual situation of the server, reducing the risk of service crash and also improving resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0035] Figure 1 It is a flowchart of a service scheduling method provided by an embodiment of the present application;

[0036] Figure 2 It is a topology diagram of a server provided by an embodiment of the present application;

[0037] Figure 3 It is a schematic diagram of a service scheduling device provided by an embodiment of the present application;

[0038] Figure 4 It is a schematic diagram of a computer-readable medium provided by an embodiment of the present application. Detailed implementation manners

[0039] As described above, when processing a large number of user service requests, in order to maintain the stability and response speed of the service, the usual practice is to evenly distribute all the service requests triggered by users to multiple servers for processing. Before this, relevant technical personnel need to determine the maximum load threshold that the server can theoretically withstand in a stress test environment. Once the real-time load of the server approaches or reaches the maximum load threshold, the server will directly reject new service requests to prevent overload.

[0040] However, in actual operation, in order to ensure the absolute stability of the server, relevant technical personnel often set a relatively conservative maximum load threshold. Although this approach effectively reduces the risk of server overload, it also means that the hardware resources of the server may not be utilized most effectively, resulting in the problem of low resource utilization rate.

[0041] After research, the inventor proposed a service scheduling method, device, equipment and medium, the method comprising: when a service request is received at a first moment, determining the hardware load information of the server within a first time window, wherein the first time window is the time window at the first moment when the time axis is divided into multiple non-overlapping time windows; when the hardware load information within the first time window is lower than the first load threshold, satisfying the service request; when the hardware load information within the first time window is higher than the first load threshold and lower than the second load threshold, adding the service request to a waiting queue; when the hardware load information within the first time window is higher than the second load threshold, downgrading the service request, or rejecting the service request. Therefore, by introducing the first load threshold and the second load threshold, the present application allows for more effective use of the server's hardware resources while ensuring the stability of the server. Specifically, when the hardware load information of the server is lower than the first load threshold, the service request will be directly satisfied; when the hardware load information is between the first load threshold and the second load threshold, the service request will be added to the waiting queue instead of being rejected immediately, thereby making more efficient use of idle resources and improving resource utilization; when the hardware load information is higher than the second load threshold, the service request can be downgraded or rejected, thereby adjusting the service quality according to the actual situation of the server, reducing the risk of service crash and also improving resource utilization.

[0042] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0043] See also Figure 1 , which is a flow chart of a service scheduling method provided by an embodiment of the present application. Figure 2 , which is a topological diagram of a server provided in an embodiment of the present application. The method includes:

[0044] S101: When a service request is received at a first moment, a first time window in which the first moment is located is determined.

[0045] The time axis refers to a continuous time series. The embodiment of the present application divides the continuous time axis into multiple non-overlapping, discrete time windows, and each time window represents a time period of a fixed length, such as 1 second, 100 milliseconds or other set duration. The first time window refers to the time window at the first moment (i.e., the moment when the service request is received).

[0046] It is understandable that by dividing the time axis into time windows, the number of service requests within the time windows can be detected more precisely, which helps to achieve smooth traffic, prevent sudden traffic surges, and optimize resource allocation.

[0047] S102: Determine the hardware load information of the server within the first time window.

[0048] The hardware load information includes the occupancy rate of the Central Processing Unit (CPU), the occupancy rate of the Graphics Processing Unit (GPU), the memory usage rate, and the network bandwidth usage, etc. It should be noted that this application does not limit the specific hardware load information.

[0049] In some specific implementation manners, the length of the time window is not fixed, but can be dynamically adjusted according to the historical load situation of the server. Specifically, the length of the time window is negatively correlated with the hardware load information of the server within the historical time window. This means that if the server has experienced a high hardware load in the past period, the time window will be shortened to increase the inspection frequency, thereby avoiding excessive pressure on the server caused by a sudden increase in traffic. If the server has experienced a low hardware load in the past period, the time window will be extended to reduce the inspection frequency, save computing resources, and at the same time allow more service requests to be processed within the same time period, improving the throughput.

[0050] In some other specific implementation manners, the server obtains the hardware load information periodically, and the acquisition frequency of the hardware load information can also be dynamically adjusted according to the historical load situation of the server. Specifically, the acquisition frequency of the hardware load information is positively correlated with the historical load information of the server. This means that if the server has experienced a high hardware load in the past period, the acquisition frequency of the hardware load information will increase. By monitoring the hardware load information of the server more frequently, any changes that may affect the service stability can be identified faster, and corresponding measures can be taken earlier, thereby improving the response speed and reliability of the server. If the server has experienced a low hardware load in the past period, the acquisition frequency of the hardware load information will decrease. In this case, the risk of server overload is small, so there is no need to check too frequently. This can not only save computing resources, but also reduce the additional burden on the server itself, indirectly improving the overall efficiency.

[0051] S103: When the hardware load information within the first time window is lower than the first load threshold, the service request is satisfied.

[0052] When the hardware load information within the first time window is lower than the first load threshold, it means that the server is currently in a relatively idle state and has sufficient computing power and resources to process new service requests. At this time, the server can directly receive and process these service requests without taking additional traffic limiting measures. It should be noted that the first load threshold can be 50%, and the present application does not limit the specific first load threshold.

[0053] Thus, not only can the waiting time of users be reduced and the user experience be improved, but also the existing hardware resources can be fully utilized without sacrificing service quality.

[0054] S104: When the hardware load information within the first time window is higher than the first load threshold and lower than the second load threshold, add the service request to the waiting queue.

[0055] When the hardware load information within the first time window is between the first load threshold and the second load threshold, it means that the server has exceeded the optimal working state but has not reached the overload state. At this time, newly arrived service requests will not be processed immediately but will be added to the waiting queue for subsequent processing. The waiting queue is a caching mechanism used to temporarily store service requests that cannot be processed immediately. Once the real-time hardware load information of the server drops below the first load threshold, the service requests in the waiting queue will be gradually processed. It should be noted that the second load threshold can be 75%, and the present application does not limit the specific second load threshold.

[0056] Thus, by introducing the waiting queue, the server is prevented from overloading due to a sudden increase in traffic, ensuring the stability of the server. Moreover, even when the hardware load is high, the server can still receive and save service requests, providing a certain degree of service guarantee for users.

[0057] S105: When the hardware load information within the first time window is higher than the second load threshold, perform degraded service on the service request, or reject the service request.

[0058] When the hardware load information within the first time window is higher than the second load threshold, it means that the server has entered a very high hardware load state. Continuing to receive new service requests may lead to a significant decline in service quality or even the crash of the server. Therefore, the server needs to take more stringent measures, such as performing degraded processing on some service requests or directly rejecting these service requests. Among them, degraded service refers to providing a simplified service result, such as reducing the image generation quality, reducing the detail level of text answers, etc., so as to reduce the burden on the server. In extreme cases, the server will directly reject some newly incoming service requests to protect the core functions and the overall stability of the service.

[0059] Thus, this application can preferentially ensure that critical services are not affected and maintain the most basic service availability. Moreover, this application can avoid server overload and prevent possible service interruptions. Although some service requests may be downgraded or rejected, for most users, the server can still provide a certain degree of service, reducing the risk of a full-scale failure.

[0060] S106: When the real-time hardware load information of the server is lower than the first load threshold, satisfy the service requests in the waiting queue until the real-time hardware load information of the server reaches the second load threshold.

[0061] When the real-time hardware load information of the server is lower than the first load threshold, it means that the server currently has enough idle resources to process more service requests. At this time, service requests can be selected from the waiting queue for processing. When the real-time hardware load information of the server reaches the second load threshold, the server will pause extracting new service requests from the waiting queue to prevent overload.

[0062] In some specific implementation manners, the specific processing principles include sorting based on priority and sorting based on the receiving time. Among them, sorting based on priority means determining the processing order according to the service priority of each service request. For example, service requests with higher priorities will be processed first. If multiple service requests have the same priority, the processing order can be determined according to the time when the service requests are added to the waiting queue (i.e., the "receiving time"). For example, service requests that enter the queue earlier will be processed first.

[0063] It should be noted that the server can also reorder the service requests in the waiting queue by combining other factors (such as user type, historical interaction behavior, etc.) to optimize the overall service quality. The specific service order is not limited in this application.

[0064] It should also be noted that this application can also provide developers with detailed data support by collecting key performance indicators (KPIs) of the entire service operation cycle and hardware load conditions in a specific time period, helping them identify and solve bottleneck problems in service operation. Specifically, key operating indicators may include queries per second (QPS), concurrency, and latency. Among them, QPS refers to the number of requests processed per unit time. A high QPS indicates that the server has a strong processing capability; conversely, the server may indicate that there is a performance bottleneck or need to be optimized. Concurrency refers to the number of requests being processed by the server at the same time, reflecting the server's concurrent processing capability. Monitoring concurrency helps understand the server's performance in multi-tasking and identify whether there are blocking or delay problems caused by excessive concurrency. Latency refers to the time difference between the issuance of a service request and the return of a response, which directly affects the user experience. Low latency means faster response speed and better user experience; high latency may be caused by server overload. This application does not limit specific key operating indicators. In addition to the above static indicators, this application also places special emphasis on tracking and recording hardware load information within a specific time period. This is because traffic patterns in different time periods may vary significantly, for example, the performance during peak hours on weekdays and low hours at night are often very different. By collecting hardware load information by time segment, these changes can be captured more accurately and targeted optimization decisions can be made accordingly.

[0065] In summary, the embodiment of the present application provides a service scheduling method. By introducing a first load threshold and a second load threshold, the present application allows for more effective use of the server's hardware resources while ensuring the stability of the server. Specifically, when the server's hardware load information is lower than the first load threshold, the service request will be directly satisfied; when the hardware load information is between the first load threshold and the second load threshold, the service request will be added to the waiting queue instead of being rejected immediately, thereby making more efficient use of idle resources and improving resource utilization; when the hardware load information is higher than the second load threshold, the service request can be downgraded or rejected, thereby adjusting the service quality according to the actual situation of the server, reducing the risk of service crashes, and also improving resource utilization.

[0066] See also Figure 3 , which is a schematic diagram of a service scheduling device provided in an embodiment of the present application. The service scheduling device 300 includes: an information determination module 301, a first scheduling module 302, a second scheduling module 303 and a third scheduling module 304;

[0067] An information determination module 301, configured to determine the hardware load information of the server within a first time window when a service request is received at a first moment, where the first time window is the time window in which the first moment is located when the time axis is divided into multiple non-overlapping time windows;

[0068] A first scheduling module 302, configured to satisfy the service request when the hardware load information within the first time window is lower than a first load threshold;

[0069] A second scheduling module 303, configured to add the service request to a waiting queue when the hardware load information within the first time window is higher than the first load threshold and lower than a second load threshold;

[0070] A third scheduling module 304, configured to perform degraded service on the service request or reject the service request when the hardware load information within the first time window is higher than the second load threshold.

[0071] In some specific implementation manners, the service scheduling apparatus 300 further includes: a fourth scheduling module;

[0072] The fourth scheduling module is specifically configured to: when the real-time hardware load information of the server is lower than the first load threshold, satisfy the service requests in the waiting queue until the real-time hardware load information reaches the second load threshold.

[0073] In some specific implementation manners, the fourth scheduling module is specifically configured to: satisfy the service requests in the waiting queue according to the service priorities or receiving moments of the service requests in the waiting queue, where the receiving moment is the moment when the service request is added to the waiting queue.

[0074] In some specific implementation manners, the hardware load information of the server within the first time window is obtained periodically, and the periodic acquisition frequency is positively correlated with the hardware load information of the server within the historical time window.

[0075] In some specific implementation manners, the length of the time window is negatively correlated with the hardware load information of the server within the historical time window.

[0076] In summary, the embodiment of the present application provides a service scheduling device. By introducing a first load threshold and a second load threshold, the present application allows the hardware resources of the server to be more fully utilized while ensuring the stability of the server. Specifically, when the hardware load information of the server is lower than the first load threshold, the service request will be directly satisfied; when the hardware load information is between the first load threshold and the second load threshold, the service request will be added to the waiting queue instead of being rejected immediately, so that idle resources can be used more effectively and resource utilization can be improved; when the hardware load information is higher than the second load threshold, the service request can be downgraded or rejected, so that the service quality can be adjusted according to the actual situation of the server, reducing the risk of service crash and also improving resource utilization.

[0077] The embodiment of the present application also provides a corresponding service scheduling device and a computer-readable medium for implementing the service scheduling method provided in the embodiment of the present application.

[0078] Among them, the service scheduling device includes a memory and a processor, the memory is used to store instructions or codes, and the processor is used to execute instructions or codes so that the device executes a service scheduling method of any embodiment of the present application.

[0079] See also Figure 4 , which is a schematic diagram of a computer-readable medium provided in an embodiment of the present application. The computer-readable medium 400 stores a computer program 411, which implements the above-mentioned Figure 1 Steps of the service scheduling method.

[0080] It should be noted that in the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0081] It should be noted that the above-mentioned machine-readable medium in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example - but not limited to - an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0082] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; it can also exist separately and not be assembled into the electronic device.

[0083] Although the subject matter has been described in language specific to structural features and / or method logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms for implementing the claims.

[0084] Although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of this application. Certain features described in the context of separate embodiments can also be implemented combinatorially in a single embodiment. On the contrary, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0085] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present application that have similar functions.

Claims

1. A service scheduling method, characterized in that: The method comprises: When a service request is received at a first moment, determining hardware load information of the server within a first time window, wherein the first time window is a time window in which the first moment is located when the time axis is divided into a plurality of non-overlapping time windows; When the hardware load information within the first time window is lower than a first load threshold, satisfying the service request; When the hardware load information in the first time window is higher than the first load threshold and lower than the second load threshold, adding the service request to a waiting queue; When the hardware load information in the first time window is higher than the second load threshold, the service request is downgraded, or the service request is rejected.

2. The method according to claim 1, characterized in that: The method further comprises: When the real-time hardware load information of the server is lower than the first load threshold, the service requests in the waiting queue are satisfied until the real-time hardware load information reaches the second load threshold.

3. The method according to claim 2, characterized in that The satisfying the service request in the waiting queue comprises: The service requests in the waiting queue are satisfied according to the service priority or the receiving time of the service requests in the waiting queue, wherein the receiving time is the time when the service request is added to the waiting queue.

4. The method according to any one of claims 1 to 3, characterized in that: The hardware load information of the server in the first time window is periodically acquired, and the acquisition frequency of the periodic acquisition is positively correlated with the hardware load information of the server in the historical time window.

5. The method according to any one of claims 1 to 3, characterized in that: The length of the time window is negatively correlated with the hardware load information of the server in the historical time window.

6. A service scheduling device, characterized in that: The device comprises: an information determination module, a first scheduling module, a second scheduling module and a third scheduling module; The information determination module is used to determine the hardware load information of the server within a first time window when a service request is received at a first moment, wherein the first time window is the time window in which the first moment is located when the time axis is divided into a plurality of non-overlapping time windows; The first scheduling module is configured to satisfy the service request when the hardware load information within the first time window is lower than a first load threshold; The second scheduling module adds the service request to a waiting queue when the hardware load information in the first time window is higher than the first load threshold and lower than the second load threshold; The third scheduling module is used to downgrade the service request or reject the service request when the hardware load information in the first time window is higher than the second load threshold.

7. The device according to claim 6, characterized in that The device further includes: a fourth scheduling module; The fourth scheduling module is specifically used to: when the real-time hardware load information of the server is lower than the first load threshold, satisfy the service requests in the waiting queue until the real-time hardware load information reaches the second load threshold.

8. The device according to claim 7, characterized in that The fourth scheduling module is specifically used to satisfy the service requests in the waiting queue according to the service priority or receiving time of the service requests in the waiting queue, wherein the receiving time is the time when the service request is added to the waiting queue.

9. A service scheduling device, characterized in that: The device comprises: a memory and a processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the service scheduling method as described in any one of claims 1 to 5.

10. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the service scheduling method according to any one of claims 1 to 5 is implemented.

Citation Information

Cited By

  • Request scheduling and streaming output adjusting method and electronic equipment

    CN122412165A