System for dynamic scheduling of delay-sensitive resources based on multi-dimensional flow time series data perception
Patent Information
- Application Number
- CN202611058055.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-16
- Publication Date
- 2026-09-29
AI Technical Summary
[0002]在分布式软件架构中,现有的服务节点管理通常采用运行指标监控与容器扩容相结合的方式,软件编排组件定期采集各服务节点的计算资源占用情况,当监测值超过预设的静态阈值时,系统自动创建并注册新的容器实例,以分担短时间内出现的高并发请求,这类方案只有在负载升高后才开始扩容,通常默认各服务实例的负载彼此独立,当在时间上高度相关的硬件状态数据和高频交易计费请求同时大量进入系统时,请求量会在短时间内急剧波动,并在多级服务调用链上逐级积压,使长尾时延不断放大;与此同时,新容器实例的初始化需要一定时间,在扩容尚未完成的这段时间内,堵塞位置可能随着负载变化在不同服务节点之间转移,响应时延也会沿调用链向上游扩散,最终造成多个服务实例大范围超时并触发熔断
[0018]1、在多维流时序数据感知的时延敏感型资源动态调度中,通过持续分析微服务调用链中的长尾时延变化,并根据时延敏感型偏差系数及时调整不同服务实例的资源配置权重,可以在突发流量到来时优先保障设备控制服务等核心服务的计算资源,本发明直接利用集群内部已有资源进行调配,无需等待新容器创建、镜像拉取和服务注册,从而缓解瞬时高并发造成的调用积压,减少时延沿调用链逐级放大的情况。
Smart Images

Figure CN122845648A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a latency-sensitive dynamic scheduling system based on multidimensional streaming time-series data awareness, belonging to the field of microservice architecture technology. Background Technology
[0002] In distributed software architectures, existing service node management typically employs a combination of operational metric monitoring and container scaling. Software orchestration components periodically collect the computing resource usage of each service node. When the monitored value exceeds a preset static threshold, the system automatically creates and registers new container instances to handle high-concurrency requests occurring within a short period. This approach only begins scaling up after the load increases, usually assuming that the load of each service instance is independent. When a large number of time-correlated hardware status data and high-frequency transaction billing requests simultaneously enter the system, the request volume will fluctuate sharply within a short period and accumulate at each level of the multi-level service call chain, continuously amplifying long-tail latency. At the same time, the initialization of new container instances takes time. During this period before scaling is complete, the congestion location may shift between different service nodes as the load changes, and the response latency will also spread upstream along the call chain, ultimately causing multiple service instances to time out on a large scale and triggering circuit breakers.
[0003] Not only do underlying hardware limitations cause delays in the collection of distributed sensing data, but software-level control methods also have shortcomings. For example, Chinese invention patent application CN112463366A discloses a method and system for automatic scaling and automatic circuit breaking of microservices for cloud-native applications. The solution relies on sub-gateways to perform retrospective statistics on the error rate or concurrency at the interface granular level to trigger container scaling or service degradation. However, such control methods explicitly rely on the ideal working condition that the loads of service replica interfaces are independent of each other. When facing the transient long-tail impact of high-concurrency traffic in the Internet of Things, its retrospective scaling mechanism cannot get rid of the inherent physical delay of the cold start of virtualized containers. The blockage may spread upstream in a cascading manner before the scaling is completed. Moreover, the solution lacks a linkage mechanism between computing power quota adjustment and distributed transaction state switching. When computing power contention leads to increased response delay of lock-holding instances, simply discarding subsequent throughput requests cannot cut off the distributed lock resource contention path. The underlying synchronous database connection resources are still exhausted for a long time, ultimately causing multiple service instances to time out on a large scale and triggering circuit breakers.
[0004] Therefore, the technical problem to be solved by this invention is how to provide a dynamic scheduling method and system for latency-sensitive resources based on multi-dimensional streaming time-series data awareness to solve the failure problem of distributed service clusters under sudden traffic. Summary of the Invention
[0005] To address the problems in the background art, the technical solution of this invention is as follows: A time-delay-sensitive resource dynamic scheduling system based on multi-dimensional stream time-series data awareness, the system comprising:
[0006] The time-series feature mapping module is used to collect the long-tail call latency feature matrix between heterogeneous service nodes in a microservice architecture within a specific communication period, and generate a difference matrix that reflects the deterioration trend of call topology cascading by performing differential comparison on the latency features of two adjacent sampling time steps.
[0007] The dynamic topology scheduling module, connected to the time-series feature mapping module, is used to send a call topology routing circuit breaker command to the network gateway layer when the growth rate of the feature variance of the difference matrix exceeds the set judgment boundary. This disconnects the call topology routing from the front end to the non-core business microservice node and switches the non-core business microservice node to the set static pseudo-response state to intercept external throughput requests.
[0008] The multi-tenant isolation scheduling module, connected to the dynamic topology scheduling module, is used to obtain the current deadlock waiting queue length of the high-frequency polling microservice instance in concurrent transactions. It converts the current deadlock waiting queue length into a stepped penalty coefficient through the multi-tenant lock conflict nonlinear penalty operator, and uses the stepped penalty coefficient to apply rigid computing power suppression to the computing power limiting adjustment module, thereby lengthening the request retry interval of the high-frequency polling microservice instance.
[0009] Preferably, the multi-tenant isolation scheduling module is also used to send a business logic degradation flag to the distributed lock arbitration module when the computing power quota of the high-frequency polling microservice instance is reduced and the duration reaches the time window threshold of 50ms; the distributed lock arbitration module is used to convert the distributed strong consistency transaction between heterogeneous service nodes into an eventual consistency state according to the business logic degradation flag, and asynchronously write the payment success callback request to the distributed message middleware module, release the billing service's synchronous database connection occupation of the core call chain, and cut off the lock resource contention path.
[0010] Preferably, the time-series feature mapping module is also used to continuously collect the first time delay feature matrix of the current sampling time step and the second time delay feature matrix of the previous sampling time step, and generate a difference matrix by subtracting the second time delay feature matrix from the first time delay feature matrix.
[0011] Preferably, the dynamic topology scheduling module is also used to continuously analyze the growth rate of the characteristic variance of the difference matrix, and to trigger the issuance of a topology routing circuit breaker command when the growth rate of the characteristic variance continuously exceeds the set judgment boundary in a series of consecutive sampling periods.
[0012] Preferably, the time-series feature mapping module is also used to calculate the latency-sensitive deviation coefficient, specifically by dividing the long-tail latency fluctuation characteristic value of the device control microservice node in the collected microservice architecture by the current time window length, and adding the quotient to the product of the system inertia correction coefficient and the load monotonicity deviation; the multi-tenant isolation scheduling module is also used to dynamically correct the time window threshold according to the latency-sensitive deviation coefficient, wherein the numerical range of the system inertia correction coefficient is 0.5 to 1.5.
[0013] Preferably, the dynamic topology scheduling module is also used to control the non-core business microservice nodes to discard external throughput requests and directly return preset static data after disconnecting the call topology route from the front end to the non-core business microservice nodes.
[0014] Preferably, the long-tail call latency feature matrix collected by the time-series feature mapping module covers the long-tail call latency feature data of communication interactions between device control microservice, billing status microservice, and queuing allocation microservice in the microservice architecture.
[0015] Preferably, the multi-tenant isolation scheduling module is also used to reduce the resource scheduling weight according to the tiered penalty coefficient, thereby limiting the number of concurrent worker threads of high-frequency polling microservice instances at the virtualization container level.
[0016] Preferably, the network gateway layer is configured with a topology routing control set; the topology routing circuit breaker command issued by the dynamic topology scheduling module acts on the topology routing control set, and by modifying the routing addressing rules, it blocks the communication distribution from the front end to non-core business microservice nodes.
[0017] Compared with the prior art, the beneficial effects of the present invention are:
[0018] 1. In the dynamic scheduling of latency-sensitive resources with awareness of multi-dimensional streaming time-series data, by continuously analyzing the long-tail latency changes in the microservice call chain and adjusting the resource configuration weights of different service instances in a timely manner according to the latency-sensitive deviation coefficient, the computing resources of core services such as device control services can be prioritized when sudden traffic arrives. This invention directly utilizes existing resources within the cluster for allocation, without waiting for new container creation, image retrieval, and service registration, thereby alleviating the call backlog caused by instantaneous high concurrency and reducing the situation where latency amplifies step by step along the call chain.
[0019] 2. When resources of non-core billing services are restricted for a certain period of time, the database connection and distributed lock resources occupied by the billing service can be released in a timely manner by converting the distributed strong consistency transaction into an eventual consistency state and handing the payment success callback request to the distributed message middleware for asynchronous processing. This avoids service instances with insufficient computing power from holding locks for a long time, reduces the risk of lock waiting and deadlock spreading between microservices, and ensures the operational stability of the entire microservice cluster.
[0020] 3. By comparing the delay feature matrices of adjacent sampling time steps and analyzing the growth rate of the characteristic variance of the difference matrix, it is helpful to identify the cascading blocking that is worsening in the call topology. When the risk continues to increase, the system can suspend the front-end calls to non-core business microservice nodes and have these nodes return preset static data, thereby reducing the occupation of computing and communication resources by non-core requests, preventing sudden traffic from continuing to squeeze the core call chain, and ensuring the continuous availability of core businesses such as device control. Attached Figure Description
[0021] Figure 1 This is a flowchart of the linkage control of each module of the system of the present invention;
[0022] Figure 2 This is a structural diagram of the module composition of the system of the present invention.
[0023] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0024] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0025] A time-delay-sensitive resource dynamic scheduling system based on multi-dimensional streaming time-series data awareness, the system comprising:
[0026] The time-series feature mapping module is used to collect the long-tail call latency feature matrix between heterogeneous service nodes in a microservice architecture within a specific communication period, and generate a difference matrix that reflects the deterioration trend of call topology cascading by performing differential comparison on the latency features of two adjacent sampling time steps.
[0027] The dynamic topology scheduling module, connected to the time-series feature mapping module, is used to send a call topology routing circuit breaker command to the network gateway layer when the growth rate of the feature variance of the difference matrix exceeds the set judgment boundary. This disconnects the call topology routing from the front end to the non-core business microservice node and switches the non-core business microservice node to the set static pseudo-response state to intercept external throughput requests.
[0028] The multi-tenant isolation scheduling module, connected to the dynamic topology scheduling module, is used to obtain the current deadlock waiting queue length of the high-frequency polling microservice instance in concurrent transactions. It converts the current deadlock waiting queue length into a stepped penalty coefficient through the multi-tenant lock conflict nonlinear penalty operator, and uses the stepped penalty coefficient to apply rigid computing power suppression to the computing power limiting adjustment module, thereby lengthening the request retry interval of the high-frequency polling microservice instance.
[0029] Preferably, the multi-tenant isolation scheduling module is also used to send a business logic degradation flag to the distributed lock arbitration module when the computing power quota of the high-frequency polling microservice instance is reduced and the duration reaches the time window threshold of 50ms; the distributed lock arbitration module is used to convert the distributed strong consistency transaction between heterogeneous service nodes into an eventual consistency state according to the business logic degradation flag, and asynchronously write the payment success callback request to the distributed message middleware module, release the billing service's synchronous database connection occupation of the core call chain, and cut off the lock resource contention path.
[0030] Preferably, the time-series feature mapping module is also used to continuously collect the first time delay feature matrix of the current sampling time step and the second time delay feature matrix of the previous sampling time step, and generate a difference matrix by subtracting the second time delay feature matrix from the first time delay feature matrix.
[0031] Preferably, the dynamic topology scheduling module is also used to continuously analyze the growth rate of the characteristic variance of the difference matrix, and to trigger the issuance of a topology routing circuit breaker command when the growth rate of the characteristic variance continuously exceeds the set judgment boundary in a series of consecutive sampling periods.
[0032] Preferably, the time-series feature mapping module is also used to calculate the latency-sensitive deviation coefficient, specifically by dividing the long-tail latency fluctuation characteristic value of the device control microservice node in the collected microservice architecture by the current time window length, and adding the quotient to the product of the system inertia correction coefficient and the load monotonicity deviation; the multi-tenant isolation scheduling module is also used to dynamically correct the time window threshold according to the latency-sensitive deviation coefficient, wherein the numerical range of the system inertia correction coefficient is 0.5 to 1.5.
[0033] Preferably, the dynamic topology scheduling module is also used to control the non-core business microservice nodes to discard external throughput requests and directly return preset static data after disconnecting the call topology route from the front end to the non-core business microservice nodes.
[0034] Preferably, the long-tail call latency feature matrix collected by the time-series feature mapping module covers the long-tail call latency feature data of communication interactions between device control microservice, billing status microservice, and queuing allocation microservice in the microservice architecture.
[0035] Preferably, the multi-tenant isolation scheduling module is also used to reduce the resource scheduling weight according to the tiered penalty coefficient, thereby limiting the number of concurrent worker threads of high-frequency polling microservice instances at the virtualization container level.
[0036] Preferably, the network gateway layer is configured with a topology routing control set; the topology routing circuit breaker command issued by the dynamic topology scheduling module acts on the topology routing control set, and by modifying the routing addressing rules, it blocks the communication distribution from the front end to non-core business microservice nodes.
[0037] Example 1: A latency-sensitive resource dynamic scheduling system based on multi-dimensional streaming time-series data awareness runs on a cloud-native cluster control platform and is deployed in a cloud-native microservice architecture for self-service bathing and smart washing machines. The architecture includes user authentication service, device status monitoring service, device control microservice, billing status microservice, queuing allocation microservice, and historical data archiving service. The billing status microservice includes an order billing service instance. During specific high-concurrency periods, the concurrent control command stream sent by the front-end application and the hardware device status reporting stream are superimposed on the timeline, forming a spatiotemporal coupling impact. Concurrent transactions between the device control microservice, billing status microservice, and queuing allocation microservice are sensitive to call latency. As latency-vulnerable transactions, they are dynamically rate-limited by the system. Conventional post-event reactive container horizontal scaling has physical delays of hundreds of milliseconds for image retrieval and service registration initialization, making it difficult to digest transient long-tail call latency in a timely manner. As a result, the service call topology network experiences cascading long-tail blocking. After service nodes become stuck, distributed transaction locks continuously occupy the database connection pool, leading to the failure of the distributed group.
[0038] Under concurrent operating conditions, the traffic state awareness unit collects the long-tail call latency feature matrix of heterogeneous service nodes within a specific communication period through the microservice mesh communication proxy layer. The matrix encompasses the long-tail call latency characteristic data of communication interactions between the device control microservice, the billing status microservice, and the queuing allocation microservice. This data is then input into the time-series feature mapping module. The module uses a first-order linear sliding window variance operator to calculate the statistical variance of the call response time within the current discrete time window. After normalization using a unified time benchmark, the long-tail latency fluctuation characteristic value of the device control microservice node is obtained. Then, a first-order difference is made between the long-tailed time delay fluctuation characteristic values of adjacent sampling time steps to obtain the load monotonicity deviation. and according to the formula Calculate the time-delay sensitive deviation coefficient ,in, It is a dimensionless time-delay sensitive deviation coefficient; This represents the characteristic value of long-tail time delay fluctuation. The current time window length. This is the load monotonicity deviation. This is the system inertia correction factor, with a value ranging from 0.5 to 1.5. In this embodiment, it is taken as 1.0. In the first term of the calculation formula, [the following is used]. and The quotient is multiplied by a standardized reference time transformation factor of 1 / ms, so that both terms in the formula are converted into dimensionless values.
[0039] Long-tail call delay feature matrix The matrix is composed of rows and columns representing the calling source service node and the called target service node, respectively. Each matrix element records the long-tail response delay time observation of data packet interaction between heterogeneous service nodes within a specific communication period. For the characteristic variance growth rate of the difference matrix, the time-series feature mapping module performs eigenvalue decomposition on the difference matrix of the current sampling time step within a continuous discrete sampling period to extract the principal component variance. Then, the principal component variance value of the current sampling time step is subtracted from the principal component variance value of the previous sampling time step, and the resulting variance difference value is divided by the principal component variance value of the previous sampling time step to obtain the characteristic variance growth rate reflecting the deterioration trend of the call topology cascading.
[0040] At that time, the sensitive deviation coefficient When the safety scheduling threshold of 15 is exceeded, the dynamic topology scheduling module sends a resource adjustment instruction to the computing power rate limiting adjustment module. The computing power rate limiting adjustment module adjusts the container orchestration architecture through the kernel controller, activates the asymmetric hard isolation zone at the kernel layer, and adjusts the relative weight ratio of kernel processors for device control microservices and device status monitoring service instances on the core call chain. The initial equilibrium quota of 25% will be adjusted to 70% of the high-priority quota, while instances of historical data archiving service and historical order billing service in the billing status microservice will be reduced to the low-priority quota. The corresponding 15% of the released processor computing resources are allocated in situ within the microservice group to limit the instance computing power of non-core business microservice nodes, and the computing power is borrowed to the device control microservice that is experiencing long-tail blocking.
[0041] The multi-tenant isolation scheduling module uses the time window threshold given by the parameter adaptive calibration unit as the initial value, and then adjusts it according to the latency-sensitive deviation coefficient. The deviation from the safety scheduling threshold of 15 dynamically adjusts the time window threshold; when the deviation increases, the time window threshold is shortened, and when the deviation decreases, it is brought back to the initial value. Under this concurrent working condition, the adjusted time window threshold is 50ms. When the computing power quota of the historical order billing service instance in the billing status microservice is reduced and the duration reaches 50ms, the multi-tenant isolation scheduling module sends a business logic degradation flag to the distributed lock arbitration module. According to the business logic degradation flag, the distributed strong consistency transaction between heterogeneous service nodes is converted into an eventual consistency state, and the payment success callback request is asynchronously written to the distributed message middleware module, releasing the synchronous database connection occupation of the core call chain by the billing status microservice and cutting off the lock resource contention path.
[0042] After the eventual consistency state transition is completed, the backlog instruction processing latency of the device control microservice is reduced to less than 15μs. Processing latency refers to the local processing time of the kernel processor executing the instruction queue in-situ offset allocation and memory-level instruction parsing after stripping the overhead of distributed lock synchronization suspension waiting and external network transmission links. After the billing state microservice releases the occupation of the core call chain synchronous database connection, the core control instructions are written to the local shared memory of the microservice mesh and pipelined operations are performed using the 70% high-priority kernel processor quota released after the low-priority instance is compressed. 15μs corresponds to the processing latency of the local computing core, and 50ms corresponds to the global network propagation and call topology routing circuit breaker switching cycle. The measurement ranges of the two are different.
[0043] The temporal feature mapping module calculates the time delay feature difference between two adjacent consecutive sampling time steps and generates a difference matrix. ,in, It is a difference matrix composed of response time differences; This is the first delay feature matrix of the current sampling time step, composed of the currently collected response time data; This is the second delay feature matrix of the previous sampling time step, which is composed of the response time data collected in the previous period.
[0044] Before the system went live, the parameter adaptive calibration unit determined the set judgment boundaries and the system's maximum tolerance length through abnormal overload boundary testing. During the testing process, high-load conditions were simulated and the collapse threshold of the microservice architecture was monitored. When the feature variance growth rate reached 20%, cascading blocking began to spread exponentially, so the judgment boundaries were set. The value is set to 0.2, and the maximum concurrent capacity of the underlying kernel database connection pool is 128 connections. After engineering safety reduction, the maximum tolerance length of the system is determined to be 50 requests. If the value is too low, penalties will be frequently triggered erroneously, and if the value is too high, the connection pool will easily run out before the penalty takes effect.
[0045] When the growth rate of the eigenvariance of the difference matrix continuously exceeds the set judgment boundary within consecutive sampling periods. When the dynamic topology scheduling module sends a call topology routing circuit breaker command to the network gateway layer, the network gateway layer, which has pre-configured the call topology routing control set, modifies the routing addressing rules of the front end pointing to the non-core business microservice node upon receiving the command, disconnects the corresponding call topology route, and stops the communication distribution from the front end to the non-core business microservice node. For external throughput requests that have already reached the non-core business microservice node, the dynamic topology scheduling module controls the node to stop calling the business processing logic and discard the request payload, directly reads the pre-stored preset static data and returns it, so that the node switches to the set static pseudo-response state; subsequent external throughput requests no longer occupy the node's business processing resources.
[0046] During the topology routing circuit breaker invocation process, the multi-tenant isolation scheduling module obtains the current deadlock waiting queue length from the distributed lock manager. And through the multi-tenant lock conflict nonlinear penalty operator, according to the formula calculate The base value, where, This is a nonlinear sensitivity adjustment operator with a value of 1.2. This represents the current deadlock wait queue length. To determine the maximum tolerable length for the system, the multi-tenant isolation scheduling module performs range quantization on the base value: when the base value is no greater than 0.1, it will... Set to 0; when the base value is greater than 0.1 and not greater than 0.3, Set it to 0.2; when the base value is greater than 0.3, then... Set to 0.5, after quantization It is used as a tiered penalty coefficient for resource scheduling.
[0047] The multi-tenant isolation scheduling module uses a tiered penalty coefficient. The resource scheduling weight of the high-frequency polling microservice instance is reduced, and a rigid computing power suppression instruction is issued to the computing power throttling module. The resource scheduling weight is represented by the ratio of the current available concurrent worker threads of the instance to the maximum concurrent thread base. In this embodiment, Multiplying the number of concurrent threads by the preset maximum number of concurrent threads (64), the resulting product is the number of threads to be reduced. The computing power limiting adjustment module reduces the number of currently available concurrent worker threads of the instance at the virtualization container layer at the kernel layer. After the number of concurrent worker threads is reduced, the ability of the high-frequency polling microservice instance to handle external throughput requests decreases. The network gateway layer detects the increase in the instance's request rejection rate and connection queue backlog, triggering the backoff algorithm and lengthening the request retry interval for synchronous requests initiated by the front-end mini-program client.
[0048] After the resource quota adjustment was completed, the call response latency of the device control microservice decreased from 324.1ms and stabilized at 14.2ms. No order data loss or deadlock failures occurred in the microservice group. As the call topology routing converged locally, the effective utilization of physical computing resources increased. The distributed system suppressed the timing fluctuations of multidimensional flow by rearranging the control flow and data flow, and maintained the global availability state by degrading local transactions.
[0049] Example 2: This example demonstrates the experimental verification of a latency-sensitive resource dynamic scheduling system based on multidimensional streaming time-series data awareness on a cloud-native container cluster simulation platform with 12 heterogeneous computing nodes. An automated load generator is used to simulate the microservice call state of a smart washing machine and a self-service bathing control gateway under concurrent pressure. The original communication data is taken from the distributed call response time series recorded by the service mesh proxy log. During the experiment, a network latency jitter noise stream with an amplitude range of 5ms to 25ms and conforming to a normal distribution is injected into the communication bus layer to test the impact of environmental interference on the scheduling results.
[0050] The time-series feature mapping module adjusts the time window length based on the transient frequency of the multi-level call topology network. When the microservice communication proxy layer detects an increase in the rate of sudden changes in call response time, the time window length... By shrinking the range of values towards the lower limit to improve the tracking frequency of cascading blockage changes, this embodiment performs discrete sampling on the distributed topology change cycle and then adjusts the time window length. Set to 10 sampling periods, system inertia correction coefficient Take 1.0 and use it in the calculation with the dimension of 1 / ms; long-tail time delay ripple characteristic value Divide by the time window length Then, it is normalized by a standardized reference time transformation factor of 1 / ms to make the calculated time delay sensitive deviation coefficient... Keep it as a dimensionless value.
[0051] The simulation environment was divided into two independent test groups. The comparison group used a static threshold-triggered container horizontal scaling strategy, while the present invention group used the system for dynamic scheduling. The two groups ran under concurrent loads of 200 requests per second, 1000 requests per second, 1200 requests per second, 3000 requests per second, and 5000 requests per second, respectively. For each load gradient, the long-tail latency fluctuation characteristic value was recorded sequentially. Load monotonicity deviation Delay-sensitive deviation coefficient Current deadlock wait queue length The continuous calculation values of the nonlinear penalty operator for multi-tenant lock conflicts, the final response latency of the device control microservice, and the connection pool failure rate.
[0052] In the prototype of this invention, the multi-tenant isolation scheduling module determines the current deadlock waiting queue length based on the current deadlock waiting queue length. With the system's maximum tolerance length The ratio calculation is a continuous value, and the nonlinear sensitivity adjustment operator is used. When the value is 1.2, and the continuous values are no greater than 0.1, the step-by-step penalty coefficient is... Set to 0; for consecutive values greater than 0.1 and not greater than 0.3, the step-by-step penalty coefficient is applied. The value is set to 0.2; for consecutive values greater than 0.3, a step-wise penalty coefficient is applied. The value is set to 0.5, and the multi-tenant isolation scheduling module then applies a tiered penalty coefficient. Reduce the resource scheduling weight of high-frequency polling microservice instances and limit the number of concurrent worker threads at the virtualization container level. Compared with the sample group that does not perform the processing, the corresponding penalty coefficient is recorded as 0.00.
[0053] Under a load of 200 requests per second, the long-tailed latency fluctuation characteristics collected from the two sample groups. Both are 12.4ms, load monotonicity deviation. Both are 1.0ms. The 12.4ms time is normalized over 10 sampling periods, and then the correction value corresponding to 1.0ms is added. The time-delay sensitive deviation coefficients obtained from the two sample groups are... Both are 2.24, which does not reach the safe scheduling threshold of 15.0. Processor resources maintain the initial balanced quota. At this time, compare the current deadlock wait queue length of the sample group. For a single request, the penalty coefficient is recorded as 0.00, the final response latency of the device control microservice is 15.3ms, and the connection pool failure rate is 0.0%; the current deadlock wait queue length of the sample group in this invention is... For the same single request, with consecutive calculated values of 0.00, the corresponding tiered penalty coefficient... The value is 0, the final response latency is 14.8ms, and the connection pool failure rate is 0.0%.
[0054] After the concurrent load was increased to 1000 requests per second, the long-tail latency fluctuation characteristics of the two sample groups were... The load monotonicity deviation has increased to 45.2ms. All were increased to 11.5ms. The 45.2ms was normalized over 10 sampling periods and then added to the correction amount corresponding to 11.5ms to obtain the time-delay sensitive deviation coefficient. Both are 16.02, exceeding the safety scheduling threshold of 15.0. Based on this, the dynamic topology scheduling module in the sample of this invention triggers resource adjustments, activating the asymmetric hard isolation zone at the cloud-native orchestration engine layer. This increases the processor quota ratio of the device control microservice from the initial equilibrium state of 25% to a high-priority quota of 70%, while compressing the quota of the historical data archiving service to 15% of the low-priority quota. This allows the released computing power to be transferred to the device control microservice. After the adjustment, the current deadlock waiting queue length of the sample of this invention is... For four requests, with consecutive values of 0.00, the corresponding tiered penalty coefficient is... The value is 0, the final response latency is 18.2ms, and the connection pool failure rate is 0.0%. The comparison sample did not perform the aforementioned resource scheduling, and its current deadlock wait queue length is... With the number of requests increased to 8 and the penalty coefficient recorded as 0.00, the final response latency increased to 58.6ms, and the connection pool failure rate was 2.1%.
[0055] Under a load of 1200 requests per second, the service call topology of the comparison sample showed abrupt response changes, and its long-tail latency fluctuation characteristic value... Increased to 118.6ms, load monotonicity deviation The time delay is 18.9ms; after normalizing 118.6ms over 10 sampling periods, it is added to the correction amount corresponding to 18.9ms to obtain the time delay sensitive deviation coefficient. The current deadlock wait queue length is 30.76, and it continues to increase as call latency accumulates. With the number of requests increased to 24, the penalty coefficient remained at 0.00. The final response latency of the device control microservice reached 142.5ms, and the connection pool failure rate rose to 12.4%. Under the same load, the long-tail latency fluctuation characteristic value of the sample group after dynamic scheduling in this invention... The load monotonicity deviation is 32.1ms. The time delay is 8.4ms; after normalizing 32.1ms over 10 sampling periods, it is added to the correction amount corresponding to 8.4ms to obtain the time delay sensitive deviation coefficient. The value is 11.61, at which point the current deadlock wait queue length is... For 5 requests, the consecutive calculated value is 0.003, falling within the range no greater than 0.1, corresponding to the stepped penalty coefficient. The final response time was 13.9ms, and the connection pool failure rate was 0.0%.
[0056] When the concurrent load increases to 3000 requests per second, the long-tail latency fluctuation characteristic value of the comparison sample is... The load monotonicity deviation reached 254.3ms. The time delay-sensitive deviation coefficient is obtained by normalizing 254.3ms over 10 sampling periods and then adding the correction amount corresponding to 32.1ms. The connection request count for the billing microservice is 57.53, indicating a persistent backlog of connection requests. The current deadlock wait queue length is [not specified]. The number of requests increased to 85; with a penalty coefficient of 0.00, the final response latency of the device control microservice increased to 298.7ms, and the connection pool failure rate reached 56.0%. The long-tail latency fluctuation characteristic values collected by the sample group of this invention under the same load... The load monotonicity deviation is 42.6ms. The time delay is 15.3ms; after normalizing 42.6ms over 10 sampling periods, it is added to the correction amount corresponding to 15.3ms to obtain the time delay sensitive deviation coefficient. The value is 19.56, exceeding the safe scheduling threshold of 15.0. The dynamic topology scheduling module continues to adjust resources, and the current deadlock wait queue length is... To limit the number of requests to 32, the continuously calculated value of the nonlinear penalty operator for multi-tenant lock conflicts is 0.12, falling within the range greater than 0.1 and not greater than 0.3, with a stepped penalty coefficient. The value is set to 0.2; based on this, the multi-tenant isolation scheduling module limits the number of concurrent worker threads for high-frequency polling microservice instances, keeping the final response latency of the device control microservice at 15.4ms and the connection pool failure rate at 0.3%.
[0057] Under a load of 5000 requests per second, the long-tail latency fluctuation characteristics of the comparison sample were analyzed. Increased to 389.1ms, load monotonicity deviation Increased to 54.6ms, the normalized 389.1ms over 10 sampling periods is then added to the correction amount corresponding to 54.6ms to obtain the delay-sensitive deviation coefficient. The current deadlock queue length is 93.51, due to high-frequency queuing requests continuously holding kernel database connection locks. When 150 requests were received, the penalty coefficient remained at 0.00; the final response latency of the device control microservice increased to 425.6ms, and the connection pool failure rate reached 98.5%.
[0058] The long-tailed time delay fluctuation characteristic values collected by the sample group of this invention under the same load The load monotonicity deviation is 51.3ms. The time delay sensitive deviation coefficient is obtained by normalizing 51.3ms over 10 sampling periods and then adding the correction amount corresponding to 22.1ms. The value is 27.23, exceeding the safety scheduling threshold of 15.0. The dynamic topology scheduling module continues to adjust the call topology routing for non-core business microservice nodes, while the multi-tenant isolation scheduling module adjusts the current deadlock waiting queue length. With the number of requests controlled to 48, the continuously calculated value of the multi-tenant lock conflict nonlinear penalty operator is 0.28, falling within the range of greater than 0.1 and not greater than 0.3. The step-wise penalty coefficient... We set the value to 0.2 and apply rigid computing power suppression to high-frequency polling microservice instances accordingly, thereby releasing the thread scheduling space of the virtualization container layer.
[0059] When the computing power quota of the billing status microservice instance is reduced and continues to reach the time window threshold of 50ms, the multi-tenant isolation scheduling module sends a business logic degradation flag to the distributed lock arbitration module. Based on the flag, the distributed lock arbitration module converts the distributed strong consistency transaction between heterogeneous service nodes into an eventual consistency state and asynchronously writes the payment success callback request to the distributed message middleware module, releasing the synchronous database connection occupied by the billing status microservice to the core call chain. After the above processing, the final response latency of the device control microservice is 18.5ms and the connection pool failure rate is 1.2%.
[0060] Example 3: The current cloud-native cluster control platform is deployed with 12 heterogeneous service nodes running on multi-core dual-processor processors. The communication proxy layer between the nodes collects transient call response time data of high-concurrency service request streams in a non-intrusive manner during the sampling period, and forms a long-tail call latency feature matrix based on this data. During periods of sudden high load, the concurrent control commands and status reporting streams from the client create a spatiotemporal coupling impact on the timeline. As a result, the multi-level call topology network experiences cascading long-tail blocking. The long-tail latency fluctuation characteristic values of heterogeneous service nodes show a discrete surge. Conventional horizontal scaling of containers has a lag in initialization cold start time, making it difficult to digest transient traffic in a timely manner. The distributed locks held by the nodes continuously occupy the database connection pool, and after the connection resources are exhausted, a system-level timeout circuit breaker failure is triggered.
[0061] The parameter adaptive calibration unit injects a baseline load traffic of 50 consecutive sampling periods into the communication proxy layer in offline mode, and statistically analyzes the long-tail call latency feature matrix generated by the multi-level call chain under different throughput intensities. Establish time delay sensitive deviation coefficient The boundary change rate curve near the abnormal overload critical point is used to calibrate the safe scheduling threshold. The traffic status sensing unit collects the long-tail call delay feature matrix for three consecutive sampling cycles. The time-series feature mapping module uses a first-order linear sliding window variance operator to calculate the time-delay sensitive deviation coefficient. ,when When the value is continuously greater than 15, the computing power rate limiting adjustment module controls the underlying cloud-native orchestration engine to activate the asymmetric hard isolation zone at the kernel layer, adjusting the relative weight ratio of the kernel processors of the device control microservice node and the device status monitoring service instance. The initial 25% quota has been increased to 70% of the high-priority quota, while the compute quota for historical data archiving services and billing status microservice instances has been reduced to the low-priority quota. The corresponding 15% utilizes the computing power released by non-latency-sensitive services within the same physical cluster to control the backlog of requests on microservice nodes.
[0062] During the same scheduling process, the dynamic topology scheduling module continuously analyzes the growth rate of the characteristic variance of the difference matrix. When the growth rate of the characteristic variance continues to exceed the set judgment boundary, the dynamic topology scheduling module issues a call topology routing circuit breaker instruction to the network gateway layer, disconnects the call topology routing from the front end to the non-core business microservice node, and switches the non-core business microservice node to the set static pseudo-response state to limit the external throughput requests from continuing to occupy its processing resources.
[0063] When the compute quota of a billing-status microservice instance remains at the low-priority quota When the duration reaches the time window threshold of 50ms, the multi-tenant isolation scheduling module sends a business logic degradation flag to the distributed lock arbitration module. Based on the flag, the distributed lock arbitration module converts the distributed strong consistency transaction between heterogeneous service nodes into an eventual consistency state and asynchronously writes the payment success callback request to the distributed message middleware module, releasing the synchronous database connection occupied by the billing status microservice to the core call chain, so that the computing power borrowing and transaction state transformation are executed sequentially in the same scheduling link.
[0064] The kernel runtime logs of the cloud-native cluster control platform show that the instruction processing latency for device control microservice nodes decreased from 324.1ms during a sudden load surge and stabilized at 14.2ms; the distributed lock manager recorded the current deadlock wait queue length. The number of requests was reduced from 150 to 48, and the failure rate of the kernel database connection pool decreased from 98.5% under the conventional expansion method to 1.2%. Under the high load of 5,000 requests per second, the distributed cluster maintained a continuous communication distribution state.
[0065] Example 4: This example combines Figures 1 to 2 This section describes a dynamic scheduling system for latency-sensitive resources based on multidimensional streaming time-series data awareness, such as... Figure 1As shown, the time-series feature mapping module collects the long-tail call latency feature matrix and generates a difference matrix. After the difference matrix is used for judgment, it is linked with the dynamic topology scheduling module. The dynamic topology scheduling module issues a call topology routing circuit breaker instruction and switches to a static pseudo-response state to intercept requests. The dynamic topology scheduling module is connected to the multi-tenant isolation scheduling module through internal system linkage. The multi-tenant isolation scheduling module obtains the current deadlock waiting queue length and applies rigid computing power suppression. It sends a business logic degradation flag to the distributed lock arbitration module. The distributed lock arbitration module converts strong consistency to eventual consistency, cuts off the lock resource contention path, and asynchronously writes to the distributed message middleware module. The distributed message middleware module receives the payment success callback request and releases the synchronous database connection occupation.
[0066] like Figure 2 As shown, the time-series feature mapping module is connected to the dynamic topology scheduling module, the dynamic topology scheduling module is connected to the multi-tenant isolation scheduling module, the multi-tenant isolation scheduling module is connected to the distributed lock arbitration module, and the distributed lock arbitration module is connected to the distributed message middleware module.
[0067] Example 5: When the cloud-native cluster control base is deployed for the first time in the field, or when the network topology of heterogeneous service nodes changes, the parameter adaptive calibration unit drives the standard test traffic generator before the system is put into operation, injecting a progressive pulse-style simulated communication load stream into the microservice mesh communication proxy layer. The traffic state perception unit extracts the baseline delay data and constructs the initial state matrix within 100 consecutive sampling steps. The time-series feature mapping module performs discrete linear fitting based on the baseline delay data to obtain the system inertia correction coefficient under deadlock-free idle state. The initial reference value ensures that parameter calibration is completed before the actual control command flow from the front-end applet is accessed, and the calibrated system inertia correction coefficient is obtained. Keep it within the range of 0.5 to 1.5.
[0068] In the discrete linear fitting process, the time-series feature mapping module uses the time window length Using the basic independent variable and its reciprocal in the fitting, the long-tail latency fluctuation characteristic value of the device control microservice node is obtained. As the dependent variable, under the baseline operating condition of no deadlock idling, the long-tail time delay variability characteristic value With time window length The reciprocal of is linearly positively correlated; the intercept of the fitted line corresponds to the inherent background noise delay of the system, and the linear slope serves as the system inertia correction coefficient. The time-series feature mapping module substitutes the discrete data points obtained from 100 consecutive sampling steps into an algebraic equation that minimizes the sum of squared first-order linear residuals, and determines the system inertia correction coefficient through iterative solution. The initial baseline value.
[0069] System inertia correction factor After the initial baseline value is locked, the parameter adaptive calibration unit inputs a monotonically increasing timeout signal with stepped damping to the billing state microservice instance, and determines the current deadlock waiting queue length based on the data returned by the distributed lock arbitration module. Measure the deadlock escalation rate when the current deadlock wait queue length is... If the deadlock threshold is exceeded continuously for a duration that reaches the time window threshold of 50ms, the parameter adaptive calibration unit will apply a step-wise penalty coefficient. The data is sent to the multi-tenant isolation scheduling module, which then applies a tiered penalty coefficient. A rigid computing power suppression is applied to the computing power limiting adjustment module. The computing power limiting adjustment module completes the rigid allocation of resource quotas through the underlying container layer kernel resource controller. After calibration, the hardware processing latency output by the test node converges to within 15μs, the node enters a zero-blocking state, and each service instance is restored to a stable communication allocation state within the safety damping boundary.
[0070] Example 6: The current automatic load generator increases the overall cluster throughput to the physical limit of 6000 requests per second, injecting abnormal state disturbances into the cloud-native cluster control platform. The network data bus continuously collects the long-tail call latency feature matrix of communication drift. The time-series feature mapping module utilizes a first-order linear sliding window variance operator to perform online self-checking filtering on continuous discrete-time sampling points, and based on the long-tail delay fluctuation characteristic values of the device control microservice nodes. Calculate the time-delay sensitive deviation coefficient .
[0071] At that time, the sensitive deviation coefficient When the safety scheduling threshold of 15 is exceeded, the computing power rate limiting adjustment module controls the cloud-native orchestration engine to activate the asymmetric hard isolation zone at the kernel layer, adjusting the relative weight ratio of kernel processors. Increased to 70%, the distributed lock manager monitors the current deadlock wait queue length in real time during resource quota adjustments. When the computing power quota of a billing status microservice instance is reduced and the duration reaches the time window threshold of 50ms, the distributed transaction dynamic circuit breaker execution unit in the multi-tenant isolation scheduling module sends a business logic degradation flag to the distributed lock arbitration module.
[0072] Based on the identifier, the distributed lock arbitration module converts the distributed strong consistency transaction between heterogeneous service nodes into an eventual consistency state, and asynchronously writes the payment success callback request to the distributed message middleware module, releasing the synchronous database connection occupied by the billing status microservice to the core call chain, reducing the overhead of second-order thread processing, and cutting off the cascading deadlock path between distributed groups.
[0073] After the system has been running continuously for 24 hours, the kernel monitoring controller performs aging and reconstruction processing on the cache data area. This reduces the allocation weight of the historical long-tail call latency feature matrix through a time decay factor, ensuring the adaptive parameter matrix maintains real-time responsiveness when operating conditions change. The multi-tenant isolation scheduling module uses a multi-tenant lock conflict nonlinear penalty operator based on the current deadlock waiting queue length. With the system's maximum tolerance length The ratio is continuously calculated as a step penalty coefficient. Furthermore, by using coefficients to apply rigid computing power suppression to the computing power limiting adjustment module, the number of concurrent worker threads of high-frequency polling microservice instances at the virtualization container layer is limited, thereby lengthening their request retry interval.
[0074] After the extreme flood peak receded, the response latency of the device control microservice decreased from 324.1ms to 14.2ms, the kernel database connection pool failure rate stabilized below 1.2%, and the dynamic utilization rate of the computing resource pool of the microservice architecture returned to the steady-state operating range.
[0075] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A time-delay-sensitive resource dynamic scheduling system based on multi-dimensional streaming time-series data awareness, characterized in that, The system includes: The time-series feature mapping module is used to collect the long-tail call latency feature matrix between heterogeneous service nodes in a microservice architecture within a specific communication period, and generate a difference matrix that reflects the deterioration trend of call topology cascading by performing differential comparison on the latency features of two adjacent sampling time steps. The dynamic topology scheduling module, connected to the time-series feature mapping module, is used to send a call topology routing circuit breaker command to the network gateway layer when the growth rate of the feature variance of the difference matrix exceeds the set judgment boundary. This disconnects the call topology routing from the front end to the non-core business microservice node and switches the non-core business microservice node to the set static pseudo-response state to intercept external throughput requests. The multi-tenant isolation scheduling module, connected to the dynamic topology scheduling module, is used to obtain the current deadlock waiting queue length of the high-frequency polling microservice instance in concurrent transactions. It converts the current deadlock waiting queue length into a stepped penalty coefficient through the multi-tenant lock conflict nonlinear penalty operator, and uses the stepped penalty coefficient to apply rigid computing power suppression to the computing power limiting adjustment module, thereby lengthening the request retry interval of the high-frequency polling microservice instance.
2. The time-delay-sensitive resource dynamic scheduling system based on multi-dimensional streaming time-series data awareness according to claim 1, characterized in that, The multi-tenant isolation scheduling module is also used to send a business logic degradation flag to the distributed lock arbitration module when the computing power quota of a high-frequency polling microservice instance is reduced and the duration reaches the time window threshold of 50ms. The distributed lock arbitration module is used to convert distributed strong consistency transactions between heterogeneous service nodes into an eventual consistent state based on the business logic degradation flag, and asynchronously write the payment success callback request to the distributed message middleware module, thereby releasing the billing service's synchronous database connection occupation of the core call chain and cutting off the lock resource contention path.
3. The time-delay-sensitive resource dynamic scheduling system based on multi-dimensional streaming time-series data awareness according to claim 1, characterized in that, The time-series feature mapping module is also used to continuously collect the first time delay feature matrix of the current sampling time step and the second time delay feature matrix of the previous sampling time step, and generate a difference matrix by subtracting the second time delay feature matrix from the first time delay feature matrix.
4. A time-delay-sensitive resource dynamic scheduling system based on multi-dimensional streaming time-series data awareness as described in claim 1, characterized in that, The dynamic topology scheduling module is also used to continuously analyze the growth rate of the characteristic variance of the difference matrix, and to trigger the issuance of a topology routing circuit breaker command when the growth rate of the characteristic variance continuously exceeds the set judgment boundary in a series of consecutive sampling periods.
5. A time-delay-sensitive resource dynamic scheduling system based on multi-dimensional streaming time-series data awareness according to claim 1, characterized in that, The time-series feature mapping module is also used to calculate the latency-sensitive deviation coefficient. Specifically, it divides the long-tail latency fluctuation feature value of the device control microservice node in the collected microservice architecture by the current time window length, and adds the quotient to the product of the system inertia correction coefficient and the load monotonicity deviation. The multi-tenant isolation scheduling module is also used to dynamically correct the time window threshold based on the delay-sensitive deviation coefficient, wherein the numerical range of the system inertia correction coefficient is 0.5 to 1.
5.
6. A time-delay-sensitive resource dynamic scheduling system based on multi-dimensional streaming time-series data awareness according to claim 1, characterized in that, The dynamic topology scheduling module is also used to control the non-core business microservice nodes to discard external throughput requests and directly return preset static data after disconnecting the call topology route from the front end to the non-core business microservice nodes.
7. A time-delay-sensitive resource dynamic scheduling system based on multi-dimensional streaming time-series data awareness according to claim 1, characterized in that, The long-tail call latency feature matrix collected by the time-series feature mapping module covers the long-tail call latency feature data of communication interactions between device control microservice, billing status microservice and queuing allocation microservice in the microservice architecture.
8. A time-delay-sensitive resource dynamic scheduling system based on multi-dimensional streaming time-series data awareness according to claim 1, characterized in that, The multi-tenant isolation scheduling module is also used to reduce resource scheduling weights based on a tiered penalty coefficient, limiting the number of concurrent worker threads of high-frequency polling microservice instances at the virtualization container level.
9. A time-delay-sensitive resource dynamic scheduling system based on multi-dimensional streaming time-series data awareness according to claim 1, characterized in that, The network gateway layer is configured with a topology routing control set; the topology routing circuit breaker command issued by the dynamic topology scheduling module acts on the topology routing control set, and by modifying the routing addressing rules, it blocks the communication distribution from the front end to non-core business microservice nodes.
Citation Information
Patent Citations
Cloud native-oriented micro-service automatic capacity expansion and shrinkage and automatic fusing method and system
CN112463366A