Hardware resource release method, device and equipment facing transaction system and medium

CN122547554BActive Publication Date: 2026-09-29中信证券股份有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611039564.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-29
Estimated Expiration
2046-07-14

AI Technical Summary

Technical Problem

仅依托进程整体CPU占用率、请求平均耗时、任务队列深度等整体运行指标进行判定,无法定位引发异常的具体功能码与用户码,因此无法针对故障功能码、异常用户实单独进行限流或熔断操作,只能对整个服务进程执行全量限流熔断处理,强制回收进程全部线程,致使正常业务资源被误回收,全量限流熔断需要停止所有线程、清空队列,后续恢复时又需重新创建线程和分配内存,这种频繁的线程回收与重建操作,造成算力资源和内存资源的浪费

Benefits of technology

[0011]本公开的上述各个实施例具有如下有益效果:通过本公开的一些实施例的面向交易系统的硬件资源释放方法,避免了全量限流熔断造成的算力和内存资源浪费。具体来说,造成算力资源和内存资源的浪费的原因在于:仅依托进程整体CPU占用率、请求平均耗时、任务队列深度等整体运行指标进行判定,无法定位引发异常的具体功能码与用户码,因此无法针对故障功能码、异常用户实单独进行限流或熔断操作,只能对整个服务进程执行全量限流熔断处理,强制回收进程全部线程,致使正常业务资源被误回收,全量限流熔断需要停止所有线程、清空队列,后续恢复时又需重新创建线程和分配内存,这种频繁的线程回收与重建操作,造成算力资源和内存资源的浪费。基于此,本公开的一些实施例的面向交易系统的硬件资源释放方法,首先,采集预设交易系统中目标服务进程的多源服务运行遥测数据。由此,可以得到原始的目标服务进程在运行时的多维度运行数据。接着,对上述多源服务运行遥测数据进行字段语义对齐映射处理,得到线程运行快照信息集。由此,可以将不同来源、不同格式的数据统一为标准的线程运行快照信息集。然后,对上述线程运行快照信息集进行服务运行异常检测处理,得到至少一个异常线程信息,其中,每个异常线程信息包括功能码、用户码、业务信息和异常类型标识。由此,可以定位到引发异常的细粒度对象(具体是哪个功能码或用户码的请求),并区分异常类型(如是崩溃还是缓慢)。然后,对于上述至少一个异常线程信息中的每个异常线程信息,执行以下硬件资源释放处理:第一步,响应于确定上述异常线程信息包括的异常类型标识为崩溃标识,重启上述目标服务进程,以及执行与上述异常线程信息包括的功能码或异常线程信息包括的用户码对应的崩溃根因熔断操作或隔离操作,以限制上述异常线程信息包括的功能码或用户码所对应的后续请求占用目标服务进程的硬件资源。第二步,响应于确定上述异常线程信息包括的异常类型标识为缓慢标识,执行与上述异常线程信息包括的功能码或用户码对应的限速操作或隔离操作。由此,可以仅对真正引发异常的后续请求进行熔断、隔离或限速,而对正常请求不做任何限制。也因为只限制了异常功能码或用户码的后续请求,因此可以避免全量限流熔断带来的正常业务资源被误回收的问题;同时,由于不需要停止所有线程、清空队列及后续的线程重建,从而避免了频繁的线程回收与重建操作,减少了算力资源和内存资源的浪费。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122547554B_ABST
    Figure CN122547554B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a hardware resource release method and device for a transaction system, an apparatus, and a medium. A specific implementation of the method includes: collecting multi-source service running telemetry data of a target service process in a preset transaction system; performing field semantic alignment mapping processing on the multi-source service running telemetry data; performing service running anomaly detection processing on a thread running snapshot information set; for each abnormal thread information in at least one abnormal thread information, performing the following hardware resource release processing: performing a crash root cause fuse operation or an isolation operation corresponding to a function code included in the abnormal thread information or a user code included in the abnormal thread information; in response to determining that an abnormal type identifier included in the abnormal thread information is a slow identifier, performing a speed limiting operation or an isolation operation corresponding to the function code included in the abnormal thread information or the user code included in the abnormal thread information. The implementation reduces the waste of computing power resources and memory resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of computer technology, and more specifically to methods, apparatuses, devices, and media for releasing hardware resources for trading systems. Background Technology

[0002] As trading systems evolve towards high concurrency, low latency, and distributed architectures, service-oriented and in-memory architectures have become mainstream. Trading requests are typically routed to different target service processes based on function codes. Within these service processes, message reception and processing are handled through thread pools (including regular threads and RDT reliable data transmission threads) and task queues. Hardware resource release for trading systems refers to a technique that, when a trading service experiences crashes, slow responses, or thread freezes, performs degradation operations on subsequent requests that caused the exception, thereby releasing the hardware resources unused by processing those requests. Currently, when releasing unused or soon-to-be-used hardware resources when a trading service experiences crashes, slow responses, or thread freezes, the common approach is to collect overall operational metrics of the service process (such as CPU utilization, average latency, and queue depth). When these metrics exceed preset thresholds, rate limiting and circuit breaking are applied to all requests corresponding to the entire service process to release occupied CPU core cycles, memory pages, and thread stacks.

[0003] However, when using the above method to release hardware resources that have been invalidated by subsequent requests that caused exceptions during processing, the following technical problems often arise: Judging solely by overall process CPU utilization, average request time, and task queue depth, it is impossible to pinpoint the specific function code and user code that caused the exception. Therefore, it is impossible to perform rate limiting or circuit breaking operations on faulty function codes or abnormal user codes individually. Instead, a full rate limiting and circuit breaking process is performed on the entire service process, forcibly reclaiming all threads of the process. This results in the wrong reclamation of normal business resources. A full rate limiting and circuit breaking process requires stopping all threads and clearing the queue. During subsequent recovery, threads need to be recreated and memory allocated again. This frequent thread reclamation and reconstruction operation results in a waste of computing and memory resources.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the present disclosure concept, and therefore may contain information that does not form prior art known to those skilled in the art. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure provide methods, apparatuses, electronic devices, and computer-readable media for releasing hardware resources in trading systems to address one or more of the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide a hardware resource release method for a trading system. The method includes: collecting multi-source service operation telemetry data of a target service process in a preset trading system; performing field semantic alignment mapping processing on the multi-source service operation telemetry data to obtain a thread operation snapshot information set; performing service operation anomaly detection processing on the thread operation snapshot information set to obtain at least one abnormal thread information, wherein each abnormal thread information includes a function code, a user code, business information, and an anomaly type identifier; for each of the at least one abnormal thread information, performing the following hardware resource release processing: in response to determining that the anomaly type identifier included in the abnormal thread information is a crash identifier, restarting the target service process, and performing a crash root cause circuit breaker operation or isolation operation corresponding to the function code or user code included in the abnormal thread information, to limit the hardware resources occupied by subsequent requests corresponding to the function code or user code included in the abnormal thread information; in response to determining that the anomaly type identifier included in the abnormal thread information is a slow identifier, performing a rate limiting operation or isolation operation corresponding to the function code or user code included in the abnormal thread information.

[0008] Secondly, some embodiments of this disclosure provide a hardware resource release device for a trading system. The device includes: a collection unit configured to collect multi-source service operation telemetry data of a target service process in a preset trading system; a field semantic alignment mapping unit configured to perform field semantic alignment mapping processing on the multi-source service operation telemetry data to obtain a thread operation snapshot information set; an anomaly detection unit configured to perform service operation anomaly detection processing on the thread operation snapshot information set to obtain at least one abnormal thread information, wherein each abnormal thread information includes a function code, a user code, business information, and an anomaly type identifier; and a hardware resource release unit configured to release resources for the above-mentioned... For each of the at least one abnormal thread information, the following hardware resource release processing is performed: In response to determining that the exception type identifier included in the above abnormal thread information is a crash identifier, the target service process is restarted, and a crash root cause circuit breaker operation or isolation operation corresponding to the function code or user code included in the above abnormal thread information is executed to limit the hardware resources occupied by subsequent requests corresponding to the function code or user code included in the above abnormal thread information; In response to determining that the exception type identifier included in the above abnormal thread information is a slow identifier, a rate limiting operation or isolation operation corresponding to the function code or user code included in the above abnormal thread information is executed.

[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0011] The above-described embodiments of this disclosure have the following beneficial effects: The hardware resource release method for trading systems in some embodiments of this disclosure avoids the waste of computing and memory resources caused by full-scale rate limiting and circuit breaking. Specifically, the waste of computing and memory resources is due to the fact that relying solely on overall process CPU utilization, average request time, task queue depth, and other overall operating indicators makes it impossible to pinpoint the specific function code and user code that caused the anomaly. Therefore, it is impossible to perform rate limiting or circuit breaking operations separately for faulty function codes and abnormal user codes; only full-scale rate limiting and circuit breaking can be performed on the entire service process, forcibly reclaiming all threads of the process. This results in the erroneous reclamation of normal business resources. Full-scale rate limiting and circuit breaking requires stopping all threads and clearing the queue. Subsequent recovery requires recreating threads and allocating memory. This frequent thread reclamation and reconstruction operation leads to a waste of computing and memory resources. Based on this, the hardware resource release method for trading systems in some embodiments of this disclosure first collects multi-source service operation telemetry data of the target service process in a preset trading system. This allows obtaining the original multi-dimensional operating data of the target service process during runtime. Next, the telemetry data from the aforementioned multi-source service operation is processed with field semantic alignment mapping to obtain a set of thread operation snapshot information. This unifies data from different sources and in different formats into a standardized set of thread operation snapshot information. Then, service operation anomaly detection processing is performed on the aforementioned thread operation snapshot information set to obtain at least one abnormal thread information. Each abnormal thread information includes a function code, a user code, business information, and an anomaly type identifier. This allows for the location of the fine-grained object that triggered the anomaly (specifically, which function code or user code's request it is) and the differentiation of the anomaly type (e.g., crash or slowness). Then, for each of the at least one abnormal thread information, the following hardware resource release processing is performed: First, in response to determining that the anomaly type identifier included in the aforementioned abnormal thread information is a crash identifier, the target service process is restarted, and a crash root cause circuit breaker or isolation operation corresponding to the function code or user code included in the aforementioned abnormal thread information is executed to limit the hardware resources occupied by subsequent requests corresponding to the function code or user code included in the aforementioned abnormal thread information. The second step involves determining that the exception type identifier included in the aforementioned abnormal thread information is a slow identifier, and then executing rate limiting or isolation operations corresponding to the function code or user code included in the abnormal thread information. This allows for circuit breaking, isolation, or rate limiting only for subsequent requests that actually caused the exception, while leaving normal requests unrestricted. Because only subsequent requests with the abnormal function code or user code are restricted, the problem of normal business resources being mistakenly reclaimed due to full-scale rate limiting and circuit breaking is avoided. Furthermore, since it eliminates the need to stop all threads, clear the queue, and subsequently rebuild threads, frequent thread reclamation and rebuilding operations are avoided, reducing the waste of computing and memory resources. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0013] Figure 1 This is a flowchart of some embodiments of a hardware resource release method for a transaction system according to the present disclosure; Figure 2 This is a schematic diagram of the structure of some embodiments of a hardware resource release device for a transaction system according to the present disclosure; Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0018] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0019] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0020] Figure 1A flow 100 of some embodiments of a hardware resource release method for trading systems according to the present disclosure is shown. This hardware resource release method for trading systems includes the following steps: Step 101: Collect multi-source service operation telemetry data of the target service process in the preset transaction system.

[0021] In some embodiments, the execution entity (e.g., a computing device) of the hardware resource release method for a transaction system can collect multi-source service operation telemetry data of a target service process in a preset transaction system. Here, the preset transaction system refers to a pre-deployed software system for processing transaction requests. The target service process refers to an independently running program instance monitored within the preset transaction system, which has an independent process identifier (PID) and memory address space, and is responsible for handling one or more types of transaction requests. The multi-source service operation telemetry data refers to raw operational status data collected from the target service process from multiple different monitoring dimensions, including thread slice data, service function code performance statistics, message queue backlog data, and thread alarm status data. The thread slice data includes individual thread slice records. Each thread slice record is a snapshot of a worker thread at the sampling time, and includes at least the thread identifier, the currently processed function code, the user code, business information, and the current request processing duration and status flag (e.g., a flag indicating busy or not busy). The function code refers to the business type code of the transaction request, such as ORDER indicating order placement. The aforementioned user code refers to a unique identifier for the client or account initiating the transaction request, such as C10086. The aforementioned business information refers to other business context information besides the function code and user code (e.g., exchange code, security code, transaction quantity, price, etc.). The aforementioned service function code performance statistics refer to request processing performance metrics aggregated by function code dimension. These metrics include the total number of requests, the number of successful requests, the number of failed requests, the minimum processing time, the maximum processing time, and the average processing time within the statistics window. The aforementioned message queue backlog data refers to the current depth of the pending request queue (i.e., the number of unprocessed requests) in the aforementioned target service process. The aforementioned thread alarm status data refers to the timeout alarm information of the worker threads in the aforementioned target service process at the sampling time. This information includes various alarm record information, each alarm record including the thread identifier and the number of consecutive timeout periods. In practice, the aforementioned execution entity (e.g., a computing device in a local operation and maintenance platform or an agent node in a distributed monitoring system) can proactively initiate data query requests to the aforementioned target service process according to a preset collection cycle (e.g., once every 5 seconds) to obtain the raw running status data collected and cached in the aforementioned target service process. Optionally, the aforementioned execution entity can collect multi-source service operation telemetry data of the target service process in the preset transaction system by calling the monitoring interface exposed by the target service process (such as an API interface based on remote procedure call or HTTP protocol).

[0022] Step 102: Perform field semantic alignment mapping on the telemetry data of the multi-source service operation to obtain the thread running snapshot information set.

[0023] In some embodiments, the aforementioned execution entity can perform field semantic alignment mapping processing on the aforementioned multi-source service operation telemetry data to obtain a set of thread operation snapshot information. Each thread operation snapshot corresponds to a complete runtime snapshot of a worker thread in the aforementioned target service process at a single sampling moment. Each thread operation snapshot includes at least: the thread identifier, the currently processed function code, the user code, business information, cumulative processing time, average time consumption, queue depth, failure rate, thread busyness, and persistence coefficient. The aforementioned cumulative processing time originates from the "current request processed time" field in the aforementioned thread slice data and is a raw metric unique to that thread. The aforementioned average time consumption directly originates from the "average time consumption" field in the aforementioned service function code performance statistics and is a raw metric aggregated by function code dimension. Multiple threads processing the same function code share the same average time consumption value. The aforementioned queue depth originates from the "current depth" field in the aforementioned message queue backlog data and is a raw metric at the target service process level; all threads share this value. The aforementioned failure rate can be the ratio of the "number of failures" to the "total number of failures" in the aforementioned service function code performance statistics. The aforementioned thread busyness is a derived metric calculated based on the status statistics of all threads in the aforementioned thread slice data. Specifically, firstly, the total number of all worker threads in the aforementioned target service process is counted, along with the number of threads in the "busy" state (i.e., processing requests). Then, the ratio of busy threads to the total number of threads is calculated, i.e., busy threads divided by the total number of threads. The aforementioned persistence coefficient originates from the "number of consecutive timeout periods" field in the aforementioned thread alarm status data. It is a unique original metric for this thread, representing the number of sampling periods in which the cumulative processing time of this thread continuously exceeds a preset threshold. In practice, the aforementioned execution entity performs field semantic alignment mapping processing on the aforementioned multi-source service operation telemetry data according to the following steps to obtain a set of thread operation snapshot information: First, it iterates through all thread slice records in the aforementioned thread slice data, counts the total number of threads (i.e., the number of thread slice records), and the number of threads marked as "busy" (e.g., the status value is running or busy). The ratio of busy threads to the total number of threads is determined as the thread busyness as a global value shared by all threads. As an example: There are 20 thread slice records, of which 18 are in a "busy" state. The thread busyness is 18 / 20 = 0.9. Next, for each thread slice record, perform the following operations: Extract the thread identifier, function code, user code, business information, and current request processing time from the record. Determine the extracted current request processing time as the cumulative processing time. Then, search for records with the same function code in the aforementioned service function code performance statistics, and extract the average processing time, number of failures, and total number of failures for that function code. Next, determine the failure rate as the ratio of the number of failures to the total number of failures. Obtain the queue depth from the aforementioned message queue backlog data.Then, the system searches for warning records in the aforementioned thread alarm status data that match the thread identifier included in the thread slice records. The number of consecutive timeout periods included in the found warning records is determined as the persistence coefficient. If no record is found, the persistence coefficient is set to 0. Next, the extracted function codes, user codes, business information, cumulative processing time, average processing time, queue depth, as well as the determined failure rate, determined thread busyness, and determined persistence coefficient are aggregated into thread runtime snapshot information. Finally, the aforementioned execution entity can determine the obtained thread runtime snapshot information into a thread runtime snapshot information set.

[0024] Step 103: Perform service operation anomaly detection processing on the thread running snapshot information set to obtain at least one abnormal thread information.

[0025] In some embodiments, the aforementioned execution entity may perform service operation anomaly detection processing on the aforementioned thread execution snapshot information set to obtain at least one abnormal thread information. Each abnormal thread information includes a function code, user code, business information, and an anomaly type identifier. Each thread execution snapshot information in the aforementioned thread execution snapshot information set includes average processing time, cumulative processing time, queue depth of the target service process, failure rate, thread busyness, and persistence coefficient. The aforementioned anomaly type identifier refers to a marker used to distinguish anomaly categories, including crash identifier, deadlock identifier, and slow identifier, corresponding to three anomaly scenarios: service process crash, thread deadlock, and slow request, respectively.

[0026] In the process of adopting technical solutions to address the technical problems mentioned above, for the application scenario: in the anomaly detection of worker threads in a trading system under high concurrency, low latency, and distributed deployment, fine-grained anomaly type identification (crash, freeze, slow) of each worker thread's runtime snapshot (i.e., thread running snapshot information) is often accompanied by the following technical problems: simple threshold judgment based on a single or a few global indicators (such as average time consumption) cannot integrate multi-dimensional risk characteristics, resulting in low distinguishability of anomalies of different physical mechanisms, easy misjudgment, misjudging a (severe) freeze as a (mild) slow, thus causing the subsequent downgrade strategy selection to be incorrect. Furthermore, since crashes, freezes, and slowdowns each exhibit different physical evolution patterns (e.g., crashes are often accompanied by a persistent increase in the coefficient and a sharp rise in the failure rate; the core characteristic of a freeze is that the cumulative processing time continuously exceeds the threshold; and slowdowns are mainly manifested in an increase in queue depth, average processing time, and thread busyness), fixed weights cannot adapt to different abnormal scenarios. This results in poor sensitivity to differences, detection accuracy, and adaptability, leading to lower accuracy in subsequent degradation operations (i.e., crash root cause circuit breaking, isolation operations, and the aforementioned rate limiting operations), causing hardware resources to be continuously and ineffectively occupied. Therefore, the following characteristics are required for this application scenario: the ability to integrate multi-dimensional risk features and configure independent weight sets for different abnormal scenarios to identify the root cause of the anomaly, avoid erroneous degradation due to misjudgment, and thus ensure the effectiveness of subsequent hardware resource release operations.

[0027] In some optional implementations of certain embodiments, the aforementioned execution entity may perform service operation anomaly detection processing on the aforementioned thread execution snapshot information set through the following steps to obtain at least one abnormal thread information: The first step is to run snapshot information for each thread in the above thread snapshot information set, and perform the following steps: The first sub-step involves extracting the function code, user code, and business information currently being processed by the thread from the aforementioned thread runtime snapshot information, as well as extracting at least one risk characteristic indicator. This at least one risk characteristic indicator includes at least one of the following: average processing time, cumulative processing time, queue depth of the target service process, failure rate, thread busyness, and persistence coefficient. In practice, the executing entity can read the function code, user code, business information, and at least one risk characteristic indicator included in the thread runtime snapshot information.

[0028] The second sub-step involves obtaining at least one first weight value corresponding to a preset crash identifier. Each of these first weight values ​​corresponds one-to-one with a risk characteristic indicator among the aforementioned risk characteristic indicators. The first weight value refers to a pre-configured weighting coefficient used to calculate the crash risk level, with one first weight value corresponding to each risk characteristic indicator. For example, in a crash scenario, the persistence coefficient and failure rate can be given higher weights because these indicators are more strongly correlated with process crashes. The preset crash identifier is a fixed string (e.g., CRASH) used to mark crash exceptions. The execution entity can read at least one first weight value from a configuration file or memory.

[0029] The third sub-step involves generating a first risk level value based on at least one of the aforementioned risk characteristic indicators and at least one of the aforementioned first weight values. In practice, the implementing entity can multiply the value of each risk characteristic indicator by its corresponding first weight value, and then sum all the products to obtain the first risk level value.

[0030] The fourth sub-step involves merging the extracted function code, user code, business information, and crash identifier into abnormal thread information when the first risk level value is greater than or equal to a preset crash threshold. The preset crash threshold is a pre-defined value; when the first risk level value reaches or exceeds this value, the current thread is determined to have a crash risk. The executing entity combines the function code, user code, business information, and crash identifier into a single abnormal thread message and outputs this message, indicating that the function code or user code processed by this thread is highly likely to cause the service process to crash.

[0031] The fifth sub-step involves obtaining at least one second weight value corresponding to the preset crash threshold, in response to the first risk level value being less than the preset crash threshold. Each of the above-mentioned second weight values ​​corresponds one-to-one with a risk characteristic indicator among the above-mentioned at least one risk characteristic indicators. The second weight value refers to a pre-configured weighting coefficient used to calculate the crash risk level value. For crash scenarios, the weights of the cumulative processing time and persistence coefficient are typically set relatively high. The preset crash threshold is a fixed string (e.g., STUCK) used to mark crash anomalies.

[0032] The sixth sub-step generates a second risk level value based on at least one of the aforementioned risk characteristic indicators and at least one of the aforementioned second weight values. This second risk level value refers to a comprehensive score used to quantify the risk of the current thread getting stuck. The calculation method is similar to the third sub-step, using a weighted summation.

[0033] The seventh sub-step involves merging the extracted function code, user code, business information, and deadlock identifier into abnormal thread information in response to the second risk level value being greater than or equal to a preset deadlock threshold. The preset deadlock threshold is a pre-defined value used to determine whether a deadlock anomaly has occurred. When the second risk level value reaches or exceeds this threshold, the executing entity combines the function code, user code, business information, and deadlock identifier into a single abnormal thread message, indicating that the request currently being processed by that thread is deadlocked (e.g., in an infinite loop or permanently blocked).

[0034] The eighth sub-step, in response to the second risk level value being less than the preset deadlock threshold, generates a third risk level value based on the preset slow flag and at least one of the aforementioned risk characteristic indicators. In practice, the executing entity can obtain at least one third weight value corresponding to the slow flag. The third weight value in the at least one third weight value corresponds one-to-one with the risk characteristic indicator in the at least one risk characteristic indicator. The at least one risk characteristic indicator and the at least one third weight value are weighted and summed to obtain the third risk level value. The third risk level value refers to a comprehensive score used to quantify the risk of slowness in the current thread. For slow scenarios, the weights of average execution time, queue depth, and thread busyness are usually set relatively high. The preset slow flag is a fixed string (e.g., "SLOW") used to mark slow exceptions.

[0035] Step nine: In response to determining that the third risk level value is greater than or equal to the preset slow threshold, the extracted function code, user code, business information, and slow identifier are merged into abnormal thread information. The preset slow threshold is a pre-defined value used to determine whether a thread is considered slow or abnormal.

[0036] The above technical solution, combined with step 104 and related content, serves as an inventive point of this disclosure, solving the technical problem of "continuous invalid occupation of hardware resources." Factors leading to continuous invalid occupation of hardware resources often include: simple threshold judgment based on a single or limited number of global indicators (such as average processing time) fails to integrate multi-dimensional risk characteristics, resulting in low distinguishability of different physical mechanism anomalies and easy misjudgment, misjudging a severe freeze as a mild slow event, thus causing incorrect selection of subsequent degradation strategies. Furthermore, since crashes, freezes, and slow events each have different physical evolution patterns (e.g., crashes are often accompanied by a persistent increase in the persistence coefficient and a sharp increase in the failure rate; the core characteristic of a freeze is that the cumulative processing time continuously exceeds the threshold; and slow events are mainly manifested in an increase in queue depth, average processing time, and thread busyness), fixed weights cannot adapt to different anomaly scenarios, resulting in poor differentiation sensitivity, detection accuracy, and adaptability. This leads to lower accuracy in subsequent degradation operations (i.e., crash root cause circuit breaking, isolation operations, and the aforementioned rate limiting operations), resulting in continuous invalid occupation of hardware resources. If the above factors are addressed, the effect of reducing the continuous and ineffective occupation of hardware resources can be achieved. To achieve this effect, firstly, for each thread runtime snapshot information set mentioned above, the following steps are performed: First, extract the function code, user code, and business information currently being processed by the corresponding thread from the above thread runtime snapshot information, as well as extract at least one risk characteristic indicator. Thus, the function code, user code, and business information currently being processed by the corresponding thread, as well as at least one risk characteristic indicator (such as average processing time, cumulative processing time, queue depth, failure rate, thread busyness, and persistence coefficient) can be extracted from the above thread runtime snapshot information. Next, obtain at least one first weight value corresponding to a preset crash identifier, wherein the first weight value among the at least one first weight value corresponds one-to-one with the risk characteristic indicator among the at least one risk characteristic indicator. Based on the at least one risk characteristic indicator and the at least one first weight value, a first risk level value is generated. Thus, for a crash scenario, the first risk level value under the crash scenario can be determined by at least one first weight value, i.e., a unique physical law (the persistence coefficient and failure rate have higher weights). Next, in response to the first risk level value being greater than or equal to a preset crash threshold, the extracted function code, user code, business information, and crash identifier are merged into abnormal thread information. Then, in response to the first risk level value being less than the preset crash threshold, at least one of the following second weight values ​​corresponding to a preset stuck identifier is obtained, wherein each of the at least one second weight value corresponds one-to-one with a risk feature indicator among the at least one risk feature indicators. Finally, a second risk level value is generated based on the at least one risk feature indicator and the at least one second weight value.Therefore, when the first risk level value is less than the crash threshold, at least one second weight value corresponding to the stuck flag (e.g., higher weights for cumulative processing time and persistence coefficient) can be obtained. This allows for a specific quantitative evaluation of the physical characteristics of the stuck scenario (cumulative processing time consistently exceeding the threshold, high persistence coefficient), avoiding the masking of typical stuck features by using the same weights as crash or slowness, thereby improving the accuracy of stuck anomaly identification. Next, in response to the second risk level value being greater than or equal to a preset stuck threshold, the extracted function code, user code, business information, and stuck flag are merged into abnormal thread information. Then, in response to the second risk level value being less than the preset stuck threshold, a third risk level value is generated based on a preset slowness flag and at least one of the aforementioned risk characteristic indicators. This allows for the generation of a third risk level value based on the physical characteristics of the slow scenario, avoiding the masking of typical slowness features by using the same weights as crash or stuckness. Then, in response to determining that the third risk level value is greater than or equal to a preset slowness threshold, the extracted function code, user code, business information, and slowness flag are merged into abnormal thread information. Finally, in step 104, corresponding hardware resource release processing (circuit breaking, isolation, or rate limiting) is performed for each generated abnormal thread information. Because of the use of multi-dimensional risk feature weighted fusion and independent configuration of weights according to scenarios, the different physical mechanisms of the three types of anomalies can be accurately distinguished. This avoids false positives and false negatives caused by misjudging a stuck state as slow or fixed weights, thus making subsequent degradation operations (crash root cause circuit breaking, isolation, and rate limiting) more accurate and ensuring the effectiveness of subsequent hardware resource release operations, reducing the continuous and ineffective occupation of hardware resources.

[0037] Step 104: For each abnormal thread information in at least one abnormal thread information, perform the following hardware resource release process: Step 1041: In response to determining that the exception type identifier included in the exception thread information is a crash identifier, restart the target service process, and execute the crash root cause circuit breaker operation or isolation operation corresponding to the function code or user code included in the exception thread information, so as to limit the hardware resources occupied by the subsequent requests corresponding to the function code or user code included in the exception thread information.

[0038] In some embodiments, the executing entity may, in response to determining that the exception type identifier included in the exception thread information is a crash identifier, restart the target service process and execute a crash root cause circuit breaker or isolation operation corresponding to the function code or user code included in the exception thread information, to limit the hardware resources occupied by subsequent requests corresponding to the function code or user code included in the exception thread information. Here, the crash identifier refers to the identifier in the exception type identifier used to mark a service process crash. In practice, the executing entity can send a startup command to the target service process through the operating system to restart it.

[0039] In some optional implementations of certain embodiments, the aforementioned execution entity may perform a circuit breaker or isolation operation corresponding to the function code or user code included in the aforementioned abnormal thread information through the following steps, in order to limit the hardware resources occupied by subsequent requests corresponding to the function code or user code included in the aforementioned abnormal thread information from the target service process: The first step is to obtain the priority information corresponding to the aforementioned function code or user code. This priority information refers to the pre-assigned importance level for each function code or user code, such as high priority (core transactions, automatic downgrading not allowed), medium priority (rate limiting allowed), and low priority (isolation or circuit breaking allowed). This priority information includes priority mapping information. Each priority mapping information includes the function code or user code, and the corresponding level (e.g., high, medium, low).

[0040] The second step, based on the aforementioned priority information, in response to determining that the aforementioned function code or user code meets the preset isolation conditions, generates isolation control information for the aforementioned function code or user code. The preset isolation conditions can be defined as a low priority level for the function code or user code, as determined from the priority information. The isolation control information refers to the instructions used to instruct the local execution engine of the target service process to perform isolation operations on subsequent requests matching the aforementioned function code or user code. In practice, the execution entity can use a preset isolation operation instruction template, filling the aforementioned function code or user code as parameters into the corresponding placeholders in the isolation operation instruction template. The execution entity determines the filled isolation operation instruction template (or converts it to binary protocol buffer format) as the isolation control information for the aforementioned function code or user code. The isolation operation instruction template is a predefined command format with placeholders, used to generate isolation control information for a specific function code or user code. The aforementioned local execution engine can be the local (transaction system) execution engine.

[0041] Third, based on the aforementioned priority information, in response to determining that the aforementioned function code or user code does not meet the preset isolation conditions, a crash root cause circuit breaker control information is generated for the aforementioned function code or user code. This crash root cause circuit breaker control information refers to an instruction used to instruct the local execution engine of the aforementioned target service process to perform a circuit breaker operation on subsequent requests matching the aforementioned function code or user code. In practice, the executing entity can use a preset circuit breaker operation instruction template, filling the aforementioned function code or user code as parameters into the corresponding placeholders in the circuit breaker operation instruction template. The executing entity determines the circuit breaker operation instruction template filled with placeholders (or converts it to a binary protocol buffer format) as the circuit breaker control information for the aforementioned function code or user code. The aforementioned circuit breaker operation instruction template refers to a pre-defined rule command format with placeholders, used to generate circuit breaker control information for a specific function code or user code.

[0042] The fourth step involves sending the generated isolation control information or crash root cause circuit breaker control information to the local execution engine of the target service process. This enables the local execution engine to perform crash root cause circuit breaker or isolation operations on subsequent requests that match the function code or user code, thereby releasing at least one of the following hardware resources that would be occupied by the subsequent requests: processor resources and memory resources. The execution entity can send the isolation control information or circuit breaker control information (as instruction data) to the local execution engine of the target service process via network protocols (e.g., Remote Procedure Call (RPC), HTTP, or a custom TCP protocol). Upon receiving the isolation control information, the execution engine can match the target with subsequent requests that match the function code or user code indicated by the isolation control information. Then, it creates an independent isolation queue (with a preset capacity) in memory and starts one or more low-priority consumer threads. When a subsequent request arrives, if the function code or user code of the subsequent request matches the target, the request is placed in the isolation queue instead of the normal shared request queue. The consumer threads retrieve requests from the isolation queue at a preset low rate (e.g., 5 requests / second) for processing. Once the isolation queue is full, subsequent matching requests are directly rejected. This ensures that core business requests in the normal shared request queue remain unaffected, thus freeing up thread stacks and CPU time slices that would otherwise be occupied by these abnormal requests. Optionally, upon receiving the crash root cause circuit breaker control information, the aforementioned execution engine, for each subsequent arriving request, if its function code or user code matches the matching target (i.e., function code or user code) contained in the crash root cause circuit breaker control information, directly returns a rejection response (e.g., HTTP 429 or a custom error code), without performing any business processing, allocating thread stacks, consuming CPU cycles, or allocating memory pages.

[0043] Step 1042: In response to determining that the exception type identifier included in the exception thread information is a slow identifier, perform a rate limiting operation or isolation operation corresponding to the function code or user code included in the exception thread information.

[0044] In some embodiments, the execution entity may, in response to determining that the exception type identifier included in the exception thread information is a slow identifier, perform a rate-limiting operation or isolation operation corresponding to the function code or user code included in the exception thread information. The exception thread information corresponds to a thread runtime snapshot in the thread runtime snapshot information set, and each thread runtime snapshot in the set includes average processing time, cumulative processing time, queue depth of the target service process, failure rate, thread busyness, and persistence coefficient.

[0045] In addressing the anomaly handling issues in the background technology using the aforementioned technical solutions, for application scenarios such as online e-commerce flash sale systems, ticketing transaction systems (e.g., ticket purchase), and real-time inventory deduction systems—systems highly sensitive to response latency—performing precise rate limiting or isolation operations on function codes or user codes identified as slow often presents the following technical problems: Due to the lack of quantitative grading of the severity of slowness, it is difficult to distinguish between mild slowness (instantaneous high pressure) and severe slowness (continuous backlog). This means that uniformly applying rate limiting may not alleviate severe congestion, while uniformly applying isolation may cause unnecessary processing delays for mild slowness. The following requirements are necessary for this application scenario: For transaction systems highly sensitive to response latency, it is necessary to be able to distinguish between mild slowness (instantaneous high pressure) and severe slowness (continuous backlog) and automatically select the appropriate operation accordingly. For example, for mild slowness, rate limiting should be implemented (directly rejecting requests exceeding the limit without introducing queuing delays), and for severe slowness, isolation should be implemented (placing requests in an independent queue for low-speed processing to avoid congestion in the normal queue), thereby reducing processing latency.

[0046] In some optional implementations of certain embodiments, the aforementioned execution entity may perform rate-limiting or isolation operations corresponding to the function code or user code included in the aforementioned abnormal thread information through the following steps: The first step is to identify the thread runtime snapshot information corresponding to the abnormal thread information in the above thread runtime snapshot information set as the target thread runtime snapshot information.

[0047] The second step involves generating a slowness anomaly severity value based on the target thread's runtime snapshot information, including average processing time, cumulative processing time, queue depth, thread busyness, and persistence coefficient. This slowness anomaly severity value is a numerical measure of the severity of slow processing in the function code or user code corresponding to the current thread. In practice, for each parameter among average processing time, cumulative processing time, queue depth, thread busyness, and persistence coefficient, the executing entity can determine at least one third weight value corresponding to that parameter as a target weight value. Next, the executing entity can multiply the target weight value by the aforementioned parameter to determine the target value. Finally, the executing entity can sum the obtained target values ​​to determine the slowness anomaly severity value.

[0048] Third, in response to a slowness anomaly value being less than a preset slowness threshold, the following rate limiting operation is performed: The first sub-step involves the local execution engine controlling the target service process maintaining a time window counter. When the number of subsequent requests corresponding to the aforementioned function code or user code exceeds a preset allowable number within a unit of time, the subsequent requests corresponding to the function code or user code are rejected to release at least one of the following hardware resources that would otherwise be occupied by the subsequent requests: processor resources and memory resources. The preset slowness threshold is a pre-set value used to distinguish between mild and severe slowness. When the slowness anomaly value is less than this threshold, it is determined to be mild slowness, and rate limiting is performed. When it is greater than or equal to this threshold, it is determined to be severe slowness, and isolation is performed. The aforementioned time window counter can be a fixed window counter or a sliding window counter. The local execution engine maintains an independent counter for each rate-limited function code or user code. At the beginning of the unit time window, the counter is reset to zero. The counter increments by 1 for each matching request arriving within the window. If the counter value has reached the preset allowable number, the subsequent requests are directly rejected (e.g., a "system busy" error code is returned) without any business processing. Rejected requests do not allocate thread stacks, CPU time slices, or memory pages, thus freeing up these hardware resources for normal requests. The preset allowable number is pre-configured based on the priority of the function code or user code and the system capacity. For example, for low-priority batch query function codes, 20 requests per second can be allowed.

[0049] Fourth, in response to a slow anomaly level value greater than or equal to a preset slowness threshold, the following isolation operations are performed: First, the local execution engine controlling the target service process places subsequent requests corresponding to the aforementioned function code or user code into a preset isolation queue, independent of the normal request queue. A consumer thread with a processing rate lower than a preset processing rate is assigned to this isolation queue to process the requests. This preset isolation queue is an independent memory queue, physically isolated from the target service process's original shared request queue. This queue has a preset capacity limit (e.g., 500). The consumer thread is one or more threads specifically designed to retrieve and process requests from the isolation queue, with their processing rate limited to below the preset processing rate (e.g., 5 requests per second). Requests in the normal request queue are processed by the existing high-priority thread pool, completely unaffected by the isolation queue. In this way, abnormal requests are confined to the isolation queue for slow processing, preventing them from consuming CPU time slices and thread stack resources of normal requests.

[0050] Second, in response to the determination that the number of requests in the aforementioned isolation queue equals the maximum capacity of the isolation queue, subsequent requests are rejected. In practice, when the aforementioned isolation queue is full (i.e., the number of requests waiting to be processed in the queue has reached the preset maximum capacity), for any newly arriving subsequent requests matching the function code or user code, the aforementioned local execution engine directly returns a rejection response.

[0051] Optionally, the aforementioned implementing entity may also perform the following steps: The first step involves continuously collecting the average execution time of the function code, the queue depth of the pending queue, and the thread busyness of the target service process at a preset period during the effective period of the aforementioned root cause circuit breaker or isolation operation. The collected average execution time of the function code, the queue depth, and the thread busyness are then subjected to fallback verification processing to obtain fallback verification information. The effective period refers to the time from when the aforementioned operation (root cause circuit breaker or isolation operation, or the aforementioned rate limiting or isolation operation) is issued to the local execution engine and begins execution, until the operation is actively revoked or automatically expires. The preset period refers to a pre-set collection time interval, such as 5 seconds or 10 seconds. The pending queue refers to the queue in the target service process that stores unprocessed request messages. Fallback verification information can indicate whether the metrics have fallen back to the normal range, such as a Boolean value ("fallback" or "not fallback").

[0052] The second step is to perform the following downgrade processing when all the corresponding fallback verification information for a consecutive preset number of cycles indicates a fallback to below the standard value: The first sub-step involves replacing each of the root cause circuit breaker operations with a rate-limiting operation using the same function code or user code. The preset number of cycles is a configurable integer, such as 3 cycles (meaning three consecutive data collections showing a decline).

[0053] The second sub-step involves replacing each isolation operation within all isolation operations with a rate-limiting operation using the same function code or user code.

[0054] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the technical problem of "causing unnecessary processing delays." Factors leading to unnecessary processing delays often include: due to the lack of quantitative grading of the severity of slowness, it is difficult to distinguish between mild slowness (instantaneous high pressure) and severe slowness (continuous backlog). This means that uniformly applying rate limiting may not alleviate severe congestion, while uniformly applying isolation may cause unnecessary processing delays for mild slowness. Solving these factors can reduce unnecessary processing delays. To achieve this, firstly, the thread runtime snapshot information corresponding to the abnormal thread information in the above-described thread runtime snapshot information set is determined as the target thread runtime snapshot information. This allows the acquisition of runtime multi-dimensional indicators (average processing time, cumulative processing time, queue depth, thread busyness, and persistence coefficient) directly associated with the abnormal thread. Then, based on the average processing time, cumulative processing time, queue depth, thread busyness, and persistence coefficient included in the target thread runtime snapshot information, a slowness anomaly value is generated. This yields a quantitative slowness severity score, which integrates multiple dimensions such as backlog pressure, time anomaly, resource busyness, and persistence. Therefore, based on the comparison between the severity value and a preset slowness threshold, the system can automatically decide whether to perform rate limiting or isolation operations: If the slowness severity value is less than the preset slowness threshold, rate limiting is performed. This controls the target service process's local execution engine to maintain a time window counter. When the number of subsequent requests corresponding to the aforementioned function code or user code exceeds a preset allowable number within a unit of time, these subsequent requests are rejected to release processor and memory resources occupied by them. Thus, for minor slowness, only requests exceeding the limit are directly rejected, reducing queuing or isolation delays and ensuring normal request responses. If the slowness severity value is greater than or equal to the preset slowness threshold, isolation is performed. This controls the target service process's local execution engine to place the subsequent requests corresponding to the aforementioned function code or user code into a preset isolation queue independent of the normal request queue, and allocates a consumer thread with a processing rate lower than the preset processing rate to process the requests in the isolation queue. When the isolation queue is full, subsequent requests are rejected. Thus, for persistently backlogged severe slowness, abnormal traffic is completely isolated to an independent queue and processed at a lower rate, preventing it from crowding out the hardware resources of the normal queue. By quantifying and classifying the severity of slowness, we can accurately distinguish between mild slowness (instantaneous high pressure) and severe slowness (continuous backlog), and adaptively select rate limiting or isolation operations accordingly. For mild slowness, rate limiting is applied, and only requests exceeding the limit are directly rejected, avoiding the queuing delay caused by isolation, thereby reducing unnecessary processing wait. For severe slowness, isolation is applied, and abnormal requests are placed in an independent queue for low-speed processing to prevent congestion from spreading, while protecting the low-latency response of normal queues.

[0055] In some optional implementations of certain embodiments, the aforementioned execution entity may perform fallback verification processing on the average time of the collected function codes, the aforementioned queue depth, and the aforementioned thread busyness through the following steps to obtain fallback verification information: The first step involves determining that the average time consumed by the aforementioned function codes is less than or equal to a preset average time consumed by function codes, the aforementioned queue depth is less than or equal to a preset queue depth, and the aforementioned thread busyness is less than or equal to a preset busyness. Information indicating a fallback to below the standard value is then identified as fallback verification information. Here, the preset average time consumed by function codes threshold is a pre-set time value, representing the upper limit of normal average time consumed when the system is healthy. The preset queue depth threshold is a pre-set numerical value, such as 200, representing the upper limit of queue depth under normal conditions. The preset thread busyness threshold is a pre-set percentage value, such as 0.6 (60%), representing the upper limit of thread busyness under normal conditions. When all three conditions are met, the execution entity generates a Boolean value indicating a fallback to below the standard value as fallback verification information.

[0056] Optionally, the aforementioned implementing entity may also perform the following steps: The first step, in response to determining that the exception type identifier included in the above-mentioned abnormal thread information is a stuck identifier, and based on the function code and business information included in the above-mentioned abnormal thread information, a stuck detection circuit breaker operation is performed. Here, the aforementioned stuck identifier refers to the value (e.g., STUCK) in the exception type identifier used to mark a thread stuck exception.

[0057] In some optional implementations of certain embodiments, the aforementioned execution entity may perform a deadlock detection circuit breaker operation based on the aforementioned abnormal thread information, including function codes and business information, through the following steps: The first step is to concatenate the function code and business information into a string or structure as the target object for deadlock detection. The function code and the various fields in the business information are then concatenated into a string in a predetermined order (e.g., "function code|exchange|securities category|trading category|securities code") or assembled into a structure containing multiple fields (such as a struct in C or a JSON object).

[0058] The second step involves recording the current time and controlling the local execution engine to use the aforementioned stuck detection target object as a matching condition, rejecting subsequent requests with the same function code and business information. The current time refers to the moment the execution entity determines the system is stuck and begins executing the circuit breaker operation. In practice, the local execution engine maintains a mapping table in memory with this target object as the key. When a subsequent request arrives, the local execution engine extracts the function code and business information from the request and concatenates them into the same string. If it matches the key in the mapping table, it directly returns a rejection response without performing any business processing.

[0059] The third step involves the local execution engine, after a preset time period following the current time, identifying any request with the same function code and business information as a probe request upon receiving such a request. This probe request is then processed normally, and its processing time is recorded. The preset time period can be a pre-defined circuit breaker duration (e.g., 60 seconds). After this period, the local execution engine no longer automatically rejects matching requests but instead marks the next arriving request with the same function code and business information as a probe request, allowing it to enter the business processing flow normally (i.e., be processed like other normal requests). Simultaneously, the local execution engine records the time elapsed from the start to the end of processing this probe request (i.e., processing duration). During this period, other subsequent requests can still be rejected, but typically only the first request is considered a probe.

[0060] The fourth step involves responding to the determination that the processing time of the aforementioned probe request is less than a preset recovery threshold, allowing subsequent requests with the same function code and business information to be processed normally. The preset recovery threshold is a pre-set duration (e.g., 50 milliseconds) representing the upper limit of the system's recovery time. If the processing time of the probe request is less than this threshold, it indicates that the problem causing the freeze has been resolved (e.g., temporary resource contention has disappeared), and the circuit breaker operation for requests with the same function code and business information will no longer be performed; that is, all subsequent requests with the same function code and business information will resume normal processing. If the processing time of the probe request is still greater than or equal to the threshold, it indicates that the freeze problem still exists, and a warning message can be sent to a preset maintenance section (e.g., the terminal equipment used by maintenance personnel).

[0061] The above-described embodiments of this disclosure have the following beneficial effects: The hardware resource release method for trading systems in some embodiments of this disclosure avoids the waste of computing and memory resources caused by full-scale rate limiting and circuit breaking. Specifically, the waste of computing and memory resources is due to the fact that relying solely on overall process CPU utilization, average request time, task queue depth, and other overall operating indicators makes it impossible to pinpoint the specific function code and user code that caused the anomaly. Therefore, it is impossible to perform rate limiting or circuit breaking operations separately for faulty function codes and abnormal user codes; only full-scale rate limiting and circuit breaking can be performed on the entire service process, forcibly reclaiming all threads of the process. This results in the erroneous reclamation of normal business resources. Full-scale rate limiting and circuit breaking requires stopping all threads and clearing the queue. Subsequent recovery requires recreating threads and allocating memory. This frequent thread reclamation and reconstruction operation leads to a waste of computing and memory resources. Based on this, the hardware resource release method for trading systems in some embodiments of this disclosure first collects multi-source service operation telemetry data of the target service process in a preset trading system. This allows obtaining the original multi-dimensional operating data of the target service process during runtime. Next, the telemetry data from the aforementioned multi-source service operation is processed with field semantic alignment mapping to obtain a set of thread operation snapshot information. This unifies data from different sources and in different formats into a standardized set of thread operation snapshot information. Then, service operation anomaly detection processing is performed on the aforementioned thread operation snapshot information set to obtain at least one abnormal thread information. Each abnormal thread information includes a function code, a user code, business information, and an anomaly type identifier. This allows for the location of the fine-grained object that triggered the anomaly (specifically, which function code or user code's request it is) and the differentiation of the anomaly type (e.g., crash or slowness). Then, for each of the at least one abnormal thread information, the following hardware resource release processing is performed: First, in response to determining that the anomaly type identifier included in the aforementioned abnormal thread information is a crash identifier, the target service process is restarted, and a crash root cause circuit breaker or isolation operation corresponding to the function code or user code included in the aforementioned abnormal thread information is executed to limit the hardware resources occupied by subsequent requests corresponding to the function code or user code included in the aforementioned abnormal thread information. The second step involves determining that the exception type identifier included in the aforementioned abnormal thread information is a slow identifier, and then executing rate limiting or isolation operations corresponding to the function code or user code included in the abnormal thread information. This allows for circuit breaking, isolation, or rate limiting only for subsequent requests that actually caused the exception, while leaving normal requests unrestricted. Because only subsequent requests with the abnormal function code or user code are restricted, the problem of wrongly reclaiming normal business resources caused by full-scale rate limiting and circuit breaking is avoided. Furthermore, since it eliminates the need to stop all threads, clear the queue, and subsequently rebuild threads, frequent thread reclamation and rebuilding operations are avoided, reducing the waste of computing and memory resources.

[0062] Further reference Figure 2 As an implementation of the methods shown in the figures, this disclosure provides some embodiments of a hardware resource release device for a trading system. These device embodiments are similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0063] like Figure 2 As shown, a hardware resource release device 200 for a trading system in some embodiments includes: a collection unit 201, a field semantic alignment mapping unit 202, an anomaly detection unit 203, and a hardware resource release unit 204. The collection unit 201 is configured to collect multi-source service operation telemetry data of a target service process in a preset trading system; the field semantic alignment mapping unit 202 is configured to perform field semantic alignment mapping processing on the multi-source service operation telemetry data to obtain a thread operation snapshot information set; the anomaly detection unit 203 is configured to perform service operation anomaly detection processing on the thread operation snapshot information set to obtain at least one abnormal thread information, wherein each abnormal thread information includes a function code, a user code, business information, and an anomaly type identifier; the hardware resource release unit 204 is configured to release each of the at least one abnormal thread information... Upon receiving abnormal thread information, the following hardware resource release procedures are performed: In response to determining that the exception type identifier included in the abnormal thread information is a crash identifier, the target service process is restarted, and a circuit breaker or isolation operation corresponding to the function code or user code included in the abnormal thread information is executed to limit the hardware resources occupied by subsequent requests corresponding to the function code or user code included in the abnormal thread information; In response to determining that the exception type identifier included in the abnormal thread information is a slow identifier, a rate limiting or isolation operation corresponding to the function code or user code included in the abnormal thread information is executed.

[0064] It is understandable that the units described in the device 200 are related to the reference. Figure 1 The steps in the method described above correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units contained therein, and will not be repeated here.

[0065] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0066] like Figure 3As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0067] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.

[0068] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.

[0069] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0070] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0071] Computer-readable media may be contained within an electronic device or may exist independently of the electronic device. A computer-readable medium carries one or more programs. When one or more programs are executed by the electronic device, the electronic device causes the following actions: It collects multi-source service operation telemetry data of a target service process in a preset transaction system; performs field semantic alignment mapping on the multi-source service operation telemetry data to obtain a set of thread operation snapshot information; performs service operation anomaly detection processing on the set of thread operation snapshot information to obtain at least one abnormal thread information, wherein each abnormal thread information includes a function code, a user code, business information, and an anomaly type identifier; for each of the at least one abnormal thread information, it performs the following hardware resource release processing: In response to determining that the anomaly type identifier included in the abnormal thread information is a crash identifier, it restarts the target service process and executes a crash root cause circuit breaker operation or isolation operation corresponding to the function code or user code included in the abnormal thread information to limit the hardware resources occupied by subsequent requests corresponding to the function code or user code included in the abnormal thread information; In response to determining that the anomaly type identifier included in the abnormal thread information is a slow identifier, it executes a rate limiting operation or isolation operation corresponding to the function code or user code included in the abnormal thread information.

[0072] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0073] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0074] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a data acquisition unit, a field semantic alignment mapping unit, an anomaly detection unit, and a hardware resource release unit. The names of these units do not necessarily limit the specific unit; for example, a data acquisition unit may also be described as "a unit that acquires multi-source service operation telemetry data of a target service process in a preset transaction system."

[0075] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0076] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of technical features, but should also cover other technical solutions formed by arbitrary combinations of technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A method for releasing hardware resources in a trading system, characterized in that: Collect multi-source service operation telemetry data of target service processes in a preset transaction system; perform field semantic alignment mapping processing on the multi-source service operation telemetry data to obtain a set of thread operation snapshot information; The thread runtime snapshot information set is subjected to service runtime anomaly detection processing to obtain at least one abnormal thread information, wherein each abnormal thread information includes a function code, a user code, business information and an anomaly type identifier; For each of the at least one abnormal thread information, perform the following hardware resource release process: In response to determining that the exception type identifier included in the abnormal thread information is a crash identifier, the target service process is restarted, and a crash root cause circuit breaker operation or isolation operation corresponding to the function code or user code included in the abnormal thread information is executed to limit the hardware resources occupied by subsequent requests corresponding to the function code or user code included in the abnormal thread information. This includes: obtaining priority information corresponding to the function code or user code; based on the priority information, in response to determining that the function code or user code meets a preset isolation condition, generating isolation control information for the function code or user code; based on the priority information, in response to determining that the function code or user code does not meet the preset isolation condition, generating crash root cause circuit breaker control information for the function code or user code; and sending the generated isolation control information or crash root cause circuit breaker control information to the local execution engine of the target service process, so that the local execution engine executes a crash root cause circuit breaker operation or isolation operation on the matched subsequent requests corresponding to the function code or user code to release at least one of the following hardware resources occupied by the subsequent requests: processor resources and memory resources. In response to determining that the exception type identifier included in the exception thread information is a slow identifier, a rate limiting operation or isolation operation corresponding to the function code or user code included in the exception thread information is executed. In response to determining that the exception type identifier included in the abnormal thread information is a dead identifier, a dead detection circuit breaker operation is performed based on the function code and business information included in the abnormal thread information; The method further includes: during the effective period of the circuit breaker operation or isolation operation for the root cause of the crash, continuously collecting the average time of function codes, the queue depth of the queue to be processed and the thread busyness of the target service process at a preset period, and performing fallback verification processing on the collected average time of function codes, the queue depth and the thread busyness to obtain fallback verification information; when the corresponding fallback verification information for a preset number of consecutive periods all indicate a fallback to below the standard value, a downgrade process is performed.

2. The method according to claim 1, wherein, The degradation process is characterized in that: For each of the root cause circuit breakers in all the root cause circuit breakers, replace the root cause circuit breaker with a rate limiting operation with the same function code or user code. For each of all isolation operations, replace the isolation operation with a rate-limiting operation of the same function code or user code.

3. The method according to claim 1, wherein, The step of performing fallback verification processing on the average time consumed by the collected function codes, the queue depth, and the thread busyness to obtain fallback verification information is characterized in that: In response to determining that the average time of the function code is less than or equal to the preset average time of the function code, the queue depth is less than or equal to the preset queue depth, and the thread busyness is less than or equal to the preset busyness, the information indicating a fallback to below the standard value is determined as fallback verification information.

4. The method according to claim 1, wherein, The method of performing a deadlock detection and circuit breaker operation based on the abnormal thread information, including function codes and business information, is characterized in that: The function code and business information are concatenated into a string or structure and used as the target object for deadlock detection. Record the current time and control the local execution engine to use the stuck detection target object as a matching condition, and refuse to process subsequent requests with the same function code and the same business information; After a preset time period following the current time, when the local execution engine receives a request with the same function code and the same business information, it identifies the request with the same function code and the same business information as a probe request, processes the probe request normally, and records the processing time of the probe request. In response to determining that the processing time of the probe request is less than a preset recovery threshold, subsequent requests with the same function code and the same business information are allowed to be processed normally.

5. A hardware resource release device for a transaction system, characterized in that: The acquisition unit is configured to acquire multi-source service operation telemetry data of the target service process in the preset transaction system; The field semantic alignment mapping unit is configured to perform field semantic alignment mapping processing on the multi-source service runtime telemetry data to obtain a thread runtime snapshot information set; An anomaly detection unit is configured to perform service operation anomaly detection processing on the thread operation snapshot information set to obtain at least one abnormal thread information, wherein each abnormal thread information includes a function code, a user code, business information, and an anomaly type identifier. The hardware resource release unit is configured to perform the following hardware resource release process for each of the at least one abnormal thread information: In response to determining that the exception type identifier included in the abnormal thread information is a crash identifier, the target service process is restarted, and a crash root cause circuit breaker operation or isolation operation corresponding to the function code or user code included in the abnormal thread information is executed to limit the hardware resources occupied by subsequent requests corresponding to the function code or user code included in the abnormal thread information. This includes: obtaining priority information corresponding to the function code or user code; based on the priority information, in response to determining that the function code or user code meets a preset isolation condition, generating isolation control information for the function code or user code; based on the priority information, in response to determining that the function code or user code does not meet the preset isolation condition, generating crash root cause circuit breaker control information for the function code or user code; and sending the generated isolation control information or crash root cause circuit breaker control information to the local execution engine of the target service process, so that the local execution engine executes a crash root cause circuit breaker operation or isolation operation on the matched subsequent requests corresponding to the function code or user code to release at least one of the following hardware resources occupied by the subsequent requests: processor resources and memory resources. In response to determining that the exception type identifier included in the exception thread information is a slow identifier, a rate limiting operation or isolation operation corresponding to the function code or user code included in the exception thread information is executed. In response to determining that the exception type identifier included in the abnormal thread information is a dead identifier, a dead detection circuit breaker operation is performed based on the function code and business information included in the abnormal thread information; The hardware resource release unit is further configured to continuously collect the average function code execution time, queue depth of the pending queue, and thread busyness of the target service process at a preset period during the effective period of the circuit breaker operation or isolation operation, and to perform fallback verification processing on the collected average function code execution time, queue depth, and thread busyness to obtain fallback verification information; when the corresponding fallback verification information for a consecutive preset number of periods all indicate a fallback to below the standard value, a downgrade process is performed.

6. An electronic device, characterized in that: One or more processors; A storage device on which one or more computer programs are stored; When the one or more computer programs are executed by the one or more processors, the one or more processors perform the method as described in any one of claims 1 to 4.

7. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • System anomaly detection method and device, storage medium and electronic equipment

    CN120029805A

  • Abnormality processing method and device, equipment, medium and product

    CN120631626A