Method, apparatus, and storage medium for monitoring operational status of a trading system

By acquiring multi-dimensional operational status information of the trading system, performing data cleaning and clustering, and using logistic regression algorithms to predict the health status of the trading system, the problem of difficulty in assessment and early warning in traditional solutions is solved, and accurate assessment and fault warning of the trading system are achieved.

CN114036027BActive Publication Date: 2026-07-21IND CONSUMER FINANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IND CONSUMER FINANCE CO LTD
Filing Date
2021-11-15
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Traditional methods for monitoring the operational status of trading systems are unable to accurately assess the health of the system and cannot provide early warnings of abnormal health conditions, making it impossible to avoid trading failures.

Method used

By acquiring multi-dimensional operational status information of the host, network devices, and applications of the trading system, data cleaning and clustering are performed to generate monitoring event data. Logistic regression algorithms are used to predict the health status of the trading system, generate transaction approval status representation data, and issue early warnings based on this data.

Benefits of technology

It enables accurate assessment and early warning of the health status of the trading system, which can prevent trading failures in advance and improve the accuracy and foresight of the system health status judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114036027B_ABST
    Figure CN114036027B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, a computing device and a computer storage medium for monitoring a running state of a transaction system. The method comprises: obtaining first running state information about a plurality of hosts, second running state information about a plurality of network devices and third running state information about an application running on the transaction system; generating monitoring event data about the running state of the transaction system; determining event attribute information of the monitoring event data; clustering the monitoring event data based on the event attribute information of the monitoring event data so as to determine monitoring event representation data corresponding to each event attribute information; generating transaction approval state representation data within a predetermined time interval; and predicting a health state of the transaction system based on the monitoring event representation data corresponding to each event attribute information and the transaction approval state representation data. The present disclosure can accurately evaluate the health state of the transaction system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to information processing, and more specifically to methods, computing devices, and computer storage media for monitoring the operational status of a trading system. Background Technology

[0002] For trading systems, especially centralized systems that need to rapidly process a large number of concurrent requests, the requirements for accuracy and stability are extremely high. This is because the normal operation of almost all transactions depends on the normal functioning of the trading system's hardware and software. Therefore, real-time monitoring of the trading system's health status, accurate assessment of its health status, and early warning or prediction of abnormal health conditions are crucial to preventing failures that could disrupt normal trading.

[0003] Traditional methods for monitoring the operational status of trading systems involve using dedicated application software (such as network management software) to monitor the trading system network and detect anomalies in real time. This assists network administrators in troubleshooting and resolving faults, thereby ensuring the normal operation of the trading system. However, the anomalies detected by dedicated application software are usually retrospective and only involve failures in the trading system's basic hardware or network links. Therefore, it is difficult to accurately assess the health status of the trading system and cannot provide early warnings for abnormal health conditions, thus failing to prevent failures that could impact trading.

[0004] In summary, the shortcomings of traditional solutions for monitoring the operational status of trading systems are: difficulty in accurately assessing the health status of the trading system and inability to provide early warnings for abnormal health conditions of the trading system. Summary of the Invention

[0005] This disclosure provides a method, computing device, and computer storage medium for monitoring the operating status of a system, which can accurately assess the health status of a trading system and provide early warnings for abnormal health statuses of the trading system.

[0006] According to a first aspect of this disclosure, a method for monitoring the operational status of a system is provided. The method includes: acquiring first operational status information about multiple hosts, second operational status information about multiple network devices, and third operational status information about applications running on a transaction system within a predetermined time interval. The transaction system includes at least multiple hosts and multiple network devices. The first operational status information includes at least device status information, operating system status information, and database status information of the hosts. The method also includes: performing data cleaning on the first, second, and third operational status information to generate monitoring event data regarding the operational status of the transaction system; determining event attribute information for the monitoring event data, the event attribute information including at least several categories such as notification category, alarm category, fault category, and production change category; clustering the monitoring event data based on the event attribute information to determine monitoring event representation data corresponding to each event attribute information; generating transaction approval status representation data within the predetermined time interval; and predicting the health status of the transaction system based on the monitoring event representation data corresponding to each event attribute information and the transaction approval status representation data.

[0007] According to a second aspect of the invention, a computing device is also provided, the device comprising: at least one processing unit; at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the computing device to perform the method of the first aspect of the present disclosure.

[0008] According to a third aspect of this disclosure, a computer-readable storage medium is also provided. This computer-readable storage medium stores a computer program that, when executed by a machine, performs the method of the first aspect of this disclosure.

[0009] In some embodiments, obtaining first operational status information about multiple hosts, second operational status information about multiple network devices, and third operational status information about applications running in the transaction system within a predetermined time interval includes: determining the third operational status information about applications running in the transaction system based on the response time of application interfaces, error codes and keywords output by application logs within the predetermined time interval.

[0010] In some embodiments, generating transaction approval status representation data within a predetermined time interval includes: generating transaction approval status representation data including: transaction approval quantity representation data, pending approval quantity representation data, approved quantity representation data, loan disbursement quantity representation data, and repayment quantity representation data based on the number of transaction approvals, the number of transactions pending approval, the number of approved transactions, and the number of repayments within the predetermined time interval.

[0011] In some embodiments, predicting the health status of a transaction system includes: based on a logistic regression algorithm, for each of multiple hosts, determining multiple first weights corresponding to various event attribute information and multiple second weights corresponding to transaction approval status representation data, wherein the multiple first weights are used to indicate the impact of various monitoring event representation data corresponding to various event attribute information on the health status of each host, and the second weights are used to indicate the impact of pending transaction quantity representation data, approved transaction quantity representation data, loan disbursement quantity representation data, and repayment quantity representation data on the health status of each host; and determining the health status of each host in the transaction system based on the monitoring event representation data corresponding to various event attribute information, the multiple first weights, the transaction approval status representation data, and the multiple second weights.

[0012] In some embodiments, generating monitoring event data about the operational status of the trading system includes: converging identical monitoring event records within consecutive sampling time intervals; filtering the current monitoring event record based on other monitoring event records associated with the current monitoring event record; and generating monitoring event data about the operational status of the trading system based on the converged and filtered monitoring event records.

[0013] In some embodiments, the method for monitoring the operating status of a transaction system further includes: sorting the health status of each determined host; determining that the health status of the current host meets predetermined conditions in response to the determination that the current host's health status is ranked before a first predetermined order; and determining that the health status of the current host does not meet predetermined conditions in response to the determination that the current host's health status is ranked after a second predetermined order.

[0014] In some embodiments, the method for monitoring the operating status of a transaction system further includes: determining whether the health status of the current host does not meet predetermined conditions; and generating a maintenance order for the current host in response to determining that the health status of the current host does not meet predetermined conditions.

[0015] In some embodiments, the method for monitoring the operating status of a transaction system further includes: operating status information about the host's central processing unit, memory, and ports, wherein the second operating status information at least indicates fault information, availability information, and operating performance information of network devices.

[0016] The summary section is provided to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or principal features of this disclosure, nor is it intended to limit the scope of this disclosure. Attached Figure Description

[0017] Figure 1A schematic diagram of a system for implementing a method for monitoring the operational status of a trading system, according to an embodiment of the present disclosure, is shown.

[0018] Figure 2 A flowchart of a method for monitoring system operating status according to an embodiment of the present disclosure is shown.

[0019] Figure 3 A flowchart illustrating a method for determining the health status of a trading system according to an embodiment of this disclosure is shown.

[0020] Figure 4 A flowchart illustrating a method for cleaning data for operational status information according to an embodiment of the present disclosure is shown.

[0021] Figure 5 A schematic diagram is shown illustrating a method for generating a maintenance order for a current host according to an embodiment of this disclosure.

[0022] Figure 6 A block diagram schematically illustrates an electronic device suitable for implementing embodiments of the present disclosure.

[0023] In the various figures, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation

[0024] Preferred embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.

[0025] The term "comprising" and its variations as used herein signify open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one example embodiment" and "one embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc., may refer to different or the same objects.

[0026] As described earlier, in traditional solutions for monitoring the operational status of trading systems, the anomalies detected by dedicated application software are usually retrospective and only involve failures in the trading system's basic hardware or network links. Therefore, it is difficult to accurately assess the health status of the trading system, provide early warnings for abnormal health conditions, and prevent failures that could impact trading from occurring in advance.

[0027] To at least partially address one or more of the aforementioned problems and other potential issues, exemplary embodiments of this disclosure propose a scheme for monitoring the operational status of a trading system. In this scheme, by acquiring multi-dimensional operational status information of the trading system's hosts, network devices, and running applications, and by performing data cleaning on this multi-dimensional operational status information, this disclosure not only comprehensively collects multi-dimensional operational information of the trading system but also ensures that the generated monitoring event data will not lead to misjudgments of the trading system's health status due to occasional factors such as network instability. Furthermore, by determining the event attribute information of the monitoring event data (including multiple categories such as notification, alarm, fault, and production change), and by clustering the monitoring event data to determine the representative data of the monitoring events, this disclosure can differentiate and comprehensively assess the impact of maintenance records on the system's health status from multiple dimensions of varying severity, thereby improving the foresight of fault event prediction and preventing trading failures from occurring in advance. Furthermore, by generating transaction approval status representation data within predetermined time intervals, and based on monitoring event representation data and transaction approval status representation data corresponding to each event attribute information, this disclosure predicts the health status of the transaction system. This disclosure can determine the health status of the transaction system based on the operational status data of the system's underlying technical environment and the status data of business transactions, thus significantly improving the accuracy of judgments regarding the system's health status and operational trends. Therefore, this disclosure can accurately assess the health status of the transaction system and provide early warnings for abnormal health conditions.

[0028] Figure 1 A schematic diagram of a system 100 for implementing a method for monitoring system operating status, according to an embodiment of the present disclosure, is shown. Figure 1 As shown, system 100 includes a computing device 110 and a trading system 130. Trading system 130 includes multiple hosts 134 and multiple network devices 132. Network devices 132 include, for example, firewalls 136, switches 138, routers 140, etc. Trading system 130 can, for example, interact with an external server (not shown) via network 150. Trading system 130 can also interact with computing device 110 via wired or wireless means.

[0029] Transaction system 130 is used, for example, to provide transaction services to users based on applications running on it. Transaction system 130 is, for example, but not limited to, a bank's business system, which can approve user transaction requests (such as loan requests).

[0030] The computing device 110 is used to monitor the operational status of the transaction system 130 in order to predict or determine the health status of the transaction system 130. Specifically, the computing device 110 can acquire first operational status information about multiple hosts 134, second operational status information about multiple network devices 132, and third operational status information about the applications running on the transaction system 130; and perform data cleaning and other processing on the first, second, and third operational status information to generate monitoring event data about the system's operational status. The computing device 110 can also determine the event attribute information of the monitoring event data; and based on the event attribute information, cluster the monitoring event data to determine the monitoring event representation data corresponding to each event attribute information. In addition, the computing device 110 can generate transaction approval status representation data within a predetermined time interval; and predict the health status of the transaction system based on the monitoring event representation data and transaction approval status representation data corresponding to each event attribute information. The computing device 110 includes, but is not limited to, server computers, multiprocessor systems, mainframe computers, and distributed computing environments including any of the above systems or devices. In some embodiments, the computing device 110 may have one or more processing units, including dedicated processing units such as image processing units (GPUs), field-programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs), as well as general-purpose processing units such as central processing units (CPUs). In some embodiments, the computing device 110 includes, for example, a running status information acquisition unit 112, a monitoring event data generation unit 114, an event attribute information determination unit 116, a monitoring event characterization data determination unit 118, a transaction approval status characterization data generation unit 120, and a transaction system health status determination unit 122.

[0031] Regarding the operation status information acquisition unit 112, it is used to acquire first operation status information about multiple hosts, second operation status information about multiple network devices, and third operation status information about the application running in the transaction system within a predetermined time interval. The transaction system includes at least multiple hosts and multiple network devices. The first operation status information includes at least the device status information, operating system status information, and database status information of the hosts.

[0032] The monitoring event data generation unit 114 is used to clean the first operating status information, the second operating status information and the third operating status information in order to generate monitoring event data about the operating status of the trading system.

[0033] Regarding the event attribute information determination unit 116, it is used to determine the event attribute information of the monitoring event data. The event attribute information includes at least several of the following: notification category, alarm category, fault category, and production change category.

[0034] Regarding the monitoring event representation data determination unit 118, it is used to cluster the monitoring event data based on the event attribute information of the monitoring event data in order to determine the monitoring event representation data corresponding to each event attribute information.

[0035] Regarding the transaction approval status representation data generation unit 12, it is used to generate transaction approval status representation data within a predetermined time interval.

[0036] Regarding the transaction system health status determination unit 122, it is used to predict the health status of the transaction system based on the monitoring event representation data and transaction approval status representation data corresponding to the attribute information of each event.

[0037] The following will combine Figure 2 A method 200 for monitoring system operating status according to embodiments of the present disclosure is described. Figure 2 A flowchart of a method 200 for monitoring system operating status according to an embodiment of the present disclosure is shown. It should be understood that method 200 can, for example, be implemented in... Figure 6 The described electronic device is executed at point 600. It can also be used in... Figure 1 The described computing device 110 performs the operation. It should be understood that method 200 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of this disclosure is not limited in this respect.

[0038] At step 202, computing device 110 acquires first operating status information about multiple hosts, second operating status information about multiple network devices, and third operating status information about the application running in the transaction system within a predetermined time interval. The transaction system includes at least multiple hosts and multiple network devices. The first operating status information includes at least the device status information, operating system status information, and database status information of the hosts.

[0039] Regarding the first operational status information, it indicates, for example, the status of the host's underlying hardware and software environment, and includes at least the device status information, operating system status information, and database status information for each host. Specifically, it includes, for example, operational status information regarding the operating system, central processing unit (CPU), memory, disk, peripherals, network, and database for each host. Examples include data on whether the operating system is functioning correctly (e.g., startup or crash), memory usage data, disk space usage data, disk read / write error data, CPU operation status data (e.g., whether application scheduling errors are caused by overload), network connectivity with other devices (e.g., other hosts), packet forwarding status data, and so on.

[0040] A method for obtaining first operating status information of multiple hosts may include, for example, configuring an agent module on each monitored host, the agent module requesting a predetermined list of monitoring items from the computing device 110, each agent module periodically collecting various data of its host based on the obtained list of monitoring items, and sending the collected data (e.g., monitoring event records) to the computing device 110 to generate first operating status information.

[0041] The second operational status information includes, for example, fault information, availability information, and operational performance information of various network devices (e.g., routers, switches, wireless network cards, firewalls). In some embodiments, the computing device 110 can also automatically discover network devices and monitor network devices based on SNMP and other protocols. This includes data on the normality of routing or forwarding paths of network devices, bandwidth utilization data, network packet loss data, etc.

[0042] The third operational status information includes, for example, information indicating the application's response time, application availability, and correctness.

[0043] A method for obtaining third runtime status information about the application running in the trading system includes, for example, determining third runtime status information about the application running in the trading system based on the response time of the application interface (e.g., the response time for a transaction request) within a predetermined time interval, error codes and keywords output by the application logs.

[0044] Application correctness information, for example, indicates whether the current application is running normally, and includes, for example, whether the process exists or whether there is a crash. A method for determining application correctness information includes, for example, the following: computing device 110 calculates the average response time (e.g., 100 milliseconds) of the application interface within a predetermined time interval (e.g., but not limited to the most recent 10 minutes); compares the calculated average response time with a predetermined response time threshold (e.g., 1 millisecond) to determine whether the application interface response time is normal based on the comparison result (e.g., if the average response time of 100 milliseconds is much greater than the predetermined response time threshold of 1 millisecond, therefore, the application interface response time is determined to be abnormal); calculates the proportion of applications with normal response times in a predetermined number of iterations; and determines the application correctness information based on the calculated proportion.

[0045] Application availability information includes data that indicates the current running status of the application. This availability information may include, for example, application middleware status data (i.e., data indicating whether application components are functioning correctly, such as components involved in forwarding, storage, and the application runtime environment), data indicating whether the application is running correctly (e.g., whether processes exist or whether there is a crash), and error status data output from application logs.

[0046] Error status data output by the application log for transaction requests is determined based on error codes and keywords in the application log output. The application log output for transaction interactions can indicate whether the interaction is normal or not. For example, if the transaction interaction is normal, the corresponding application log output is a general information code; if the transaction interaction is abnormal, the corresponding application log output is an error code. Error codes can also indicate the severity of the application abnormality. For example, computing device 110 counts the number of abnormal transaction interactions based on keywords; compares the number of abnormal transaction interactions with a first application abnormality threshold and a second application abnormality threshold; if the number of abnormal transaction interactions is greater than or equal to the first application abnormality threshold and less than the second application abnormality threshold, a first error code indicating a lower-level abnormality is generated; if the number of abnormal transaction interactions is greater than or equal to the second application abnormality threshold, a second error code indicating a higher-level fault is generated.

[0047] In step 204, the computing device 110 performs data cleaning on the first operating status information, the second operating status information, and the third operating status information to generate monitoring event data regarding the operating status of the trading system. By employing the above methods, the generated monitoring event data regarding the operating status of the trading system can be prevented from causing misjudgments about the health status of the trading system due to status data generated by occasional factors.

[0048] Methods for cleaning operational status information include, for example, the following: the computing device 110 converges identical monitoring event records within consecutive sampling time intervals (it should be understood that identical monitoring event records refer to records indicating the same monitoring event and associated host, network device, or application); filters the current monitoring event record based on other monitoring event records associated with it; and generates monitoring event data about the operational status of the trading system based on the converged and filtered monitoring event records. The following will combine... Figure 4 The detailed method 400 for cleaning data for operational status information will not be elaborated here.

[0049] In step 206, the computing device 110 determines the event attribute information of the monitored event data. The event attribute information includes at least several of the following: notification category, alarm category, fault category, and production change category. By distinguishing and monitoring notification category, alarm category, fault category, and production change category, this disclosure can identify and address monitored events in their early stages, preventing serious faults that could affect transactions.

[0050] A method for determining event attribute information of monitoring event data includes, for example, the computing device 110 comparing the monitoring event data with predefined set values ​​associated with each monitoring event to determine the event attribute information of the monitoring event data based on the comparison results. Notification categories indicate monitoring events that do not require attention. Alarm categories indicate monitoring events that have not yet affected or only slightly affected transaction processing but require attention. Fault categories indicate monitoring events that have already affected transaction processing and require maintenance. Production change categories indicate monitoring events related to software changes or maintenance. For example, monitoring events involving adjustments to operating system data or parameters (e.g., the number of processes) or updates to application packages.

[0051] For example, if the host's disk space utilization is below 80%, the monitoring event for disk space utilization falls under the notification category. If the host's disk space utilization is above 80% but below 90%, the monitoring event for disk space utilization falls under the alarm category. If the host's disk space utilization is above 90%, the monitoring event for disk space utilization falls under the fault category.

[0052] In step 208, the computing device 110 clusters the monitoring event data based on the event attribute information to determine the monitoring event representation data corresponding to each event attribute information. By generating individual monitoring event representation data corresponding to different event attribute information through clustering of the monitoring event data, this disclosure can efficiently represent different degrees of monitoring event concentration without the need for complex processing of a large number of monitoring event records. The quantitative representation data includes, for example, monitoring event representation data corresponding to the event attribute information of notification category, alarm category, fault category, and production change category.

[0053] A method for determining the monitoring event representation data corresponding to each event attribute information includes, for example, sorting the number of monitoring events associated with each host corresponding to each event attribute information, so as to generate monitoring event representation data for each device corresponding to each event attribute information based on the sorting results.

[0054] The following uses formula (1) as an example to illustrate a method for determining the monitoring event representation data corresponding to the notification category event attribute information of the xth host.

[0055]

[0056] In formula (1) above, N represents the number of hosts. G1(x) represents the monitoring event representation data corresponding to the event attribute information of notification category for the x-th host. M1 represents the order in which the x-th host is ranked according to the order of the monitoring event representation data corresponding to the event attribute information of notification category for the N hosts from low to high.

[0057] It should be understood that, based on a method similar to the above formula (1), monitoring event representation data corresponding to the event attribute information of alarm category, monitoring event representation data corresponding to the event attribute information of fault category, and monitoring event representation data corresponding to the event attribute information of production change category can be determined for each host.

[0058] In step 210, computing device 110 generates transaction approval status representation data for a predetermined time interval.

[0059] The method for generating transaction approval status representation data includes, for example, the following: the computing device 110 generates transaction approval status representation data including: transaction approval quantity representation data, pending approval quantity representation data, approval quantity representation data, loan disbursement quantity representation data, and loan disbursement quantity representation data based on the number of transaction approvals, the number of pending approvals, the number of approved approvals, the number of loan disbursements, and the number of loan repayments within a predetermined time interval.

[0060] The calculation method for the number of transactions pending approval can be based on the following formula (2).

[0061]

[0062] In the above formula (2), This represents the number of transactions approved by the i-th host for a single transaction object k within each predetermined time interval. This represents the total number of approvals granted by the i-th host for a single transaction object k within each predetermined time interval. This represents the number of loans disbursed by the i-th host for a single transaction object k within each predetermined time interval. This represents the number of approval rejections made by the i-th host for a single transaction object k within each predetermined time interval. This represents the total number of pending transactions for all transaction objects within each predetermined time interval. n represents the total number of all transaction objects.

[0063] The following example uses the number of transaction approvals to illustrate how to generate transaction approval status data.

[0064] The calculation method for the transaction approval quantity representation data of the xth host includes, for example, calculating the average of the transaction approval quantity representation data of the xth host in the current predetermined time interval and the previous predetermined time interval for a single transaction object and all transaction objects, respectively, so as to generate the transaction approval quantity representation data of the xth host.

[0065] The transaction approval quantity representation data of the trading system is calculated based on the transaction approval quantity representation data of each host and the corresponding contribution of each host.

[0066] The following example, using formula (3), illustrates a method for generating transaction approval quantity representation data for the xth device.

[0067]

[0068] In formula (3) above, 'a' represents the variable representing the number of transactions approved. θ j The function weight factor represents the j-th element (e.g., a training sample element or a test element) under the i-th feature vector, i.e., the model parameter. The number of transactions approved represents the data for the i-th feature vector and the j-th element. a (x) represents the transaction approval count data of m elements in the training set (or test set). Let x represent the data model function that characterizes the number of transactions approved, which is the i-th feature vector and the j-th element of the training set (or test set) X, under the influence of the variable 'a' representing the number of transactions approved. j (i) This represents the j-th element of the i-th feature vector in the training (or test) set X. m represents the number of training sample elements (or test elements) in the training (or test) set. n represents the number of feature vectors.

[0069] Transaction approval quantity representation data model function The loss function J(θ) is processed, for example, but not limited to, based on the following formula (4).

[0070]

[0071] In the above formula (4), y (i) h represents the true value of the i-th feature vector in the transaction approval quantity representation data model function.θ (x (i) ) represents the predicted value of the x-th feature vector in the transaction approval quantity representation data model function. m represents the number of training sample elements (or test elements) in the training set (or test set).

[0072] It should be understood that, using a method similar to that used to calculate the transaction approval count for device x, the pending transaction count, approved transaction count, loan disbursement count, and repayment count for device x can be calculated. The following explanation, in conjunction with formulas (5) to (8), illustrates the calculation method for generating the pending transaction count, approved transaction count, loan disbursement count, and repayment count for device x.

[0073]

[0074]

[0075]

[0076]

[0077] In formulas (5) to (8) above, b represents the number of transactions pending approval. c represents the number of transactions pending approval. d represents the number of loans disbursed. e represents the number of repayments made. This represents the number of pending transactions, where the i-th feature vector and the j-th element represent the data. b (x) represents the number of pending transactions with m elements. This represents the data model function representing the number of pending transactions in the i-th feature vector and the j-th element under the influence of the variable b, which is the number of pending transactions. The number of approvals represented by the i-th feature vector and the j-th element is used to characterize the data. c (x) represents the number of approved data for m elements. This represents the data model function that characterizes the number of approvals for the i-th feature vector and the j-th element of the training set (or test set) X, under the influence of the factor variable c, which represents the number of approvals. The data represents the number of loan disbursements for the i-th feature vector and the j-th element. d (x) represents the number of loan disbursements for m elements. This represents the data model function that describes the number of loan disbursements for the i-th feature vector and the j-th element of the training set (or test set) X, under the influence of the variable d, which is the number of loan disbursements. The data represents the number of repayments for the i-th feature vector and the j-th element.e (x) represents the number of repayments for m elements. This represents the data model function that characterizes the number of repayments for the i-th feature vector and the j-th element of the training (or test) set X, under the influence of the variable d (number of repayments). θ j x represents the function weight factor of the j-th element under the i-th feature vector, i.e., the model parameter. j (i) This represents the j-th element of the i-th feature vector in the training set (or test set) X. m represents the number of elements in the training set (or test set). n represents the number of feature vectors.

[0078] In step 212, the computing device 110 predicts the health status of the transaction system based on the monitoring event representation data and transaction approval status representation data corresponding to each event attribute information.

[0079] Methods for determining the health status of a transaction system include: Computing device 110, based on a logistic regression algorithm, determines multiple first weights corresponding to various event attribute information and multiple second weights corresponding to transaction approval status representation data for each of the multiple hosts. The multiple first weights are used to indicate the impact of various monitoring event representation data corresponding to various event attribute information on the health status of each host. The second weights are used to indicate the impact of pending transaction quantity representation data, approved transaction quantity representation data, loan disbursement quantity representation data, and repayment quantity representation data on the health status of each host. Furthermore, based on the monitoring event representation data corresponding to various event attribute information, the multiple first weights, the transaction approval status representation data, and the multiple second weights, the health status of each host in the transaction system is determined. The following will combine... Figure 3 The specific methods for determining the health status of a trading system are explained in section 300, and will not be repeated here.

[0080] This solution acquires multi-dimensional operational status information of the trading system's hosts, network devices, and running applications, and performs data cleaning on this multi-dimensional operational status information. This disclosure not only enables comprehensive multi-dimensional collection of the trading system's operational information but also ensures that the generated monitoring event data will not lead to misjudgments of the trading system's health status due to occasional factors such as network instability. Furthermore, by determining the event attribute information of the monitoring event data (including multiple categories such as notification, alarm, fault, and production change), and by clustering the monitoring event data to determine the representative data of the monitoring events, this disclosure can differentiate and comprehensively assess the impact of maintenance records on the system's health status from multiple dimensions of varying severity. This improves the foresight of fault event prediction and helps prevent trading failures from occurring in advance. Furthermore, by generating transaction approval status representation data within predetermined time intervals, and based on monitoring event representation data and transaction approval status representation data corresponding to each event attribute information, this disclosure predicts the health status of the transaction system. This disclosure can determine the health status of the transaction system based on the operational status data of the system's underlying technical environment and the status data of business transactions, thus significantly improving the accuracy of judgments regarding the system's health status and operational trends. Therefore, this disclosure can accurately assess the health status of the transaction system and provide early warnings for abnormal health conditions.

[0081] The following will combine Figure 3 A method 300 for determining the health status of a transaction system according to embodiments of the present disclosure is described. Figure 3 A flowchart of a method 300 for determining the health status of a trading system according to an embodiment of the present disclosure is shown. It should be understood that method 300 can, for example, be implemented in... Figure 6 The described electronic device is executed at point 600. It can also be used in... Figure 1 The described computing device 110 performs the operation. It should be understood that method 300 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of this disclosure is not limited in this respect.

[0082] In step 302, the computing device 110, based on the logistic regression algorithm, determines multiple first weights corresponding to each event attribute information and multiple second weights corresponding to the transaction approval status representation data for each of the multiple hosts. The multiple first weights are used to indicate the impact of each monitoring event representation data corresponding to each event attribute information on the health status of each host. The second weights are used to indicate the impact of the pending transaction quantity representation data, the approved quantity representation data, the loan disbursement quantity representation data, and the repayment quantity representation data on the health status of each host.

[0083] In step 304, the computing device 110 determines the health status of each host in the transaction system based on the monitoring event representation data corresponding to each event attribute information, multiple first weights, transaction approval status representation data, and multiple second weights.

[0084] The following describes the method for determining the health status of each host in conjunction with formula (9).

[0085]

[0086] In formula (8) above, J1 represents the weight corresponding to the event attribute information of notification category. J2 represents the weight corresponding to the event attribute information of alarm category. J3 represents the weight corresponding to the event attribute information of fault category. J4 represents the weight corresponding to the event attribute information of production change category. The first weight includes J1, J2, J3 and J4. a The weights corresponding to the number of transaction approvals represent the data. b The weights represent the data corresponding to the number of transactions pending approval. c The number of approvals represents the weight of the data. d This represents the weight corresponding to the number of loan disbursements. e The first weight represents the weight corresponding to the number of repayments. The second weight includes J. a J b J c J d and J e Y represents the health status of the x-th host. G1(x) represents the monitoring event representation data corresponding to the event attribute information of notification category. G2(x) represents the monitoring event representation data corresponding to the event attribute information of alarm category. G3(x) represents the monitoring event representation data corresponding to the event attribute information of fault category. G4(x) represents the monitoring event representation data corresponding to the event attribute information of production change category.

[0087] The following will combine Figure 4 A method 400 for cleaning data for operational status information according to embodiments of the present disclosure is described. Figure 4 A flowchart of a method 400 for data cleaning of runtime status information according to an embodiment of the present disclosure is shown. It should be understood that method 400 can, for example, be used in... Figure 6 The described electronic device is executed at point 600. It can also be used in... Figure 1 The described computing device 110 performs the operation. It should be understood that method 400 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of this disclosure is not limited in this respect.

[0088] At step 402, computing device 110 converges the records of the same monitoring events within consecutive sampling time intervals.

[0089] For example, computing device 110 determines that there are identical first and second records within consecutive sampling time intervals. The first record is, for example, a monitoring event record indicating that the host disk utilization was too high (e.g., the host disk utilization data is 80%, higher than the utilization threshold) in the previous sampling time interval (e.g., 5 minutes); the second record is, for example, a monitoring event record indicating that the host disk utilization data is still 80% in the next sampling time interval, indicating that the host disk utilization was too high; then computing device 110 converges the first record and the second record into a single monitoring event record regarding the host disk utilization being too high.

[0090] At step 404, computing device 110 filters the current monitoring event record based on other monitoring event records associated with the current monitoring event record.

[0091] For example, if computing device 110 determines that the third record is a monitoring event record indicating that the process of an application on a certain host crashes at a certain moment; computing device 110 also determines that the monitoring event record indicating the traffic transmission between the host where the application resides and other hosts is normal, and the monitoring event record indicating the number of approved transactions corresponding to the application is normal; and the monitoring event record indicating the response time of the application interface is also normal, then computing device 110 determines that the third record is invalid and thus filters out the third record.

[0092] At step 406, computing device 110 generates monitoring event data about the system's operating status based on converged and filtered monitoring event records.

[0093] By employing the above methods, this disclosure can more effectively prevent the generated monitoring event data on system operating status from causing misjudgments about the health status of the trading system due to alarm data generated by occasional factors such as network instability.

[0094] The following will combine Figure 5 A method 500 for generating maintenance orders for a current host is described according to embodiments of the present disclosure. Figure 5 A flowchart of a method 500 for generating a maintenance order for a current host, according to an embodiment of the present disclosure, is shown. It should be understood that method 500 can, for example, be implemented in... Figure 6 The described electronic device is executed at point 600. It can also be used in... Figure 1 The described computing device 110 performs the operation. It should be understood that method 500 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of this disclosure is not limited in this respect.

[0095] At step 502, computing device 110 sorts the health status of each determined host.

[0096] At step 504, computing device 110 determines whether the current host's health status is ranked before a first predetermined order. In some embodiments, the first predetermined order corresponds, for example, to the 20th percentile of all hosts.

[0097] At step 506, if the computing device 110 determines that the current host's health status is ranked before the first predetermined order, it determines that the current host's health status meets the predetermined conditions.

[0098] At step 508, if the computing device 110 determines that the current host's health status is not ranked before a first predetermined order, it determines whether the current host's health status is ranked after a second predetermined order. In some embodiments, the second predetermined order corresponds, for example, to the 60th percentile of the ranking of all hosts.

[0099] At step 510, if the computing device 110 determines that the current host's health status is ranked after the second predetermined order, it determines that the current host's health status does not meet the predetermined conditions. If the computing device 110 determines that the current host's health status is not ranked after the second predetermined order, it jumps to step 512 and determines that the current host's health status needs attention.

[0100] In step 514, computing device 110 determines whether the current host's health status does not meet predetermined conditions. If it is determined that the current host's health status does not fail to meet predetermined conditions, the process proceeds to step 518, where corresponding indication information is presented based on the current host's health status. For example, indication information indicating good health is presented for the top 20% of the health status rankings. Indication information indicating that attention is needed is presented for the 20% to 40% of the health status rankings.

[0101] In step 516, if the computing device 110 determines that the health status of the current host does not meet predetermined conditions, a maintenance order is generated for the current host. For example, a maintenance order is generated for the bottom 60% of hosts in terms of health status.

[0102] By employing the above methods, this disclosure can automatically handle maintenance orders for faulty hosts and present corresponding indication signals for hosts in a healthy state and hosts in a state requiring attention.

[0103] Figure 6 A block diagram schematically illustrates an electronic device (or computing device) 600 suitable for implementing embodiments of the present disclosure. Device 600 may be used to implement... Figures 2 to 5The devices shown are those for methods 200, 300, 400, and 500. (As shown...) Figure 6 As shown, device 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 602 or loaded from storage unit 608 into random access memory (RAM) 603. The RAM may also store various programs and data required for the operation of device 600. The CPU, ROM, and RAM are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0104] Multiple components in device 600 are connected to input / output (I / O) 605, including an input unit 606, an output unit 607, and a storage unit 608. A central processing unit 601 executes the various methods and processes described above, such as methods 200, 300, 400, and 500. For example, in some embodiments, methods 200, 300, 400, and 500 may be implemented as computer software programs stored in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed onto device 600 via ROM and / or communication unit 609. When the computer program is loaded into RAM and executed by the CPU, one or more operations of methods 200, 300, 400, and 500 described above may be performed. Alternatively, in other embodiments, the CPU may be configured to perform one or more actions of methods 200, 300, 400, and 500 by any other suitable means (e.g., by means of firmware).

[0105] It should be further noted that this disclosure can be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0106] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0107] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0108] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as C or similar languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0109] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or step diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each step in the flowchart illustrations and / or step diagrams, as well as combinations of steps in the flowchart illustrations and / or step diagrams, can be implemented by computer-readable program instructions.

[0110] These computer-readable program instructions can be provided to a processor in a voice interaction device, a general-purpose computer, a special-purpose computer, or a processing unit of another programmable data processing device, thereby producing a machine such that, when executed by the processing unit of the computer or other programmable data processing device, these instructions create means for implementing the functions / actions specified in one or more steps of the flowchart and / or diagram of steps. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing device, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more steps of the flowchart and / or diagram of steps.

[0111] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more steps of a flowchart and / or a diagram of steps.

[0112] The flowcharts and step diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each step in a flowchart or step diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions indicated in the step may occur in a different order than those indicated in the drawings. For example, two consecutive step diagrams may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each step in the step diagrams and / or flowcharts, and combinations of steps in the step diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0113] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0114] The above are merely optional embodiments of this disclosure and are not intended to limit this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for monitoring the operational status of a trading system, comprising: The system acquires first operating status information about multiple hosts, second operating status information about multiple network devices, and third operating status information about the applications running in the transaction system within a predetermined time interval. The transaction system includes at least the multiple hosts and the multiple network devices. The first operating status information includes at least the device status information, operating system status information, and database status information of the hosts. Data cleaning is performed on the first, second, and third operating status information to generate monitoring event data on the operating status of the trading system. The generation of monitoring event data regarding the operational status of the trading system includes: convergence of identical monitoring event records within consecutive sampling time intervals; filtering of the current monitoring event record based on other monitoring event records associated with it, wherein the filtering includes: determining the current monitoring event record as invalid and filtering it when the host traffic monitoring event record, transaction approval monitoring event record, and application interface response time monitoring event record associated with the current monitoring event record all indicate normal; Determine the event attribute information of the monitoring event data. The event attribute information includes at least several of the following: notification category, alarm category, fault category, and production change category. The notification category is used to indicate monitoring events that do not need attention. The alarm category indicates monitoring events that have not yet affected or have only affected a small amount of transaction processing but need attention. The fault category indicates monitoring events that have already affected transaction processing and require maintenance. The production change category indicates monitoring events that involve software changes or maintenance. Based on the event attribute information of the monitoring event data, clustering is performed on the monitoring event data to determine the monitoring event representation data corresponding to each event attribute information. The clustering includes: for each host, counting the number of monitoring events corresponding to each event attribute information, sorting the number of monitoring events corresponding to the same event attribute for each host, and using the sorting order of each host as the monitoring event representation data of that host under that event attribute. Generate transaction approval status representation data within a predetermined time interval; and Based on the monitoring event representation data and transaction approval status representation data corresponding to the attribute information of each event, the health status of the transaction system is predicted.

2. The method according to claim 1, wherein obtaining first operating status information about multiple hosts, second operating status information about multiple network devices, and third operating status information about the application running in the transaction system within a predetermined time interval includes: Based on the response time of the application interface within a predetermined time interval, the error codes and keywords output by the application log, the third runtime status information of the application running in the trading system is determined.

3. The method according to claim 1, wherein generating transaction approval status representation data within a predetermined time interval includes: Based on the number of transactions approved, the number of transactions pending approval, the number of approved transactions, the number of loans disbursed, and the number of repayments within a predetermined time interval, the transaction approval status representation data is generated, which includes: transaction approval quantity representation data, transaction pending approval quantity representation data, number of approved transactions representation data, number of loans disbursed, and number of repayments representation data.

4. The method of claim 3, wherein predicting the health status of the trading system comprises: Based on the logistic regression algorithm, for each of the multiple hosts, multiple first weights corresponding to each event attribute information and multiple second weights corresponding to the transaction approval status representation data are determined. The multiple first weights are used to indicate the impact of each monitoring event representation data corresponding to each event attribute information on the health status of each host. The second weights are used to indicate the impact of pending transaction quantity representation data, approved transaction quantity representation data, loan disbursement quantity representation data, and repayment quantity representation data on the health status of each host. Based on the monitoring event representation data corresponding to each event attribute information, multiple first weights, transaction approval status representation data, and multiple second weights, the health status of each host in the transaction system is determined.

5. The method according to claim 1, wherein generating monitoring event data regarding the operational status of the trading system further includes: Based on converged and filtered monitoring event records, monitoring event data about the operational status of the trading system is generated.

6. The method according to claim 5, further comprising: Sort the health status of each host according to its condition. Determine whether the current host's health status is prioritized before the first predetermined order; In response to the order in which the health status of the current host is determined to be prior to the first predetermined order, the health status of the current host is determined to meet the predetermined conditions; In response to determining that the current host's health status is not ranked before the first predetermined order, determine whether the current host's health status is ranked after the second predetermined order; as well as In response to the fact that the order in which the health status of the current host is determined is after the second predetermined order, it is determined that the health status of the current host does not meet the predetermined conditions.

7. The method of claim 6, further comprising: Determine if the current host's health status does not meet the predetermined conditions; as well as In response to the determination that the health status of the current host does not meet the predetermined conditions, a maintenance order for the current host is generated.

8. The method according to claim 1, wherein the first operating status information further includes at least: Regarding the operating status information of the host's central processing unit, memory, and ports, the second operating status information at least indicates the network device's fault information, availability information, and operating performance information.

9. A computing device, comprising: At least one processing unit; At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the computing device to perform the method of any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, which, when executed by a machine, implements the method of any one of claims 1 to 8.