Node abnormal skipping detection method and device, equipment, medium and product

By building an abnormal skip detection model based on a decision tree and utilizing the recent skipping status and date information of the node, abnormal skipping can be automatically identified, solving the problems of low efficiency and high cost of manual judgment and achieving efficient and accurate node anomaly detection.

CN120602374APending Publication Date: 2025-09-05AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510968706.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In the existing technology, the identification of abnormal node skipping mainly relies on manual judgment, resulting in low processing efficiency and accuracy and high labor costs.

Method used

By building an abnormal skip detection model and using machine learning technology to analyze the recent skipping conditions and date information of nodes, abnormal skipping can be automatically identified. The model training is based on the decision tree algorithm and combines date and business attribute information to generate abnormal detection results.

Benefits of technology

It improves the recognition accuracy and processing efficiency of abnormal node skipping, reduces labor costs, and realizes automated anomaly detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602374A_ABST
    Figure CN120602374A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a node exception skipping detection method and device, equipment, a medium and a product. The method comprises the following steps: when a skipped node has a corresponding abnormal skipping detection model, acquiring node skipping information of a target day of the skipped node and date information and service attribute information of the current day, and inputting the node skipping information of the target day and the date information and the service attribute information of the current day into the abnormal skipping detection model, and obtaining a first anomaly detection result output by the anomaly skipping detection model. Wherein the target day is before the current day, the time interval between the target day and the current day is smaller than a preset time interval, and the business attribute information of the current day is used for indicating whether business needing to be processed exists in the current day; and the first anomaly detection result is used for indicating whether the skipped node is abnormal. The method is used for achieving the technical effects of improving the detection efficiency and accuracy and reducing the labor cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, equipment, medium and product for detecting abnormal node skipping. Background Art

[0002] In batch job scheduling for production application systems, nodes in the job plan may need to be skipped due to the influence of upstream and downstream data or business operations. Therefore, job scheduling generally supports skipping nodes in a specific batch for a particular day. However, in real-world applications, manual configuration errors or scheduling system anomalies can cause nodes to be mistakenly skipped, leading to task execution errors or failures. Therefore, it is very necessary to identify whether nodes have been abnormally skipped.

[0003] Currently, determining whether a node has been abnormally skipped is primarily done manually. Specifically, when a node is skipped, expert experience is used to manually determine whether it was an abnormal skip. However, this manual judgment method results in low processing efficiency and accuracy, and high labor costs. Summary of the Invention

[0004] The embodiments of the present application provide a node abnormal skipping detection method, apparatus, equipment, medium and product to improve the processing efficiency and accuracy of determining whether a node is abnormally skipped and reduce labor costs.

[0005] In a first aspect, an embodiment of the present application provides a method for detecting abnormal node skipping, comprising:

[0006] When a corresponding abnormal skip detection model exists for a skip node, obtaining node skip information of a target day of the skip node and date information and business attribute information of the current day, where the target day is before the current day and the time interval between the target day and the current day is less than a preset time interval, and the business attribute information of the current day is used to indicate whether there is business that needs to be processed on the current day;

[0007] Inputting the node skipping information of the target day, the date information of the current day, and the service attribute information into the abnormal skipping detection model, and obtaining a first abnormality detection result output by the abnormal skipping detection model, wherein the first abnormality detection result is used to indicate whether the skipped node is an abnormal skip;

[0008] In which, the abnormal skip detection model is trained based on the node skip information of the first sample day of the sample node and the date information and business attribute information of the second sample day, the first sample day is before the second sample day, and the time interval between the first sample day and the second sample day is less than the preset time interval.

[0009] In a possible implementation, the date information of the current day includes at least one of the following: the current day date, a week attribute, and a holiday attribute.

[0010] In a possible implementation, when a corresponding abnormal skip detection model exists for the skip node, before obtaining the node skip information of the target day of the skip node and the date information and service attribute information of the day, the method further includes:

[0011] Acquire multiple training samples for each sample node, where the training samples include node skipping information of the first sample day and date information and service attribute information of the second sample day of the sample node;

[0012] If the data amount of the multiple training samples of the sample node is less than the preset training sample amount, calculating the historical skip probability of the sample node based on the historical skip information of the sample node;

[0013] Accordingly, when the skip node does not have a corresponding abnormal skip detection model, the method further includes:

[0014] Obtaining a historical skip probability corresponding to the skipped node;

[0015] A second abnormality detection result is output according to the historical skip probability and the preset skip probability, where the second abnormality detection result is used to indicate whether the skipped node is an abnormal skip.

[0016] In a possible implementation, outputting a second anomaly detection result according to the historical skip probability and the preset skip probability includes:

[0017] If the historical skip probability is greater than or equal to the preset skip probability, outputting a second abnormality detection result indicating that the skipped node is skipped normally;

[0018] If the historical skip probability is less than the preset skip probability, a second abnormality detection result indicating that the skipped node is abnormally skipped is output.

[0019] In one possible implementation, the method further includes:

[0020] If the data amount of the multiple training samples of the sample node is greater than or equal to the preset training sample amount, determining a label for each training sample, wherein the label of the training sample is used to indicate whether the sample node has skipped processing on the second sample day;

[0021] Model training is performed based on the multiple training samples and the label of each training sample to generate the abnormal skip detection model.

[0022] In a possible implementation, the abnormal skip detection model is a decision tree model.

[0023] In a second aspect, an embodiment of the present application provides a node abnormal skipping detection device, comprising:

[0024] an acquisition module, configured to, when a corresponding abnormal skip detection model exists for a skip node, acquire node skip information of a target day for the skip node, as well as date information and business attribute information of the current day, wherein the target day is before the current day and the time interval between the target day and the current day is less than a preset time interval, and the business attribute information of the current day is used to indicate whether there is any business that needs to be processed on the current day;

[0025] an input module, configured to input the node skipping information of the target day, the date information of the current day, and the service attribute information into the abnormal skip detection model, and obtain a first abnormality detection result output by the abnormal skip detection model, wherein the first abnormality detection result is used to indicate whether the skipped node is an abnormal skip;

[0026] In which, the abnormal skip detection model is trained based on the node skip information of the first sample day of the sample node and the date information and business attribute information of the second sample day, the first sample day is before the second sample day, and the time interval between the first sample day and the second sample day is less than the preset time interval.

[0027] In a possible implementation, the date information of the current day includes at least one of the following: the current day date, a week attribute, and a holiday attribute.

[0028] In a possible embodiment, the node abnormal skip detection device further includes a calculation module. When a corresponding abnormal skip detection model exists for the skipped node, before obtaining the node skip information of the target day of the skipped node and the date information and business attribute information of the day, the acquisition module is further used to obtain multiple training samples for each sample node, the training samples including the node skip information of the first sample day and the date information and business attribute information of the second sample day of the sample node;

[0029] The calculation module is configured to calculate a historical skip probability of the sample node based on historical skip information of the sample node if the data amount of the multiple training samples of the sample node is less than a preset training sample amount;

[0030] Correspondingly, when there is no corresponding abnormal skip detection model for the skip node, the acquisition module is further used to obtain the historical skip probability corresponding to the skip node;

[0031] The output module is further configured to output a second anomaly detection result based on the historical skip probability and the preset skip probability, where the second anomaly detection result is configured to indicate whether the skipped node is an abnormal skip.

[0032] In a possible implementation, the output module is specifically configured to:

[0033] If the historical skip probability is greater than or equal to the preset skip probability, outputting a second abnormality detection result indicating that the skipped node is skipped normally;

[0034] If the historical skip probability is less than the preset skip probability, a second abnormality detection result indicating that the skipped node is abnormally skipped is output.

[0035] In a possible implementation, the node abnormal skipping detection device further includes a training module, which is configured to:

[0036] If the data amount of the multiple training samples of the sample node is greater than or equal to the preset training sample amount, determining a label for each training sample, wherein the label of the training sample is used to indicate whether the sample node has skipped processing on the second sample day;

[0037] Model training is performed based on the multiple training samples and the label of each training sample to generate the abnormal skip detection model.

[0038] In a possible implementation, the abnormal skip detection model is a decision tree model.

[0039] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory, a processor;

[0040] The memory stores computer-executable instructions;

[0041] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.

[0042] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementation methods of the first aspect.

[0043] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.

[0044] The embodiments of the present application provide a node abnormal skip detection method, apparatus, device, medium and product, in which: when a corresponding abnormal skip detection model exists for a skipped node, the node skip information of the target day of the skipped node and the date information and business attribute information of the current day are obtained, the node skip information of the target day and the date information and business attribute information of the current day are input into the abnormal skip detection model, and the first abnormal detection result output by the abnormal skip detection model is obtained. Wherein, the target day is before the current day, and the time interval between the target day and the current day is less than the preset time interval, and the business attribute information of the current day is used to indicate whether there is any business that needs to be processed on the current day. The first abnormal detection result is used to indicate whether the skipped node is an abnormal skip, and the abnormal skip detection model is trained based on the node skip information of the first sample day of the sample node and the date information and business attribute information of the second sample day. The first sample day is before the second sample day, and the time interval between the first sample day and the second sample day is less than the preset time interval. In this technology, the abnormal skip detection model can determine whether it meets the pattern characteristics of normal skipping based on the node skipping information of the target day of the skipped node and the date information and business attribute information of the day, and then identify abnormal skipping that does not conform to conventional logic, and output the first abnormality detection result, replacing the manual processing process, solving the problem of low accuracy caused by the inability to guarantee rigor in manual processing, reducing labor costs and improving processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0046] Figure 1 Schematic diagram of the process of node abnormal skipping detection method provided by this application Figure 1 ;

[0047] Figure 2 Schematic diagram of the process of node abnormal skipping detection method provided by this application Figure 2 ;

[0048] Figure 3 A flowchart of the model training process for the node abnormal skipping detection method provided in this application;

[0049] Figure 4 A flowchart of the model reasoning process of the node abnormal skipping detection method provided by this application;

[0050] Figure 5 This is a schematic diagram of the structure of the node abnormal skipping detection device provided by this application;

[0051] Figure 6 This is a schematic diagram of the structure of the electronic device provided in this application.

[0052] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0053] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0054] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0055] First, the professional terms involved in this application are explained:

[0056] Decision tree model: A decision tree is a supervised machine learning algorithm with a tree structure that is widely used in classification, prediction, rule extraction and other fields. Each of its leaf nodes corresponds to a target value, and non-leaf nodes correspond to the division on a certain attribute. The samples are divided into several subsets according to their different values ​​on this attribute. The core issue of constructing a decision tree is how to select appropriate attributes to split the samples at each step. Learning and constructing a decision tree from training samples with known predicted values ​​is a top-down, divide-and-conquer process. Commonly used decision tree algorithms include the Iterative Dichotomiser 3 (ID3) algorithm, the C4.5 algorithm, and the Classification and Regression Trees (CART) algorithm.

[0057] Batch: In computer science, batch processing is a method of executing jobs that allows a system to automatically process large amounts of data or perform a series of predefined operations without human intervention. Batch processing generally has the following characteristics: 1. Non-interactive: Batch processing typically runs without direct user intervention; once a task begins, the system automatically completes all steps; 2. High throughput: Batch processing is suitable for processing large amounts of data; 3. Scheduled execution: Batch processing tasks are typically scheduled during periods of low system load (such as nighttime) to avoid disrupting daily business operations that require user interaction.

[0058] Batch job: A series of operations that are pre-arranged. A batch job may contain multiple batch nodes, and there may be execution order dependencies between the nodes.

[0059] Batch node: A batch node represents the atomic operation in a batch job. One batch node corresponds to one batch program.

[0060] Next, the application background involved in this application is explained:

[0061] In production application systems, data processing is primarily divided into two modes: online operations and batch processing. Online operations refer to real-time interactive operations with direct user participation, such as online transactions or information queries. Batch processing, on the other hand, involves large-scale data operations automatically executed by the system according to preset rules, often with periodic or triggered characteristics. Typical batch processing scenarios include scheduled tasks (such as daily bank transaction settlements and e-commerce inventory status updates), data processing (such as monthly bill generation and sales report calculations), and file operations (such as reconciliation file generation and log archiving). These batch processing processes serve as backend support processes, ensuring the real-time requirements of online operations while maintaining system data integrity and ensuring a closed-loop business model.

[0062] In batch job scheduling for production application systems, jobs typically consist of multiple nodes with dependencies, which are executed sequentially according to a predetermined schedule (such as a fixed date each day or month). However, in actual business operations, nodes in the job plan may need to be skipped due to the influence of upstream and downstream data or business operations. Therefore, job scheduling generally supports skipping batch nodes for a specific day. To this end, scheduling systems are typically designed with the ability to skip nodes for specific dates. This flexible scheduling mechanism not only ensures the regular execution logic of batch jobs, but also adapts to the special needs of business scenarios.

[0063] While node skipping during batch job scheduling can flexibly adapt to business changes, it can also lead to nodes being mistakenly skipped due to configuration errors or system anomalies. These unexpected node skipping events are categorized as "abnormal skipping." When critical nodes are abnormally skipped, data processing on subsequent dependent nodes can fail, leading to data processing interruptions or incorrect results. Therefore, to ensure batch job stability and data processing accuracy, it is essential to establish an effective abnormal skipping detection mechanism to identify whether nodes have been abnormally skipped.

[0064] Currently, whether a node is abnormally skipped is mainly determined by manual processing. Specifically, when a node is skipped, whether it is an abnormal skip is manually determined based on expert experience.

[0065] However, identifying whether a node has been abnormally skipped requires considering multiple factors, such as the date of operation, historical operation history, and batch business scenarios. Manual judgment relies on expert experience and expertise, requiring a certain level of business knowledge about batch nodes, which can't guarantee accuracy or efficiency. Furthermore, the large number of batch nodes requires significant labor costs and low efficiency.

[0066] In other words, the existing technology has the problems of low processing efficiency and accuracy, and high labor costs.

[0067] Based on this, the technical concept of the present application is as follows: taking into account the recent skipping of nodes, as well as the date information and business attribute information of the day, will determine whether the node is skipped. For example, on holidays (date information of the day), the node does not need to process business, and the node can be skipped at this time; on the interest payment date (business attribute information of the day), the node needs to process business, and the node cannot be skipped at this time; if the node has been in a skipping state for nearly a month (recent skipping situation), the node can be skipped at this time. Therefore, the recent skipping of nodes, as well as the date information and business attribute information of the day, can be analyzed to allow the machine learning model to automatically capture the pattern features of normal skipping (such as system silence on holidays, periodic business lows, etc.), thereby identifying abnormal skipping that does not conform to conventional logic, which not only retains the key dimensions of manual research and judgment, but also realizes long-term automatic monitoring through the model, effectively solving the problems of low accuracy and low efficiency in manual processing, and reducing labor costs.

[0068] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0069] Figure 1 Schematic diagram of the process of node abnormal skipping detection method provided by this application Figure 1 ,like Figure 1 As shown, the method can be implemented by the following steps:

[0070] S11. When a corresponding abnormal skip detection model exists for a skip node, node skip information of a target day for skipping the node, as well as date information and service attribute information of the day, are obtained.

[0071] The execution subject of the embodiments of the present application is an electronic device, which can be a terminal device, such as a laptop computer, a desktop computer, a tablet computer, etc., or a server. In actual applications, whether the electronic device is a terminal device or a server can be determined based on actual conditions and is not specifically limited to this.

[0072] In this step, a skipped node refers to a node that is skipped in a batch of nodes. This node needs to be checked for node abnormality to verify whether the skipping is due to an error or other unexpected circumstances. In practical applications, a corresponding abnormal skip detection model can be pre-deployed for one or more nodes in the batch of nodes. When the node is a skipped node, the corresponding abnormal skip detection model is called to perform node abnormality skip detection on the skipped node.

[0073] It should be understood that the abnormal skip detection model corresponding to each node can be deployed locally on the electronic device, and the abnormal skip detection model corresponding to each node can also be deployed on other devices (such as the cloud). It can be determined based on actual conditions, and the embodiments of the present application do not impose specific restrictions on this.

[0074] The target date is before the current date, and the time interval between the target date and the current date is less than a preset time interval.

[0075] Exemplarily, the preset time interval may be 1 month, 2 months, or 3 months, etc., which may be pre-set according to actual conditions, and the embodiments of the present application do not impose any specific restrictions on this.

[0076] Among them, the number of target days can be 1 or more than 1, which can be pre-set according to actual conditions, and the embodiment of the present application does not impose specific restrictions on this.

[0077] That is, the target day may be at least one day in the month, two months, or three months before the current day, and the target day is a date adjacent to the current day.

[0078] The business attribute information of the day is used to indicate whether there is any business that needs to be processed on that day.

[0079] For example, the business attribute information of the day may include whether the day is an interest payment day.

[0080] The node skipping information of the target day is used to indicate whether the skipped node is skipped on the target day.

[0081] Optionally, the date information of the day includes at least one of the following: the date of the day, a week attribute, and a holiday attribute.

[0082] It should be understood that the week attribute refers to the day of the week, the holiday attribute refers to whether the day is a holiday, and the date of the day refers to the specific date of the day (January 11, 2000) as well as the month (January), year (2000), and day (11th) of the day.

[0083] It should be understood that the abnormal skip detection model is trained based on the node skip information of the first sample day of the sample node and the date information and business attribute information of the second sample day. The first sample day is before the second sample day, and the time interval between the first sample day and the second sample day is less than the preset time interval.

[0084] It should be understood that during model training, each node in a batch of nodes is a sample node. The node skip information for the first sample day, as well as the date information and service attribute information for the second sample day, can be obtained from the node's historical operation data, which can be an operation log. The first sample day can refer to concepts related to the target day, and the second sample day can refer to concepts related to the current day. The first sample day is an adjacent date to the second sample day.

[0085] It should be understood that the execution subject of the model training process is an electronic device, which can be the same as or different from the electronic device that executes the node abnormal skip detection method. That is, the electronic device can first perform model training to obtain and deploy the abnormal skip detection model corresponding to each node, and then perform node abnormal skip detection through the corresponding abnormal skip detection model when it is detected that the node is a skip node; the electronic device can also obtain the abnormal skip detection model corresponding to each node from other devices in advance and deploy it locally, and then perform node abnormal skip detection through the corresponding abnormal skip detection model when it is detected that the node is a skip node; the electronic device can also send the node skip information of the target day of the skip node and the date information and business attribute information of the day to the electronic device where the abnormal skip detection model is deployed when it detects that the node is a skip node, and obtain the first abnormal detection result returned by the electronic device, and the first abnormal detection result is used to indicate whether the skip node is an abnormal skip.

[0086] S12: Input the node skipping information of the target day, the date information of the current day, and the service attribute information into an abnormal skipping detection model to obtain a first abnormality detection result output by the abnormal skipping detection model.

[0087] In this step, the node skipping information of the target day, as well as the date information and business attribute information of the day, are the input data required by the abnormal skipping detection model. After obtaining the node skipping information of the target day, as well as the date information and business attribute information of the day, these data can be input into the abnormal skipping detection model to obtain the first abnormality detection result output by the abnormal skipping detection model. The first abnormality detection result is used to indicate whether the skipped node is an abnormal skip.

[0088] The abnormal skip detection model may be a decision tree model.

[0089] The node abnormal skip detection method provided by the embodiment of the present application, when there is a corresponding abnormal skip detection model for the skipped node, obtains the node skip information of the target day of the skipped node and the date information and business attribute information of the current day, inputs the node skip information of the target day and the date information and business attribute information of the current day into the abnormal skip detection model, and obtains the first abnormal detection result output by the abnormal skip detection model. Wherein, the target day is before the current day, and the time interval between the target day and the current day is less than the preset time interval, and the business attribute information of the current day is used to indicate whether there is any business that needs to be processed on the current day. The first abnormal detection result is used to indicate whether the skipped node is an abnormal skip, and the abnormal skip detection model is trained based on the node skip information of the first sample day of the sample node and the date information and business attribute information of the second sample day. The first sample day is before the second sample day, and the time interval between the first sample day and the second sample day is less than the preset time interval. In this technology, the abnormal skip detection model can determine whether it meets the pattern characteristics of normal skipping based on the node skipping information of the target day of the skipped node and the date information and business attribute information of the day, and then identify abnormal skipping that does not conform to conventional logic, and output the first abnormality detection result, replacing the manual processing process, solving the problem of low accuracy caused by the inability to guarantee rigor in manual processing, reducing labor costs and improving processing efficiency.

[0090] When there is no corresponding abnormal skip detection model for the skip node, the historical skip probability corresponding to the skip node may be obtained, and then the second abnormality detection result may be output according to the historical skip probability and the preset skip probability.

[0091] The second abnormality detection result is used to indicate whether the skipped node is an abnormal skip.

[0092] Specifically, the historical skip probability can be calculated based on the historical skip information of the sample node.

[0093] For example, assuming that the historical skip information includes historical skip situations of the sample node in the past three months (90 days), of which a total of 30 days are skipped, the historical skip probability is 30%.

[0094] Exemplarily, the preset skip probability may be 50%, 60% or 70%, etc., which may be preset according to actual conditions, and the embodiments of the present application do not impose any specific limitation on this.

[0095] In one possible implementation, the above “outputting the second anomaly detection result according to the historical skip probability and the preset skip probability” can be implemented as follows:

[0096] If the historical skip probability is greater than or equal to the preset skip probability, a second abnormality detection result indicating that the skipped node is normally skipped is output; if the historical skip probability is less than the preset skip probability, a second abnormality detection result indicating that the skipped node is abnormally skipped is output.

[0097] Continuing with the above example, if the preset skip probability is 50%, since the historical skip probability (30%) is less than the preset skip probability (50%), a second abnormality detection result indicating that the skipped node is an abnormal skip is output.

[0098] In the above embodiment, the absence of a corresponding abnormal skip detection model for a skip node may be because the skip node does not meet the model deployment conditions, or it may be because the abnormal skip detection model corresponding to the skip node is still being trained and has not yet been fully trained. In this case, by quantifying historical behavior and converting expert experience into a computable probability threshold, we avoid the efficiency bottleneck of relying entirely on manual review while ensuring the objective consistency of the judgment criteria. This allows us to still perform node abnormal skip detection on the skip node even when the corresponding abnormal skip detection model is not deployed, thus ensuring system stability while reducing the pressure of manual operation and maintenance.

[0099] In practical applications, a corresponding abnormal skip detection model can be deployed for each node in a batch of nodes. It is also possible to pre-judge whether each node meets the model deployment conditions, so as to deploy the corresponding abnormal skip detection model for the nodes that meet the model deployment conditions.

[0100] Exemplarily, the model deployment condition may be that the data volume of the node's training samples is greater than a preset training sample volume.

[0101] It should be understood that Figure 2 The illustrated embodiment illustrates how to determine whether each node meets the model deployment conditions, and the corresponding subsequent processing under different situations (meeting the model deployment conditions and not meeting the model deployment conditions).

[0102] Figure 2Schematic diagram of the process of node abnormal skipping detection method provided by this application Figure 2 ,like Figure 2 As shown, before S11, the method may further include the following steps:

[0103] S21. Obtain multiple training samples for each sample node.

[0104] Each training sample includes node skipping information of the first sample day of the sample node and date information and service attribute information of the second sample day.

[0105] S22: If the data volume of the multiple training samples of the sample node is less than the preset training sample volume, calculate the historical skip probability of the sample node based on the historical skip information of the sample node.

[0106] It should be understood that the preset training sample size is the standard for judging data sparsity. That is to say, when the data volume of multiple training samples of a sample node is less than the preset training sample size, it means that the data is sparse at this time, and the accuracy of the model obtained by model training based on the multiple training samples cannot be guaranteed; on the contrary, when the data volume of multiple training samples of a sample node is greater than or equal to the preset training sample size, it means that the data is relatively rich at this time, and model training can be performed based on the multiple training samples, and the accuracy of the model obtained by the model training is higher.

[0107] It should be understood that the preset training sample size can be determined in advance based on expert experience or relevant experimental data, and the embodiments of the present application do not impose specific limitations on this.

[0108] S23: If the data volume of the multiple training samples of the sample node is greater than or equal to the preset training sample volume, determine the label of each training sample.

[0109] The label of the training sample is used to indicate whether the sample node has skipped processing on the second sample day.

[0110] S24: Perform model training based on multiple training samples and the label of each training sample to generate an abnormal skip detection model.

[0111] In an embodiment of the present application, whether the sample node meets the model deployment conditions is determined based on the training sample size of each sample node to ensure the accuracy of the abnormal skip detection model obtained by model training.

[0112] Next, the model training process and model reasoning process involved in the node abnormality skipping detection solution are explained through two specific examples.

[0113] Figure 3 A flow chart of the model training process of the node abnormal skipping detection method provided in this application is as follows: Figure 3As shown, the model training process can be achieved through the following steps:

[0114] S301: Obtain historical operation data of all sample nodes in a batch of nodes.

[0115] The historical operation data includes the operation data of each sample node on each historical date.

[0116] S302: For the current sample node, determine whether the data volume of multiple training samples of the sample node is greater than or equal to a preset training sample volume.

[0117] If the data volume of the multiple training samples of the sample node is greater than or equal to the preset training sample volume, execute S303; if the data volume of the multiple training samples of the sample node is less than the preset training sample volume, execute S309.

[0118] S303: Segment the historical operation data by date, and obtain sub-operation data for each historical date.

[0119] It should be understood that each historical date is the second sample date.

[0120] S304: Mark the sample node on each historical date based on whether the sample node is skipped on each historical date.

[0121] Optionally, a 0 / 1 value may be used as a label. For example, 0 may be used to indicate that the sample node was not skipped on the historical date, and 1 may be used to indicate that the sample node was skipped on the historical date.

[0122] It should be understood that the sample node is marked on each historical date according to the sub-operation data of each historical date.

[0123] S305: Extract date information and business attribute information of each historical date.

[0124] It should be understood that binary features (whether it is a holiday, whether it is an interest payment date, etc.) can be represented by 0 / 1 values, and other features (the year, month, day, etc. of the historical date) can be represented by data.

[0125] It should be understood that the date information and business attribute information of each historical date are extracted from the sub-operation data of each historical date.

[0126] S306: Extract node skipping information close to the historical date.

[0127] It should be understood that the near historical date is the first sample day. Exemplarily, the near historical date may be seven days before the historical date.

[0128] Optionally, node skip information may be represented by a 0 / 1 value.

[0129] Optionally, when there are multiple adjacent historical dates, the node skipping information of the adjacent historical dates can be represented by a 0 / 1 sequence.

[0130] It should be understood that the date information and business attribute information of each historical date are extracted from the sub-operation data of each historical date.

[0131] S307 : For each historical date, generate a training sample based on the node skipping information adjacent to the historical date and the date information and business attribute information of the historical date.

[0132] For each historical date, the node skipping information close to the historical date and the date information and business attribute information of the historical date are aligned by date to generate training samples. Specifically,

[0133] For example, the node skipping information for nodes adjacent to a historical date can be used as a column of data, the date information for the historical date can be used as a column of data, and the business attribute information for the historical date can be used as a column of data to construct a feature matrix. Each row of data in the feature matrix corresponds to a historical date, that is, each row of data includes the date information, business attribute information, and the node skipping information for the historical date adjacent to the historical date.

[0134] It should be understood that each row of data in the feature matrix is ​​a training sample, that is, the feature matrix is ​​a set of training samples for the sample node.

[0135] It should be understood that the label obtained by marking the sample node on each historical date in S304 is the label of the training sample.

[0136] S308: Input the training samples into the decision tree model for training to obtain an abnormal skip detection model.

[0137] Alternatively, the CART decision tree algorithm can be used for model training. The input to the model learning is the training sample, and the output is the predicted value corresponding to the training sample. CART uses the Gini coefficient to measure information impurity, selecting the node with the smallest Gini coefficient as the node feature. The CART decision tree is a binary tree, meaning that each node has only two branches.

[0138] It should be understood that model training can also be performed using other algorithms, and the embodiments of the present application do not impose specific limitations on this.

[0139] S309: Calculate the historical skip probability of the sample node according to the historical skip information of the sample node.

[0140] Exemplarily, the historical skip information may include the skipping status of the sample node within a certain time range, such as the past three years.

[0141] S310: Serialize and save the abnormal skip detection model or historical skip probability of the sample node.

[0142] Afterwards, the next sample node is used as the new current sample node, and steps S302 to S310 are repeated until all sample nodes are traversed.

[0143] It should be understood that for a sample node with a corresponding abnormal skip detection model, the historical skip probability of the sample node can also be calculated simultaneously, so that when the abnormal skip detection model fails, node abnormal skip detection can be performed based on the historical skip probability.

[0144] Figure 4 A flow chart of the model reasoning process of the node abnormal skipping detection method provided in this application is as follows: Figure 4 As shown, the model inference process can be achieved through the following steps:

[0145] S401: Determine skip nodes in batch nodes.

[0146] S402: Determine whether there is a corresponding abnormal skip detection model for the skip node.

[0147] If the skip node has a corresponding abnormal skip detection model, execute S403; if the skip node does not have a corresponding abnormal skip detection model, execute S406.

[0148] It should be understood that the abnormal skip detection model is Figure 3 The embodiment shown is obtained by model training.

[0149] S403: Load the abnormal skip detection model corresponding to the skip node.

[0150] S404: Extract and concatenate the node skipping information of the target day of the skipped node, the date information of the day, and the business attribute information to generate input data for the abnormal skipping detection model.

[0151] S405: Input the input data into the abnormal skip detection model to obtain a first abnormality detection result output by the abnormal skip detection model.

[0152] S406: Obtain the historical skip probability corresponding to the skip node.

[0153] S407: Determine whether the historical skip probability is greater than or equal to the preset skip probability.

[0154] If the historical skip probability is greater than or equal to the preset skip probability, execute S408; if the historical skip probability is less than the preset skip probability, execute S409.

[0155] S408: Output a second abnormality detection result indicating that the skipped node is normally skipped.

[0156] S409: Output a second abnormality detection result indicating that the skipped node is abnormally skipped.

[0157] In summary, this application includes two parts: offline training and real-time detection. Offline training directly calculates the historical skip probability for sample nodes with insufficient training sample data. For sample nodes with sufficient training sample data, a decision tree model is used for model training to obtain the corresponding abnormal skip detection model. Specifically, a feature matrix is ​​constructed based on the historical operation data of the sample node. Whether the sample node was skipped on each historical date is used as the predicted value corresponding to that historical date. After extracting features from the historical operation data through feature engineering, a decision tree model is trained. The features include date features, business attribute features, and features of skipping on nearby dates. The decision path of the decision tree model can be explained and described based on the meaning of the features. During real-time detection, for skipped nodes with an abnormal skip detection model, the date information and business attribute information of the skipped node on the current day, as well as the node skip information on nearby dates, are obtained, and features are extracted. The extracted features are input into the abnormal skip detection model, and the first abnormal detection result of the skipped node is output in real time. For skipped nodes without an abnormal skip detection model, the historical skip probability is used as the confidence level to determine whether the skip is abnormal.

[0158] In other words, the node abnormal skipping detection solution provided by this application has the following technical effects:

[0159] 1. It can efficiently and automatically identify abnormal skipping of nodes, avoiding problems such as manual configuration errors, scheduling anomalies, etc. that cause node skipping and affect the normal operation of batches.

[0160] 2. It does not rely on expert experience or batch job business knowledge, but objectively identifies abnormal skipping of nodes based on the historical operation data of batch nodes.

[0161] 3. Explore the relationship between node skipping and multi-dimensional data such as the current date, business, and upcoming dates to ensure that node skipping at monthly, quarterly, or specific date intervals can be accurately identified.

[0162] 4. The abnormal skip detection model is a decision tree algorithm with good recognition effect and strong model interpretability.

[0163] 5. For abnormal skipping situations, clear and explainable judgment criteria can be given, providing a good user experience.

[0164] 6. Before modeling a sample node, determine whether the sample node meets the model deployment requirements. If the sample node does not have sufficient training data, meaning it does not meet the model deployment requirements, the anomaly skip detection model for that sample node will not be trained. Instead, the historical skip probability will be used as the basis for generating a second anomaly detection result, effectively improving the accuracy of the anomaly skip detection model.

[0165] Figure 5 This is a schematic diagram of the structure of the node abnormal skipping detection device provided by this application, as shown in Figure 5 As shown, the node abnormal skipping detection device 50 provided in this embodiment includes:

[0166] The acquisition module 51 is used to obtain the node skipping information of the target day of the skipped node and the date information and business attribute information of the current day when there is a corresponding abnormal skipping detection model for the skipped node. The target day is before the current day, and the time interval between the target day and the current day is less than the preset time interval. The business attribute information of the current day is used to indicate whether there is any business that needs to be processed on the current day.

[0167] The input module 52 is used to input the node skipping information of the target day and the date information and business attribute information of the day into the abnormal skipping detection model, and obtain the first abnormality detection result output by the abnormal skipping detection model. The first abnormality detection result is used to indicate whether the skipped node is an abnormal skip.

[0168] Among them, the abnormal skip detection model is trained based on the node skip information of the first sample day of the sample node and the date information and business attribute information of the second sample day. The first sample day is before the second sample day, and the time interval between the first sample day and the second sample day is less than the preset time interval.

[0169] In a possible implementation, the date information of the current day includes at least one of the following: the current day's date, a week attribute, and a holiday attribute.

[0170] In one possible embodiment, the node abnormal skip detection device 50 also includes a calculation module. When there is a corresponding abnormal skip detection model for the skipped node, before obtaining the node skip information of the target day of the skipped node and the date information and business attribute information of the day, the acquisition module 51 is also used to obtain multiple training samples for each sample node. The training samples include the node skip information of the first sample day of the sample node and the date information and business attribute information of the second sample day.

[0171] The calculation module is used to calculate the historical skip probability of the sample node based on the historical skip information of the sample node if the data amount of multiple training samples of the sample node is less than the preset training sample amount.

[0172] Correspondingly, when there is no corresponding abnormal skip detection model for the skip node, the acquisition module 51 is further configured to acquire the historical skip probability corresponding to the skip node.

[0173] The output module is further configured to output a second anomaly detection result based on the historical skip probability and the preset skip probability, where the second anomaly detection result is used to indicate whether the skipped node is an abnormal skip.

[0174] In a possible implementation, the output module is specifically configured to:

[0175] If the historical skip probability is greater than or equal to the preset skip probability, a second abnormality detection result indicating that the skipped node is a normal skip is output.

[0176] If the historical skip probability is less than the preset skip probability, a second abnormality detection result indicating that the skipped node is an abnormal skip is output.

[0177] In a possible implementation, the node abnormal skipping detection device 50 further includes a training module, which is configured to:

[0178] If the data volume of multiple training samples of the sample node is greater than or equal to the preset training sample volume, a label of each training sample is determined, and the label of the training sample is used to indicate whether the sample node has skipped processing on the second sample day.

[0179] The model is trained based on multiple training samples and the labels of each training sample to generate an abnormal skip detection model.

[0180] In a possible implementation, the abnormal skip detection model is a decision tree model.

[0181] The node abnormal skipping detection device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effects are similar, and are not described in detail in this embodiment.

[0182] Figure 6 This is a schematic diagram of the structure of the electronic device provided in this application. Figure 6 As shown, the electronic device 60 provided in this embodiment includes: at least one processor 601 and a memory 602. Optionally, the electronic device 60 further includes a communication component 603. The processor 601, the memory 602 and the communication component 603 are connected via a bus 604.

[0183] During the specific implementation process, at least one processor 601 executes the computer-executable instructions stored in the memory 602, so that the at least one processor 601 performs the above method.

[0184] The specific implementation process of the processor 601 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.

[0185] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), etc. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly implemented by a hardware processor or implemented by a combination of hardware and software modules in the processor.

[0186] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory.

[0187] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be classified into address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.

[0188] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0189] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.

[0190] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0191] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist in a device as discrete components.

[0192] The division of units is merely a logical functional division; actual implementations may employ alternative divisions, such as combining or integrating multiple units or components into another system, or omitting or disabling certain features. Furthermore, any direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units, either through an interface, electrical, mechanical, or other means.

[0193] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0194] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0195] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, ROM, RAM, disk or optical disk, etc. Various media that can store program code.

[0196] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0197] Finally, it should be noted that those skilled in the art will readily identify other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed herein. The present invention is not limited to the precise structure described above and illustrated in the accompanying drawings, and various modifications and variations may be made without departing from the scope thereof. The scope of the present invention is limited solely by the appended claims.

Claims

1. A method for detecting abnormal node skipping, characterized in that: include: When a corresponding abnormal skip detection model exists for a skip node, obtaining node skip information of a target day of the skip node and date information and business attribute information of the current day, where the target day is before the current day and the time interval between the target day and the current day is less than a preset time interval, and the business attribute information of the current day is used to indicate whether there is business that needs to be processed on the current day; Inputting the node skipping information of the target day, the date information of the current day, and the service attribute information into the abnormal skipping detection model, and obtaining a first abnormality detection result output by the abnormal skipping detection model, wherein the first abnormality detection result is used to indicate whether the skipped node is an abnormal skip; In which, the abnormal skip detection model is trained based on the node skip information of the first sample day of the sample node and the date information and business attribute information of the second sample day, the first sample day is before the second sample day, and the time interval between the first sample day and the second sample day is less than the preset time interval.

2. The method according to claim 1, characterized in that The date information of the current day includes at least one of the following: the current day's date, a weekday attribute, and a holiday attribute.

3. The method according to claim 1 or 2, characterized in that When the skip node has a corresponding abnormal skip detection model, before obtaining the node skip information of the target day of the skip node and the date information and service attribute information of the day, the method further includes: Acquire multiple training samples for each sample node, where the training samples include node skipping information of the first sample day and date information and service attribute information of the second sample day of the sample node; If the data amount of the multiple training samples of the sample node is less than the preset training sample amount, calculating the historical skip probability of the sample node based on the historical skip information of the sample node; Accordingly, when the skip node does not have a corresponding abnormal skip detection model, the method further includes: Obtaining a historical skip probability corresponding to the skipped node; A second abnormality detection result is output according to the historical skip probability and the preset skip probability, where the second abnormality detection result is used to indicate whether the skipped node is an abnormal skip.

4. The method according to claim 3, characterized in that Outputting a second anomaly detection result according to the historical skip probability and the preset skip probability includes: If the historical skip probability is greater than or equal to the preset skip probability, outputting a second abnormality detection result indicating that the skipped node is skipped normally; If the historical skip probability is less than the preset skip probability, a second abnormality detection result indicating that the skipped node is abnormally skipped is output.

5. The method according to claim 3, characterized in that The method further comprises: If the data amount of the multiple training samples of the sample node is greater than or equal to the preset training sample amount, determining a label for each training sample, wherein the label of the training sample is used to indicate whether the sample node has skipped processing on the second sample day; Model training is performed based on the multiple training samples and the label of each training sample to generate the abnormal skip detection model.

6. The method according to claim 1, 2, 4 or 5, characterized in that The abnormal skip detection model is a decision tree model.

7. A node abnormal skipping detection device, characterized in that: include: an acquisition module, configured to, when a corresponding abnormal skip detection model exists for a skip node, acquire node skip information of a target day for the skip node, as well as date information and business attribute information of the current day, wherein the target day is before the current day and the time interval between the target day and the current day is less than a preset time interval, and the business attribute information of the current day is used to indicate whether there is any business that needs to be processed on the current day; an input module, configured to input the node skipping information of the target day, the date information of the current day, and the service attribute information into the abnormal skip detection model, and obtain a first abnormality detection result output by the abnormal skip detection model, wherein the first abnormality detection result is used to indicate whether the skipped node is an abnormal skip; In which, the abnormal skip detection model is trained based on the node skip information of the first sample day of the sample node and the date information and business attribute information of the second sample day, the first sample day is before the second sample day, and the time interval between the first sample day and the second sample day is less than the preset time interval.

8. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.

10. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 6 when the computer program is executed by a processor.