Risk judgment method and device, computer equipment and storage medium

By performing risk vocabulary recognition and abnormal sequence recognition on the system log and user operation log, combining association analysis and clustering technology, a risk determination decision tree and knowledge base are built, which solves the problem of difficult to effectively identify and determine risk operations in the existing technology, and achieves more efficient and accurate risk determination.

CN120068060APending Publication Date: 2025-05-30KANG JIAN INFORMATION TECH (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510227797.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the fields of financial technology and digital medical, it is difficult to effectively identify and determine potential risk operations, especially because the construction of a risk vocabulary relies on manual experience and lack of accuracy, making it difficult to fully cover all possible risk scenarios.

Method used

By obtaining system log information and preset risk vocabulary database, risk vocabulary recognition is carried out; user operation logs and standard operation sequence templates are obtained to perform abnormal sequence recognition; risk vocabulary information and abnormal sequence results are correlated and analyzed and clustered, risk determination tree and knowledge base are built, and risk determination is finally made on real-time system logs and operation logs.

Benefits of technology

It realizes effective risk determination of users' operation logs on the system, improves the accuracy and efficiency of risk determination, and can more comprehensively identify and deal with potential risk operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068060A_ABST
    Figure CN120068060A_ABST
Patent Text Reader

Abstract

The embodiment of the invention belongs to the field of data processing, and relates to a risk judgment method and device, computer equipment and a storage medium, and the method comprises the following steps: carrying out risk vocabulary recognition on system log information according to a preset risk vocabulary library to obtain risk vocabulary information; performing abnormal sequence identification according to the user operation log and the standard operation sequence template to obtain an abnormal sequence result; performing association analysis on the risk vocabulary information and the abnormal sequence result to obtain association degree information; clustering the risk vocabulary information and the abnormal sequence result according to the association degree information to obtain a risk abnormal cluster; constructing a risk judgment decision-making tree according to the risk anomaly clustering cluster, and generating a risk judgment knowledge base according to a risk judgment rule of the risk judgment decision-making tree; and performing risk judgment on the real-time system log and the real-time operation log according to the risk judgment knowledge base to obtain a risk judgment result. According to the method, effective risk judgment can be carried out on the operation log of the user on the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, specifically to the field of fintech, and particularly to a risk determination method, apparatus, computer device, and storage medium. Background Art

[0002] In the fields of fintech and digital healthcare, with the increasing complexity of enterprise information systems and the explosive growth of data volume, how to quickly and accurately identify potential risk operations from a vast amount of log data has become an important technical problem to be solved urgently. These two industries are not only related to the capital security and privacy protection of users, but also directly affect the accuracy and security of medical services. Therefore, the requirements for log analysis technology are particularly strict.

[0003] Currently, log analysis based on a predefined risk vocabulary library is a common method for identifying risk operations. This method discovers potential risk behaviors by matching keywords in the logs. However, this method exposes two major limitations in practical applications. First, the construction of the risk vocabulary library highly depends on manual experience, which is not only time-consuming and laborious, but also limited by personal cognitive scope and industry experience, resulting in a relatively limited coverage of the risk vocabulary library and being difficult to comprehensively cover all possible risk scenarios.

[0004] Secondly, relying solely on risk vocabulary matching to identify risk operations also greatly reduces its accuracy. Because in actual business operations, some hidden risk behaviors may evade detection by means of splitting execution, disguising as normal operations, etc. These complex risk operation sequences are often difficult to discover through simple keyword matching, thus bringing great potential risks to the information security of enterprises. Summary of the Invention

[0005] The purpose of the embodiments of the present application is to propose a risk determination method, apparatus, computer device, and storage medium to solve the problem of being unable to effectively determine the risk of user operation logs on the system.

[0006] To solve the above technical problems, the embodiments of the present application provide a risk determination method, which adopts the following technical solutions:

[0007] Obtain system log information and a preset risk vocabulary library, and perform risk vocabulary recognition on the system log information according to the preset risk vocabulary library to obtain risk vocabulary information;

[0008] Obtain user operation logs and a standard operation sequence template, and perform abnormal sequence recognition according to the user operation logs and the standard operation sequence template to obtain an abnormal sequence result;

[0009] Perform a correlation analysis on the risk vocabulary information and the abnormal sequence results to obtain correlation degree information;

[0010] Generate an effective weight vector based on the correlation degree information, and cluster the risk vocabulary information and the abnormal sequence results according to the effective weight vector to obtain a risk abnormal clustering cluster;

[0011] Construct a risk determination decision tree based on the risk abnormal clustering cluster, and generate a risk determination knowledge base according to the risk determination rules of the risk determination decision tree;

[0012] Obtain real-time system logs and real-time operation logs, and perform risk determination on the real-time system logs and the real-time operation logs according to the risk determination knowledge base to obtain a risk determination result.

[0013] Further, the step of obtaining system log information and a preset risk vocabulary library, and performing risk vocabulary recognition on the system log information according to the preset risk vocabulary library specifically includes:

[0014] Obtain an information extraction identifier and a vocabulary library extraction identifier, and extract the system log information and the preset risk vocabulary library from the database according to the information extraction identifier and the vocabulary library extraction identifier;

[0015] Perform text preprocessing on the system log information to obtain standard system log information;

[0016] Use the preset risk vocabulary library as a feature dictionary, and match the standard system log information and the feature dictionary according to a string matching algorithm to obtain risk vocabulary information.

[0017] Further, the step of obtaining user operation logs and a standard operation sequence template, and performing abnormal sequence recognition on the user operation logs and the standard operation sequence template to obtain abnormal sequence results specifically includes:

[0018] Obtain a log extraction identifier and a template extraction identifier, and extract the user operation logs and the standard operation sequence template from the database according to the log extraction identifier and the template extraction identifier;

[0019] Analyze the user operation logs based on a sequence pattern mining algorithm to obtain actual operation sequence information;

[0020] Calculate the edit distance between the actual operation sequence information and the standard operation sequence template to obtain an edit distance value;

[0021] Determine whether the edit distance value is greater than or equal to a preset distance threshold;

[0022] If the edit distance value is greater than or equal to the preset distance threshold, the actual operation sequence information is regarded as an abnormal sequence, and the abnormal sequence result is output.

[0023] Further, the step of performing correlation analysis on the risk vocabulary information and the abnormal sequence result to obtain correlation degree information specifically includes:

[0024] Performing correlation analysis on the risk vocabulary information and the abnormal sequence result based on a preset correlation rule mining algorithm to obtain correlation rule information;

[0025] Calculating the support, confidence, and lift of the correlation rule information, and screening the correlation rule information according to the support, the confidence, and the lift to obtain effective correlation rules;

[0026] Calculating the correlation degree between the risk vocabulary information and the abnormal sequence result according to the effective correlation rules, and generating corresponding correlation degree information.

[0027] Further, the step of generating an effective weight vector according to the correlation degree information and clustering the risk vocabulary information and the abnormal sequence result according to the effective weight vector to obtain a risk abnormal clustering cluster specifically includes:

[0028] Selecting the risk vocabulary and abnormal operation sequences with positive correlation from the risk vocabulary information and the abnormal sequence result according to the correlation degree information;

[0029] Calculating the weights of the risk vocabulary and the abnormal operation sequences according to the information gain algorithm to obtain the effective weight vector;

[0030] Performing feature extraction on the risk vocabulary and the abnormal operation sequences to obtain a risk vocabulary feature vector and an operation sequence feature vector;

[0031] Performing clustering analysis based on a preset clustering algorithm according to the risk vocabulary feature vector, the operation sequence feature vector, and the effective weight vector to obtain the risk abnormal clustering cluster.

[0032] Further, the step of constructing a risk determination decision tree according to the risk abnormal clustering cluster and generating a risk determination knowledge base according to the risk determination rules of the risk determination decision tree specifically includes:

[0033] Constructing the risk determination decision tree for the risk abnormal clustering cluster based on the decision tree algorithm;

[0034] Obtaining the paths from the root node to the leaf nodes of the risk determination decision tree as the risk determination rules to obtain a risk determination rule set;

[0035] Store the risk determination rule set in a predetermined knowledge base for preservation to obtain the risk determination knowledge base.

[0036] Further, the steps of obtaining real-time system logs and real-time operation logs, and performing risk determination on the real-time system logs and the real-time operation logs according to the risk determination knowledge base to obtain a risk determination result specifically include:

[0037] Obtain a system log record table and an operation log record table, and extract the real-time system logs and the real-time operation logs from the system log record table and the operation log record table according to the current timestamp;

[0038] Input the real-time system logs and the real-time operation logs into the risk determination knowledge base for matching to obtain the risk determination result.

[0039] To solve the above technical problems, an embodiment of the present application further provides a risk determination device, which adopts the following technical solutions:

[0040] An information acquisition module, configured to acquire system log information and a preset risk vocabulary library, and perform risk vocabulary recognition on the system log information according to the preset risk vocabulary library to obtain risk vocabulary information;

[0041] A sequence recognition module, configured to acquire user operation logs and a standard operation sequence template, and perform abnormal sequence recognition on the user operation logs and the standard operation sequence template to obtain an abnormal sequence result;

[0042] An association analysis module, configured to perform association analysis on the risk vocabulary information and the abnormal sequence result to obtain association degree information;

[0043] An information clustering module, configured to generate an effective weight vector according to the association degree information, and perform clustering on the risk vocabulary information and the abnormal sequence result according to the effective weight vector to obtain a risk abnormal clustering cluster;

[0044] A knowledge base generation module, configured to construct a risk determination decision tree according to the risk abnormal clustering cluster, and generate a risk determination knowledge base according to the risk determination rules of the risk determination decision tree;

[0045] A risk determination module, configured to acquire real-time system logs and real-time operation logs, and perform risk determination on the real-time system logs and the real-time operation logs according to the risk determination knowledge base to obtain a risk determination result.

[0046] To solve the above technical problems, an embodiment of the present application further provides a computer device, which adopts the following technical solutions:

[0047] A computer device includes a memory and a processor. Computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of the risk determination method described in any one of the above are implemented.

[0048] To solve the above technical problems, an embodiment of the present application also provides a computer-readable storage medium, adopting the following technical solutions:

[0049] A computer-readable storage medium has computer-readable instructions stored thereon. When the computer-readable instructions are executed by a processor, the steps of the risk determination method described in any one of the above are implemented.

[0050] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects: In this embodiment, system log information and a preset risk vocabulary library are obtained, and risk vocabulary recognition is performed on the system log information according to the preset risk vocabulary library to obtain risk vocabulary information; user operation logs and a standard operation sequence template are obtained, and abnormal sequence recognition is performed according to the user operation logs and the standard operation sequence template to obtain an abnormal sequence result; correlation analysis is performed on the risk vocabulary information and the abnormal sequence result to obtain correlation degree information; an effective weight vector is generated according to the correlation degree information, and clustering is performed on the risk vocabulary information and the abnormal sequence result according to the effective weight vector to obtain a risk abnormal clustering cluster; a risk determination decision tree is constructed according to the risk abnormal clustering cluster, and a risk determination knowledge base is generated according to the risk determination rules of the risk determination decision tree; real-time system logs and real-time operation logs are obtained, and risk determination is performed on the real-time system logs and the real-time operation logs according to the risk determination knowledge base to obtain a risk determination result. Thereby, effective risk determination of the operation logs of users on the system is realized, and the accuracy and efficiency of risk determination are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] To more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0052] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;

[0053] Figure 2 A flowchart according to an embodiment of the risk determination method of the present application;

[0054] Figure 3 is Figure 2Flowchart of a specific implementation of step S10;

[0055] Figure 4 is Figure 2 Flowchart of a specific implementation of step S20;

[0056] Figure 5 is Figure 2 Flowchart of a specific implementation of step S30;

[0057] Figure 6 is Figure 2 Flowchart of a specific implementation of step S40;

[0058] Figure 7 is Figure 2 Flowchart of a specific implementation of step S50;

[0059] Figure 8 is Figure 2 Flowchart of a specific implementation of step S60;

[0060] Figure 9 Structural schematic diagram of an embodiment of a risk determination device according to the present application;

[0061] Figure 10 Structural schematic diagram of an embodiment of a computer device according to the present application. Specific implementation

[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.

[0063] Referring to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an unrelated or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0064] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings.

[0065] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. The terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0066] The user may use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.

[0067] The terminal device 101 may be various electronic devices having a display screen and supporting web browsing. In addition to the laptop computer 1011, the tablet computer 1012, or the mobile phone 1013, the terminal device 101 may also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop portable computer, a desktop computer, etc.

[0068] The server 103 may be a server providing various services, such as a background server supporting the pages displayed on the terminal device 101.

[0069] It should be noted that the risk determination method provided by the embodiments of this application is generally executed by the server / terminal device. Correspondingly, the risk determination device is generally disposed in the server / terminal device.

[0070] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in

[0071] are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers. Figure 2 Continuing to refer to

[0072] Step S10, obtain system log information and a preset risk vocabulary library, and identify risk vocabulary in the system log information according to the preset risk vocabulary library to obtain risk vocabulary information;

[0073] In this embodiment, the system log information is a file or data stream that records various events, errors, warnings, etc. generated during the operation of the system. This information includes a timestamp, a log level, log content, etc. The preset risk vocabulary library is a set containing keywords or phrases related to risks, and this part of the vocabulary is associated with the security risks, threats, or vulnerabilities of the system. The risk vocabulary information is the risk vocabulary identified after comparing the preset risk vocabulary library with the system log information.

[0074] Step S20, obtain user operation logs and a standard operation sequence template, and identify abnormal sequences according to the user operation logs and the standard operation sequence template to obtain an abnormal sequence result;

[0075] In this embodiment, the user operation logs record all the operation processes and operation results of the user in the system, which can be automatically recorded by the system and retrieved from the recorded database. The standard operation sequence template defines the sequence that the user should follow when performing normal operations in the system. These templates are usually designed based on business logic and system functions and are used to compare with the user operation logs to identify abnormal sequences. The abnormal sequence result is the sequence identified as abnormal after comparing and identifying according to the user operation logs and the standard operation sequence template. The abnormal sequence result may include detailed information of the abnormal sequence, the type of abnormality, the time when the abnormality occurred, etc.

[0076] Step S30, perform correlation analysis on the risk vocabulary information and the abnormal sequence result to obtain correlation degree information;

[0077] In this embodiment, correlation analysis is performed according to the correlation rules between the risk vocabulary information and the abnormal sequence result. The correlation rules can be obtained based on a preset correlation mining algorithm for the risk vocabulary information and the abnormal sequence result. The preset correlation mining algorithm can adopt the AprioriAll algorithm. The correlation degree information includes indicators such as the premise, conclusion, support degree, confidence degree, lift degree, and correlation degree of the rule.

[0078] Step S40, generate an effective weight vector according to the correlation degree information, and cluster the risk vocabulary information and the abnormal sequence result according to the effective weight vector to obtain a risk abnormal clustering cluster;

[0079] In this embodiment, the effective weight vector is used to represent the importance of different risk words and abnormal sequences in cluster analysis. This vector is constructed by quantifying the correlation information (i.e., the degree of correlation between risk words and abnormal sequences). Each weight vector corresponds to a specific risk word or abnormal sequence, and the magnitude of the weight vector reflects the influence of this element in the clustering process. The risk anomaly clustering cluster is the result of cluster analysis, which represents a set of similar risk words and abnormal sequences. Among them, the risk anomaly clustering cluster can include risk words related to fraud behavior and abnormal operation sequences, as well as abnormal behaviors caused by system failures or misconfigurations.

[0080] Step S50: Construct a risk determination decision tree based on the risk anomaly clustering cluster, and generate a risk determination knowledge base according to the risk determination rules of the risk determination decision tree;

[0081] In this embodiment, a risk determination decision tree is constructed based on the risk anomaly clustering cluster using the decision tree algorithm. Among them, the internal nodes of the risk determination decision tree model represent the features corresponding to the operation sequences, and the leaf nodes represent the risk determination results. The risk determination knowledge base includes a series of risk determination rules, which are extracted from the various paths of the risk determination decision tree. Among them, each rule describes a specific combination of conditions (such as the frequency of risk words, the type of abnormal sequences, etc.), and gives the risk type or label to be determined under these conditions. In addition to the risk determination rules, the risk determination knowledge base also includes the feature definitions and risk type labels for risk determination. The feature definitions are the key information extracted from the risk anomaly clustering cluster and are used for testing and comparison in the rules, including the name, data type, value range, etc. of the features. The risk type labels can include fraud risk, system risk, operation risk, etc.

[0082] Step S60: Obtain real-time system logs and real-time operation logs, and perform risk determination on the real-time system logs and the real-time operation logs according to the risk determination knowledge base to obtain a risk determination result.

[0083] In this embodiment, the real-time system log is the log information automatically generated by the system during its operation, recording the system status, events, and errors, including information such as timestamps, log levels, log sources, and log contents. The real-time operation log is the log information recording the operation behaviors performed by users in the system, including information such as timestamps, user IDs, operation types, operation results, and associated data. The risk determination result is the result obtained by analyzing and judging the log data according to the information in the real-time system log and the real-time operation log, combined with the rules in the risk determination knowledge base, including information such as risk type, risk level, risk description, and associated logs.

[0084] In this embodiment, the above method can be applied to a medical service system, where the operation logs are monitored by generating a risk determination knowledge base to perform corresponding risk determinations. Specifically, in this embodiment, the medical service system can be one or more of a medical insurance system and a disease insurance system. System logs and user operation logs are log data recorded in the medical system. System log information, a preset risk vocabulary library, user operation logs, and a standard operation sequence template are all stored in the medical insurance system and the disease insurance system and are obtained from the databases of the above systems. The risk determination result is generated by the above systems through the method of this embodiment and is stored in the databases of the above systems.

[0085] In this embodiment, system log information and a preset risk vocabulary library are obtained, and risk vocabulary identification is performed on the system log information according to the preset risk vocabulary library to obtain risk vocabulary information; user operation logs and a standard operation sequence template are obtained, and abnormal sequence identification is performed according to the user operation logs and the standard operation sequence template to obtain an abnormal sequence result; correlation analysis is performed on the risk vocabulary information and the abnormal sequence result to obtain correlation degree information; an effective weight vector is generated according to the correlation degree information, and clustering is performed on the risk vocabulary information and the abnormal sequence result according to the effective weight vector to obtain a risk abnormal clustering cluster; a risk determination decision tree is constructed according to the risk abnormal clustering cluster, and a risk determination knowledge base is generated according to the risk determination rules of the risk determination decision tree; real-time system logs and real-time operation logs are obtained, and risk determinations are performed on the real-time system logs and the real-time operation logs according to the risk determination knowledge base to obtain risk determination results. Thus, effective risk determinations can be made on the operation logs of users on the system, and the accuracy and efficiency of risk determinations are improved.

[0086] Reference Figure 3 , in some alternative implementation manners of this embodiment, step S10 includes the following steps:

[0087] Step S101, obtain an information extraction identifier and a vocabulary library extraction identifier, and extract the system log information and the preset risk vocabulary library from the database according to the information extraction identifier and the vocabulary library extraction identifier;

[0088] In this embodiment, the information extraction identifier is the unique identifier information corresponding to the system log information in the database, and traversal matching is performed in the database through the information extraction identifier to extract the corresponding system log information. The vocabulary library extraction identifier is the unique identifier information corresponding to the preset risk vocabulary library in the database, and traversal matching is performed in the database through the vocabulary library extraction identifier to extract the corresponding preset risk vocabulary library.

[0089] Step S102, perform text preprocessing on the system log information to obtain standard system log information;

[0090] In this embodiment, the text preprocessing of the system log information includes steps such as removing unnecessary characters (such as timestamps, log levels, etc.), word segmentation, and removing stop words. By performing the above text preprocessing on the system log information, standard system log information can be effectively obtained.

[0091] Step S103, use the preset risk vocabulary library as a feature dictionary, and match the standard system log information and the feature dictionary according to the string matching algorithm to obtain risk vocabulary information.

[0092] In this embodiment, the preset risk vocabulary library contains all the words considered to be risk-related. This part of the vocabulary is used to find potential risk points in the standard system log information. The string matching algorithm is an algorithm for finding specific patterns (risk vocabulary in this embodiment) in the text. Specifically, in this embodiment, the brute-force matching algorithm is used to match the standard system log information and the feature dictionary to identify the risk vocabulary information in the standard system log information.

[0093] In this embodiment, by obtaining the information extraction identifier and the vocabulary library extraction identifier, the system log information and the preset risk vocabulary library are extracted from the database according to the information extraction identifier and the vocabulary library extraction identifier; perform text preprocessing on the system log information to obtain standard system log information; use the preset risk vocabulary library as a feature dictionary, and match the standard system log information and the feature dictionary according to the string matching algorithm, so as to effectively obtain the risk vocabulary information identified according to the system log information and the preset risk vocabulary library, which is convenient for subsequent correlation analysis processing.

[0094] Reference Figure 4 , in some alternative implementation manners of this embodiment, step S20 includes the following steps:

[0095] Step S201, obtain the log extraction identifier and the template extraction identifier, and extract the user operation log and the standard operation sequence template from the database according to the log extraction identifier and the template extraction identifier;

[0096] In this embodiment, the log extraction identifier is the unique identifier information corresponding to the user operation log in the database. By traversing and matching with the log extraction identifier in the database, the corresponding user operation log can be extracted. The template extraction identifier is the unique identifier information corresponding to the standard operation sequence template in the database. By traversing and matching with the template extraction identifier in the database, the corresponding standard operation sequence template can be extracted.

[0097] Step S202: Analyze the user operation log based on the sequence pattern mining algorithm to obtain the actual operation sequence information;

[0098] In this embodiment, the sequence pattern mining algorithm can adopt the AprioriAll algorithm to analyze the user operation log. The steps of analyzing the user operation log based on the sequence pattern mining algorithm include: determining the threshold of the minimum support (min_support), which will determine which item sets are considered frequent, that is, they appear frequently enough in the transaction database. Using the AprioriAll algorithm, start from the single-item set (a set containing only one item) and gradually generate item sets containing more items. In each step, calculate the support of the current item set and compare it with the set minimum support threshold. Retain those item sets whose support exceeds the threshold as frequent item sets and continue to generate item sets containing more items. Repeat the above process until no new frequent item sets can be generated. Identify the actual operation sequence: Among the generated frequent item sets, identify the item sets representing the actual operation sequence of the user. This sequence includes a series of operations executed by the user in order, such as "log in to the system" -> "query orders" -> "submit orders". Analyze the user operation log through the above sequence pattern mining algorithm to obtain the actual operation sequence information presented in sequence form.

[0099] Step S203: Calculate the edit distance between the actual operation sequence information and the standard operation sequence template to obtain the edit distance value;

[0100] In this embodiment, the edit distance is a measurement method to measure the degree of difference between two strings, which considers the costs required for operations such as insertion, deletion, and replacement. The calculation of the edit distance can be achieved by calculating the Levenshtein distance. The calculated edit distance value is used to reflect the degree of difference between the actual operation sequence and the standard operation sequence. A smaller edit distance value indicates that the actual operation is closer to the standard operation, while a larger value indicates a larger deviation.

[0101] Step S204: Determine whether the edit distance value is greater than or equal to the preset distance threshold;

[0102] In this embodiment, the preset distance threshold is the standard threshold for determining whether the actual operation sequence information is an abnormal sequence, which can be obtained by analyzing historical data. By comparing the edit distance value with the preset distance threshold, it can effectively determine whether the edit distance value is greater than or equal to the preset distance threshold.

[0103] Step S205, if the edit distance value is greater than or equal to the preset distance threshold, then regard the actual operation sequence information as an abnormal sequence, and output the abnormal sequence result;

[0104] In this embodiment, the abnormal sequence is marked by adding a first identifier corresponding to the abnormal sequence to the operation sequence corresponding to the actual operation sequence information, and the marked abnormal sequences are combined and output as the abnormal sequence result.

[0105] Step S206, if the edit distance value is less than the preset distance threshold, then regard the actual operation sequence information as a non-abnormal sequence, and output the non-abnormal sequence result.

[0106] In this embodiment, the abnormal sequence is marked by adding a second identifier corresponding to the non-abnormal sequence to the operation sequence corresponding to the actual operation sequence information, and the marked abnormal sequences are combined and output as the non-abnormal sequence result.

[0107] In this embodiment, by obtaining a log extraction identifier and a template extraction identifier, the user operation log and the standard operation sequence template are extracted from the database according to the log extraction identifier and the template extraction identifier; the user operation log is analyzed based on a sequence pattern mining algorithm to obtain actual operation sequence information; an edit distance calculation is performed on the actual operation sequence information and the standard operation sequence template to obtain an edit distance value; it is judged whether the edit distance value is greater than or equal to a preset distance threshold; if the edit distance value is greater than or equal to the preset distance threshold, then regard the actual operation sequence information as an abnormal sequence, and output the abnormal sequence result; if the edit distance value is less than the preset distance threshold, then regard the actual operation sequence information as a non-abnormal sequence, and output the non-abnormal sequence result. Thus, it is effectively realized to identify the abnormal sequence in the user operation log according to the standard operation sequence template, so as to facilitate subsequent correlation analysis and processing.

[0108] Reference Figure 5 , in some optional implementation manners of this embodiment, step S30 includes the following steps:

[0109] Step S301, perform a correlation analysis on the risk vocabulary information and the abnormal sequence result based on a preset correlation rule mining algorithm to obtain correlation rule information;

[0110] In this embodiment, the preset association rule mining algorithm can adopt the AprioriAll algorithm to perform association analysis on the risk vocabulary information and the abnormal sequence results. The steps of performing association analysis based on the AprioriAll algorithm include setting parameters: Minimum Support: Set a threshold for screening frequent item sets. Only the item sets with a support greater than or equal to this threshold are considered frequent. Minimum Confidence: Set a threshold for screening association rules. Only the rules with a confidence greater than or equal to this threshold are considered valid. Generate frequent 1-item sets: Calculate the support of each risk vocabulary information and abnormal sequence result, and screen out the frequent 1-item sets. Iteratively generate frequent k-item sets (k>1): Use the frequent k-1 item sets to generate candidate k-item sets. Calculate the support of the candidate k-item sets, and screen out the frequent k-item sets. Repeat this process until no new frequent item sets can be generated. For each frequent item set, generate all possible association rules. Calculate the confidence of each rule, and screen out the rules with a confidence greater than or equal to the minimum confidence threshold. Perform association analysis on the risk vocabulary information and the abnormal sequence results through the above preset association rule mining algorithm to effectively obtain the association rule information.

[0111] Step S302, calculate the support, confidence, and lift of the association rule information, and screen the association rule information according to the support, the confidence, and the lift to obtain valid association rules;

[0112] In this embodiment, the support is obtained by calculating the frequency of each association rule appearing in the dataset. Among them, the higher the support, the more common the rule is in the dataset. The confidence is obtained by calculating the probability that the conclusion (such as a certain abnormal sequence) also appears when the premise of the rule (such as a certain risk vocabulary) appears. Among them, the higher the confidence, the stronger the prediction ability of the premise of the rule for the conclusion. The lift is obtained by calculating the promotion effect of the premise of the rule on the appearance of the conclusion. Among them, a lift greater than 1 indicates a positive correlation between the premise and the conclusion, a lift equal to 1 indicates no correlation, and a lift less than 1 indicates a negative correlation. Screening is performed by setting the support threshold, confidence threshold, and lift threshold corresponding to the support, confidence, and lift. When the support, confidence, and lift of the association rule information are all greater than or equal to the corresponding support threshold, confidence threshold, and lift threshold, the association rule information is used as a valid association rule.

[0113] Step S303, calculate the association degree between the risk vocabulary information and the abnormal sequence result according to the valid association rule, and generate the corresponding association degree information.

[0114] In this embodiment, the comprehensive correlation index can be obtained by calculating the weighted average of the support, confidence, and lift of the effective association rules, and the correlation between the risk vocabulary information and the abnormal sequence result is mapped by the correlation index. After calculating the correlation, each effective association rule and its corresponding correlation index can be sorted out to generate correlation information.

[0115] In this embodiment, the risk vocabulary information and the abnormal sequence result are subjected to association analysis based on a preset association rule mining algorithm to obtain association rule information; the support, confidence, and lift of the association rule information are calculated, and the association rule information is screened according to the support, the confidence, and the lift to obtain effective association rules; the correlation between the risk vocabulary information and the abnormal sequence result is calculated according to the effective association rules, and the corresponding correlation information is generated. Thus, the correlation information representing the association correspondence between the risk vocabulary information and the abnormal sequence result can be effectively obtained to facilitate subsequent clustering processing.

[0116] Continuing to refer to Figure 6 , in some alternative implementation manners of this embodiment, step S40 includes the following steps:

[0117] Step S401, screening out the risk vocabulary and abnormal operation sequences with positive association from the risk vocabulary information and the abnormal sequence result according to the correlation information;

[0118] In this embodiment, according to the correlation information, the risk vocabulary and abnormal operation sequences with positive association (that is, the lift is greater than 1) are screened out from the risk vocabulary information and the abnormal sequence result. The risk vocabulary and abnormal operation sequences with positive association in this part indicate that there is a certain degree of positive correlation causal relationship or correlation between the occurrence of the risk vocabulary and the abnormal operation sequence.

[0119] Step S402, calculating the weights of the risk vocabulary and the abnormal operation sequences according to the information gain algorithm to obtain the effective weight vector.

[0120] In this embodiment, information gain is a method for measuring the importance of features, which is calculated based on the amount of information provided by features in the dataset. For each risk vocabulary and abnormal operation sequence, the information gain for distinguishing different categories (such as normal and abnormal) can be calculated. Then, according to the results of the information gain, a weight is assigned to each risk vocabulary and abnormal operation sequence to form an effective weight vector. Among them, the greater the weight, the higher the importance of the feature in distinguishing different categories.

[0121] Step S403, performing feature extraction on the risk vocabulary and the abnormal operation sequences to obtain a risk vocabulary feature vector and an operation sequence feature vector;

[0122] In this embodiment, the feature extraction of risk words is achieved by mapping each risk word to a unique index (such as using the bag-of-words model or the TF-IDF method) to obtain the risk word feature vector. The operation sequence feature vector is obtained by converting the abnormal operation sequence into a feature vector of a series of operation codes or operation types. Among them, if the operation sequence contains multiple steps, each step can be considered as a feature and encoded using the one-hot encoding or embedding method.

[0123] Step S404: Based on a preset clustering algorithm, perform clustering analysis according to the risk word feature vector, the operation sequence feature vector, and the effective weight vector to obtain the risk anomaly clustering clusters.

[0124] In this embodiment, the preset clustering method can adopt the K-means clustering method. By splicing and fusing the risk word feature vector and the operation sequence feature vector into a unified feature vector, and then inputting this unified feature vector into the K-means clustering method, the effective weight vector is used as a weighting factor when calculating the distance or similarity in the clustering process. For example, when calculating the Euclidean distance between two samples, a weighted distance formula can be used, where the contribution of each feature is adjusted according to its weight. Thus, the clustering analysis operation is effectively implemented to obtain the risk anomaly clustering clusters, and each risk anomaly clustering cluster represents a specific risk pattern or abnormal behavior.

[0125] In this embodiment, the positive-associated risk words and abnormal operation sequences are screened out from the risk word information and the abnormal sequence result according to the correlation information; the weights of the risk words and the abnormal operation sequences are calculated according to the information gain algorithm to obtain the effective weight vector; the feature extraction of the risk words and the abnormal operation sequences is performed to obtain the risk word feature vector and the operation sequence feature vector; based on a preset clustering algorithm, clustering analysis is performed according to the risk word feature vector, the operation sequence feature vector, and the effective weight vector, so as to obtain the risk anomaly clustering clusters formed by effective clustering according to the weights and features corresponding to the risk word information and the abnormal sequence result, which is convenient for constructing the risk determination decision tree in the subsequent step.

[0126] Continue to refer to Figure 7 , in some alternative implementation manners of this embodiment, step S50 includes the following steps:

[0127] Step S501: Based on the decision tree algorithm, construct the risk determination decision tree for the risk anomaly clustering clusters;

[0128] In this embodiment, the decision tree algorithm can be the ID3 algorithm. The steps of constructing a risk determination decision tree for the risk anomaly clustering cluster based on the decision tree algorithm include: creating an empty decision tree structure. Selecting the operation sequence feature with the largest information gain (or Gini coefficient, entropy, etc.) in the data set as the splitting feature of the current node. Taking the selected feature as the current internal node. Dividing the data set into multiple subsets according to the feature values of the current node. Repeating the above process for each subset until the stopping condition is met (such as the subset purity reaches the threshold, the maximum tree depth is reached, etc.). When no further splitting is possible, setting the current node as a leaf node and assigning a risk determination result label. Through the above decision tree algorithm, the decision tree is constructed to obtain the risk determination decision tree.

[0129] Step S502: Obtaining the path from the root node to the leaf node of the risk determination decision tree as the risk determination rule to obtain a risk determination rule set;

[0130] In this embodiment, the steps of obtaining the risk determination rule include: starting from the root node of the decision tree, traversing the entire tree until the leaf node. For each path from the root node to the leaf node, record the feature tests and test results along the way. Among them, the leaf node usually represents a category (such as "high risk", "low risk", etc.), so the end point of the path is the risk determination result. Converting each path into a risk determination rule, and the form of the rule can be a statement of "if... then...", where the "if" part contains the feature tests and test results, and the "then" part contains the risk determination result. After obtaining all the risk determination rules from the risk determination decision tree, organize the risk determination rules to obtain a risk determination rule set.

[0131] Step S503: Storing the risk determination rule set in a predetermined knowledge base for preservation to obtain the risk determination knowledge base.

[0132] In this embodiment, the predetermined knowledge base is a preset knowledge base for storing risk determination rules, and this predetermined knowledge base can be a database, a file system, or other storage mechanisms. When the risk determination rule set is stored in the predetermined knowledge base, the current knowledge base storing the risk determination rule set is used as the risk determination knowledge base.

[0133] In this embodiment, by constructing the risk determination decision tree for the risk anomaly clustering cluster based on the decision tree algorithm; obtaining the path from the root node to the leaf node of the risk determination decision tree as the risk determination rule to obtain a risk determination rule set; storing the risk determination rule set in a predetermined knowledge base for preservation, a risk determination knowledge base that can provide an effective risk determination reference is obtained, so as to facilitate subsequent risk determination processing.

[0134] Continue to refer toFigure 8 , in some alternative implementation manners of this embodiment, step S60 includes the following steps:

[0135] S601, obtain a system log record table and an operation log record table, and extract the real-time system log and the real-time operation log from the system log record table and the operation log record table according to the current timestamp;

[0136] In this embodiment, the system log record table contains log information related to the overall operation status of the system, such as system startup, shutdown, error reporting, performance monitoring, etc. The operation log record table records the specific operations of users in the system, such as login, logout, data modification, transaction execution, etc. It can be obtained by reading the log information saved in the system. The current timestamp is the current date and time, serving as the time reference for extracting real-time logs. By using SQL queries or other database query techniques, log records close to the current timestamp (or within a specified time period, which is from the time of the current timestamp to one minute before in this embodiment) are extracted from the system log record table and the operation log record table. Then, the queried log records are organized into an easily processable format to form the real-time system log and the real-time operation log.

[0137] S602, input the real-time system log and the real-time operation log into the risk determination knowledge base for matching to obtain the risk determination result.

[0138] In this embodiment, before inputting the real-time system log and the real-time operation log into the risk determination knowledge base, preprocessing needs to be performed first. This preprocessing includes extracting key information, converting formats, etc., to ensure their compatibility with the rules in the knowledge base. The preprocessed real-time logs are input into the risk determination knowledge base to traverse the log records and check whether they meet the conditions of any risk determination rules to complete the matching of the rules in the knowledge base. When the matching result is obtained, the corresponding risk determination result is obtained from the knowledge base according to the matching result. The risk determination result includes a risk level (such as "high risk", "medium risk", "low risk"), a risk type (such as "fraudulent behavior", "system anomaly", etc.), and possible risk mitigation measures or suggestions.

[0139] This embodiment effectively obtains the risk determination result for the current real-time operation of the user by obtaining the system log record table and the operation log record table, extracting the real-time system log and the real-time operation log from the system log record table and the operation log record table according to the current timestamp, and inputting the real-time system log and the real-time operation log into the risk determination knowledge base for matching.

[0140] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disc, a Read-Only Memory (ROM), or a Random Access Memory (RAM), etc.

[0141] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0142] Further reference Figure 9 to Figure 1 As an implementation of the method shown above, an embodiment of a risk determination device is provided in this application. This device embodiment corresponds to the method embodiment shown in Figure 1 and can be specifically applied to various electronic devices.

[0143] As Figure 9 shown, the risk determination device 700 described in this embodiment includes: an information acquisition module 701, a sequence recognition module 702, a correlation analysis module 703, an information clustering module 704, a knowledge base generation module 705, and a risk determination module 706. Among them:

[0144] The information acquisition module 701 is used to acquire system log information and a preset risk vocabulary library, and perform risk vocabulary recognition on the system log information according to the preset risk vocabulary library to obtain risk vocabulary information;

[0145] The sequence recognition module 702 is used to acquire user operation logs and a standard operation sequence template, and perform abnormal sequence recognition according to the user operation logs and the standard operation sequence template to obtain an abnormal sequence result;

[0146] The correlation analysis module 703 is used to perform correlation analysis on the risk vocabulary information and the abnormal sequence result to obtain correlation degree information;

[0147] An information clustering module 704, configured to generate an effective weight vector according to the relevance information, and cluster the risk vocabulary information and the abnormal sequence result according to the effective weight vector to obtain a risk abnormal clustering cluster;

[0148] A knowledge base generation module 705, configured to construct a risk determination decision tree according to the risk abnormal clustering cluster, and generate a risk determination knowledge base according to the risk determination rules of the risk determination decision tree;

[0149] A risk determination module 706, configured to obtain real-time system logs and real-time operation logs, and perform risk determination on the real-time system logs and the real-time operation logs according to the risk determination knowledge base to obtain a risk determination result.

[0150] In this embodiment, by adopting the above risk determination device, it is possible to obtain change operation information and migration operation information, encapsulate the change operation information and the migration operation information into a data change transaction; execute the data change transaction, and record transaction execution information; detect whether the transaction execution information contains abnormal execution information; if the transaction execution information contains the abnormal execution information, perform Schema reverse restoration and migration data rollback respectively according to the transaction execution information to obtain a restored database; perform consistency verification on the restored database according to a preset consistency rule, and after the consistency verification passes, mark the restored database as a repaired state. Thus, it effectively realizes effective risk determination operations on the database, improves the efficiency and security of risk determination, and ensures the real-time, accurate and effective database data.

[0151] To solve the above technical problems, an embodiment of the present application further provides a computer device. For details, please refer to Figure 10 , Figure 10 which is the basic structural block diagram of the computer device in this embodiment.

[0152] The computer device 8 includes a memory 81, a processor 82, and a network interface 83 that are communicatively connected to each other via a system bus. It should be noted that only the computer device 8 with components 81 - 83 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, those skilled in the art of the present technology can understand that the computer device here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0153] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through means such as a keyboard, a mouse, a remote control, a touchpad, or a voice control device.

[0154] The memory 81 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), random access memories (RAMs), static random access memories (SRAMs), read-only memories (ROMs), electrically erasable programmable read-only memories (EEPROMs), programmable read-only memories (PROMs), magnetic memories, magnetic disks, optical disks, etc. In some embodiments, the memory 81 can be an internal storage unit of the computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the memory 81 can also be an external storage device of the computer device 8, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 8. Of course, the memory 81 can also include both the internal storage unit and the external storage device of the computer device 8. In this embodiment, the memory 81 is generally used to store the operating system and various application software installed on the computer device 8, such as computer-readable instructions of the risk determination method. In addition, the memory 81 can also be used to temporarily store various data that have been output or will be output.

[0155] In some embodiments, the processor 82 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 82 is generally used to control the overall operation of the computer device 8. In this embodiment, the processor 82 is used to run the computer-readable instructions stored in the memory 81 or process data, such as running the computer-readable instructions of the risk determination method.

[0156] The network interface 83 may include a wireless network interface or a wired network interface, and the network interface 83 is generally used to establish a communication connection between the computer device 8 and other electronic devices.

[0157] In this embodiment, by using the above computer device, system log information and a preset risk vocabulary library can be obtained, and risk vocabulary recognition is performed on the system log information according to the preset risk vocabulary library to obtain risk vocabulary information; user operation logs and a standard operation sequence template are obtained, and abnormal sequence recognition is performed according to the user operation logs and the standard operation sequence template to obtain an abnormal sequence result; correlation analysis is performed on the risk vocabulary information and the abnormal sequence result to obtain correlation degree information; an effective weight vector is generated according to the correlation degree information, and clustering is performed on the risk vocabulary information and the abnormal sequence result according to the effective weight vector to obtain a risk abnormal clustering cluster; a risk determination decision tree is constructed according to the risk abnormal clustering cluster, and a risk determination knowledge base is generated according to the risk determination rules of the risk determination decision tree; real-time system logs and real-time operation logs are obtained, and risk determination is performed on the real-time system logs and the real-time operation logs according to the risk determination knowledge base to obtain a risk determination result. Thus, effective risk determination of the operation logs of the user on the system is realized, and the accuracy and efficiency of risk determination are improved.

[0158] The present application also provides another implementation manner, that is, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to execute the steps of the risk determination method as described above.

[0159] In this embodiment, by using the above-mentioned computer-readable storage medium, system log information and a preset risk vocabulary library can be obtained, and risk vocabulary recognition is performed on the system log information according to the preset risk vocabulary library to obtain risk vocabulary information; user operation logs and a standard operation sequence template are obtained, and abnormal sequence recognition is performed according to the user operation logs and the standard operation sequence template to obtain an abnormal sequence result; correlation analysis is performed on the risk vocabulary information and the abnormal sequence result to obtain correlation degree information; an effective weight vector is generated according to the correlation degree information, and clustering is performed on the risk vocabulary information and the abnormal sequence result according to the effective weight vector to obtain a risk abnormal clustering cluster; a risk determination decision tree is constructed according to the risk abnormal clustering cluster, and a risk determination knowledge base is generated according to the risk determination rules of the risk determination decision tree; real-time system logs and real-time operation logs are obtained, and risk determination is performed on the real-time system logs and the real-time operation logs according to the risk determination knowledge base to obtain a risk determination result. Thus, effective risk determination of the operation logs of users on the system is realized, and the accuracy and efficiency of risk determination are improved.

[0160] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0161] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The preferred embodiments of the present application are shown in the drawings, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields shall be within the scope of the patent protection of the present application by the same token.

[0162] The non-company software tools or components that appear in the embodiments of the present application are only introduced by way of example and do not represent actual use.

Claims

1. A risk determination method, characterized in that: The steps include: Acquire system log information and a preset risk vocabulary library, and perform risk vocabulary recognition on the system log information according to the preset risk vocabulary library to obtain risk vocabulary information; Obtaining a user operation log and a standard operation sequence template, performing abnormal sequence identification according to the user operation log and the standard operation sequence template, and obtaining an abnormal sequence result; Performing correlation analysis on the risk vocabulary information and the abnormal sequence results to obtain correlation information; Generate an effective weight vector according to the association information, and cluster the risk vocabulary information and the abnormal sequence results according to the effective weight vector to obtain a risk abnormality cluster; Constructing a risk determination decision tree according to the risk anomaly clustering clusters, and generating a risk determination knowledge base according to the risk determination rules of the risk determination decision tree; A real-time system log and a real-time operation log are obtained, and risk determination is performed on the real-time system log and the real-time operation log according to the risk determination knowledge base to obtain a risk determination result.

2. The risk determination method according to claim 1, characterized in that: The step of obtaining system log information and a preset risk vocabulary library, and performing risk vocabulary identification on the system log information according to the preset risk vocabulary library to obtain risk vocabulary information specifically includes: Acquire an information extraction identifier and a vocabulary library extraction identifier, and extract the system log information and the preset risk vocabulary library from a database according to the information extraction identifier and the vocabulary library extraction identifier; Performing text preprocessing on the system log information to obtain standard system log information; The preset risk vocabulary library is used as a feature dictionary, and the standard system log information is matched with the feature dictionary according to a string matching algorithm to obtain risk vocabulary information.

3. The risk determination method according to claim 1, characterized in that: The step of obtaining a user operation log and a standard operation sequence template, performing abnormal sequence identification according to the user operation log and the standard operation sequence template, and obtaining an abnormal sequence result specifically includes: Obtaining a log extraction identifier and a template extraction identifier, and extracting the user operation log and the standard operation sequence template from a database according to the log extraction identifier and the template extraction identifier; Analyze the user operation log based on a sequence pattern mining algorithm to obtain actual operation sequence information; Performing edit distance calculation on the actual operation sequence information and the standard operation sequence template to obtain an edit distance value; Determine whether the edit distance value is greater than or equal to a preset distance threshold; If the edit distance value is greater than or equal to the preset distance threshold, the actual operation sequence information is used as an abnormal sequence, and the abnormal sequence result is output.

4. The risk determination method according to claim 1, characterized in that: The step of performing correlation analysis on the risk vocabulary information and the abnormal sequence results to obtain correlation information specifically includes: Performing association analysis on the risk vocabulary information and the abnormal sequence results based on a preset association rule mining algorithm to obtain association rule information; Calculating the support, confidence and lift of the association rule information, and screening the association rule information according to the support, confidence and lift to obtain valid association rules; The association degree between the risk vocabulary information and the abnormal sequence result is calculated according to the effective association rule, and corresponding association degree information is generated.

5. The risk determination method according to claim 1, characterized in that: The step of generating an effective weight vector according to the association information, and clustering the risk vocabulary information and the abnormal sequence results according to the effective weight vector to obtain a risk abnormality cluster specifically includes: Filtering positively associated risk words and abnormal operation sequences from the risk word information and the abnormal sequence results according to the association degree information; Calculating the weights of the risk vocabulary and the abnormal operation sequence according to an information gain algorithm to obtain the effective weight vector; Performing feature extraction on the risk vocabulary and the abnormal operation sequence to obtain a risk vocabulary feature vector and an operation sequence feature vector; Based on a preset clustering algorithm, cluster analysis is performed according to the risk vocabulary feature vector, the operation sequence feature vector, and the effective weight vector to obtain the risk anomaly cluster.

6. The risk determination method according to claim 1, characterized in that: The step of constructing a risk determination decision tree according to the risk anomaly clustering clusters, and generating a risk determination knowledge base according to the risk determination rules of the risk determination decision tree specifically includes: Constructing the risk determination decision tree for the risk anomaly clusters based on a decision tree algorithm; Obtaining a path from a root node to a leaf node of the risk determination decision tree as the risk determination rule to obtain a risk determination rule set; The risk determination rule set is stored in a predetermined knowledge base for preservation to obtain the risk determination knowledge base.

7. The risk determination method according to claim 1, characterized in that: The step of obtaining the real-time system log and the real-time operation log, performing risk determination on the real-time system log and the real-time operation log according to the risk determination knowledge base, and obtaining the risk determination result specifically includes: Obtain a system log record table and an operation log record table, and extract the real-time system log and the real-time operation log from the system log record table and the operation log record table according to a current timestamp; The real-time system log and the real-time operation log are input into the risk determination knowledge base for matching to obtain the risk determination result.

8. A risk determination device, characterized in that: include: An information acquisition module is used to acquire system log information and a preset risk vocabulary library, and perform risk vocabulary recognition on the system log information according to the preset risk vocabulary library to obtain risk vocabulary information; A sequence identification module is used to obtain a user operation log and a standard operation sequence template, perform abnormal sequence identification according to the user operation log and the standard operation sequence template, and obtain an abnormal sequence result; An association analysis module, used to perform association analysis on the risk vocabulary information and the abnormal sequence results to obtain association degree information; An information clustering module, used to generate an effective weight vector according to the correlation information, and cluster the risk vocabulary information and the abnormal sequence results according to the effective weight vector to obtain a risk abnormality cluster; A knowledge base generation module, used to construct a risk determination decision tree according to the risk anomaly clustering clusters, and generate a risk determination knowledge base according to the risk determination rules of the risk determination decision tree; The risk determination module is used to obtain the real-time system log and the real-time operation log, perform risk determination on the real-time system log and the real-time operation log according to the risk determination knowledge base, and obtain a risk determination result.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the risk determination method according to any one of claims 1 to 7 when executing the computer-readable instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the risk determination method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Material whole-process dynamic management and control system and method

    CN120931209A

  • A material whole-process dynamic management and control system and method

    CN120931209B