A Real-time Reconciliation Method and System Based on a Rule Engine
By calculating the dimension importance of reconciliation data and correcting the information gain of decision tree, a more scientific and flexible rule engine is built, which solves the problems of difficulty and insufficient flexibility in maintaining the existing rule engine, and achieves more efficient reconciliation data abnormal identification and rule maintenance.
Patent Information
- Application Number
- CN202510405618.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-02
AI Technical Summary
In actual applications, existing rule engines have problems such as difficult rules maintenance, insufficient flexibility, strong data dependence, and manual rules writing are prone to errors, resulting in high maintenance costs of reconciliation data, which limits the process of supply chain service companies to improve reconciliation efficiency.
By calculating the degree of discreteness and abnormal sensitivity of each dimension of the reconciliation data, the importance of each dimension is determined, and the information gain of the decision tree node is corrected based on this, a decision tree is generated and a rule engine for reconciliation data is constructed.
It improves the scientific construction of the rule engine, reduces the difficulty of maintaining rules, improves the accuracy of abnormal identification of reconciliation data, reduces maintenance costs, and improves the flexibility of the rule engine.
Smart Images

Figure CN119904324B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a real-time reconciliation method and system based on a rule engine. Background Art
[0002] Since the business system of supply chain service enterprises covers multiple links such as procurement, logistics, and cross-border, its unique reconciliation scenarios present significant characteristics such as high-frequency transactions, multiple participants, and complex business rules; this complex business model causes enterprises to face practical pain points such as high labor costs and great operation difficulties during the reconciliation process. Therefore, enterprises usually adopt a rule engine to improve the level of reconciliation automation.
[0003] Chinese patent document with publication number CN117350880B discloses a full-link reconciliation method, device, equipment and medium based on the snowflake algorithm, including: responding to a cross-border business reconciliation request and obtaining bill source data from the business side; parsing the obtained bill source data and converting it into target bill data; calling a reconciliation rule engine to obtain corresponding reconciliation rules and select corresponding reconciliation algorithms, and performing a full-link reconciliation task on the target bill data to obtain a reconciliation result; transmitting the data with unequal reconciliation to the corresponding exception handling system for verification and processing; generating a unique target ID based on the optimized snowflake algorithm to associate the corresponding bills with the reconciliation result, automatically penetrating the full-link data of the information flow and the capital flow according to the unique target ID in the reconciliation result, and tracing the whole process of the corresponding bill flowing between the capital flow and the information flow according to the unique target ID to achieve a complete reconciliation closed-loop, improving the reconciliation work efficiency and data accuracy of the full-link cross-border business.
[0004] However, the existing rule engines have the following main defects in practical applications: First, the rule maintenance is difficult, and the change of business rules requires frequent intervention and modification by technical personnel; second, the system flexibility is insufficient, and it is difficult to quickly adapt to the changing reconciliation requirements; third, the dependence on data is too strong, resulting in limited applicability of the rule engine; and manual rule writing is prone to errors, and the rules need to be frequently modified when the business changes, making the rule engine for reconciliation data have a high maintenance cost, restricting the process of supply chain service enterprises to improve the reconciliation efficiency. Summary of the Invention
[0005] In order to solve the problems that the construction of the rule engine has defects such as difficult rule maintenance, insufficient flexibility, and strong data dependence; and manual rule writing is prone to errors, and the rules need to be frequently modified when the business changes, making the rule engine for reconciliation have a high maintenance cost, restricting the process of supply chain service enterprises to improve the reconciliation efficiency, the present invention provides a real-time reconciliation method and system based on a rule engine.
[0006] First aspect, the present invention provides a real-time reconciliation method based on a rule engine, adopting the following technical solution:
[0007] A real-time reconciliation method based on a rule engine includes: obtaining reconciliation data of an enterprise including multiple dimensions, and determining normal reconciliation data and abnormal reconciliation data according to the results of each reconciliation data; for each numerical dimension among the multiple dimensions of the reconciliation data, determining the data dispersion degree of the numerical dimension according to the standard deviation, interquartile range, and maximum value of the reconciliation data of the numerical dimension; determining the abnormal sensitivity of the numerical dimension according to the means of the normal reconciliation data and the abnormal reconciliation data in the numerical dimension; for each text dimension among the multiple dimensions of the reconciliation data, determining the data dispersion degree of the text dimension according to the data dispersion degree of the numerical dimension, the information entropy of the reconciliation data of the text dimension, the minimum value, and the cosine similarity between the sentence vectors of the reconciliation data; determining the abnormal sensitivity of the text dimension according to the cosine similarity between the means of the sentence vectors of the normal reconciliation data and the abnormal reconciliation data in the text dimension; recording the product of the data dispersion degree and the abnormal sensitivity of each dimension as the importance of each dimension; in the process of constructing a rule engine for the reconciliation data using a decision tree, correcting the information gain of each node of the decision tree according to the importance, and obtaining the corrected information gain for data partitioning of each node using each dimension; based on the corrected information gain, generating a decision tree and constructing a rule engine for the reconciliation data.
[0008] The present invention calculates the dispersion degree of each dimension of the reconciliation data. The dispersion degree can intuitively reveal which dimensions are more potentially valuable in distinguishing normal and abnormal, providing a basis for subsequent adjustment according to the importance of the dimensions; by combining the dispersion degree and abnormal sensitivity of each dimension to obtain the importance of each dimension, it is possible to evaluate the contribution of each dimension in distinguishing normal and abnormal data. Through the importance, the feature weights can be automatically adjusted, making the generation of the decision tree more in line with the requirements of the rule engine for the actual reconciliation business scenario; by adjusting the calculation of the information gain through the importance of each dimension, generating a decision tree through the corrected information gain, and constructing a rule engine for the reconciliation data according to the decision tree. Introducing the importance of the dimension on the basis of the traditional information gain can make the decision tree pay more attention to those dimensions that are more sensitive to the detection of reconciliation anomalies when selecting split attributes, making the generated rules closer to the actual reconciliation scenario, and thus improving the accuracy of abnormal identification of the reconciliation data.
[0009] Further, the reconciliation data includes a unique identifier, a transaction amount, a transaction status, a timestamp, and a keyword field.
[0010] Further, the data dispersion degree of the numerical dimension satisfies: ; where is the data dispersion degree of the th numerical dimension, is the standard deviation of the reconciliation data for the th numerical dimension, is the interquartile range of the reconciliation data for the th numerical dimension, is the maximum value of the reconciliation data for the th numerical dimension.
[0011] The present invention comprehensively considers the standard deviation, interquartile range and maximum value, can comprehensively evaluate the dispersion degree of the reconciliation data, and reflects the volatility, variability and centrality of the data; by quantifying the dispersion degree of the numerical dimension, it can more sensitively capture abnormal data points. When the dispersion degree significantly deviates from the normal range, the system can identify and alarm in time, improving the accuracy and real-time performance of the reconciliation.
[0012] Furthermore, the abnormal sensitivity of the numerical dimension satisfies: ; where is the abnormal sensitivity of the th numerical dimension, is the mean value of the normal reconciliation data in the th numerical dimension, is the mean value of the abnormal reconciliation data in the th numerical dimension, is the maximum value function, is the absolute value symbol.
[0013] The present invention can quantify the significance of the abnormality by calculating the mean difference between the normal and abnormal reconciliation data. A high sensitivity means that there is a large difference between the abnormal data and the normal data, which helps to identify the abnormality more quickly and accurately; using the ratio of the absolute difference of the means to the maximum mean can effectively eliminate the difference in data scale, making the sensitivities between different numerical dimensions comparable, and helping to identify the importance of abnormal data in different dimensions in the overall dataset.
[0014] Furthermore, the sentence vector is obtained by transforming the reconciliation data in the text dimension using Sentence - BERT.
[0015] Furthermore, the data dispersion degree of the text dimension satisfies: ; where is the data dispersion degree of the th text dimension, is the minimum value among the data dispersion degrees of all numerical dimensions, is the information entropy of the reconciliation data of the th text dimension, is the th in the The mean of the cosine similarities between the sentence vectors of one copy of the reconciliation data and the rest of the reconciliation data is the quantity of the reconciliation data is the natural exponential function is the linear normalization function
[0016] The present invention combines the information entropy and the mean of the cosine similarities between sentence vectors, which can comprehensively reflect the diversity and complexity of text data, and helps to deeply analyze the characteristics of the reconciliation data; by introducing the information entropy, the diversity of the text dimension can be quantified; the linear normalization processing of the cosine similarity of the sentence vectors ensures that the influence of the extreme similarity on the result is reasonably controlled, making the calculated degree of dispersion more robust
[0017] Further, the abnormal sensitivity of the text dimension satisfies ; where is the abnormal sensitivity of the th text dimension is the cosine similarity between the means of the sentence vectors of the normal reconciliation data and the abnormal reconciliation data in the th text dimension is the natural exponential function
[0018] The present invention quantifies the similarity between the normal reconciliation data and the abnormal reconciliation data, making the calculation of the abnormal sensitivity more intuitive and clear, and the value of the sensitivity directly reflects the difference degree between the normal data and the abnormal data; using the natural exponential function to transform the similarity can assign higher weights to the data with lower similarity (larger abnormalities), making it more sensitive to larger deviations in anomaly detection, thereby enhancing the anomaly detection ability of the system; through the exponential processing of the cosine similarity, different degrees of abnormal situations in the text dimension can be effectively captured, enabling the model to more flexibly adapt to data changes
[0019] Further, the corrected information gain satisfies ; where is the corrected information gain for partitioning the data of node using the th dimension is the information entropy of the reconciliation data within node is the number of branch nodes obtained by partitioning the data of node using the th dimension is the information entropy of the th branch node obtained by partitioning the data of node using the th dimension is the node The quantity of the reconciled data included is the pair node Using the th dimension to perform data partitioning, the quantity of the reconciled data included in the th branch node is the th dimension's importance degree is the natural exponential function
[0020] Furthermore, the method for generating a decision tree based on the corrected information gain and constructing a rule engine for the reconciled data includes: replacing the information gain during the generation of the decision tree with the corrected information gain, selecting the dimension corresponding to the maximum value of the corrected information gain for partitioning the branch nodes of the decision tree until all samples on all leaf nodes of the decision tree belong to the same category, completing the generation of the decision tree; using each node of the decision tree as a judgment condition, each branch as a judgment result, and each leaf node as a reconciliation rule, traversing the decision tree, extracting all reconciliation rules and forming a rule set, implementing the construction of the rule engine, and completing the real-time reconciliation based on the rule engine
[0021] In a second aspect, the present invention provides a real-time reconciliation system based on a rule engine, adopting the following technical solution
[0022] A real-time reconciliation system based on a rule engine includes: a processor and a memory, where the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned real-time reconciliation method based on a rule engine is implemented
[0023] By adopting the above technical solution, the above-mentioned real-time reconciliation method based on a rule engine is generated into a computer program and stored in the memory to be loaded and executed by the processor, thereby manufacturing a terminal device according to the memory and the processor, which is convenient to use
[0024] The present invention has the following technical effects
[0025] A rule engine is constructed by calculating the importance of each dimension and modifying the information gain of the decision tree node accordingly, making the construction of the rule engine more scientifically based, reducing the difficulty of rule maintenance compared with the traditional method. Because it is error-prone to write rules manually in the traditional way and the rules need to be frequently modified when the business changes, while the present invention determines the key factors for rule construction based on the characteristics of the data itself, making the rules more stable, reducing the situation of frequently modifying rules due to business changes, and reducing the difficulty of rule maintenance; constructing a rule engine based on the relevant characteristics (such as data dispersion degree, anomaly sensitivity, etc.) of multiple dimensions of data (including numerical dimension and text dimension) can better adapt to different types of data and business scenarios. When the business changes, the rule engine can adjust the rules more flexibly according to the changes in data characteristics, rather than completely relying on manual and frequent rule modification like the traditional method, thus improving the flexibility of the rule engine; in the process of constructing the rule engine, multiple dimension characteristics of the data and the importance of each dimension are comprehensively considered, and the data is analyzed more comprehensively and deeply, rather than simply relying on one or several data characteristics, which can more accurately capture the rules and anomalies in the data, reduce the over-reliance on a single data characteristic, and make the rule engine more robust when processing data; due to reducing the difficulty of rule maintenance and improving the flexibility, the adjustment of rules is more efficient and accurate when the business changes, reducing the workload of manual writing and modifying rules, and thus reducing the maintenance cost of the reconciliation rule engine. Description of the Drawings
[0026] Figure 1 It is the flowchart of the method in an embodiment of the real-time reconciliation method based on a rule engine of the present invention. Detailed Embodiment
[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0028] An embodiment of the present invention discloses a real-time reconciliation method based on a rule engine. Refer to Figure 1 , including steps S1 - S6:
[0029] S1: Obtain the reconciliation data of the enterprise including multiple dimensions, and determine the normal reconciliation data and abnormal reconciliation data according to the results of each reconciliation data.
[0030] It should be noted that the internal reconciliation system of the enterprise collects structured (the structure includes account ID, the difference between the amount in the business system and the amount in the financial system, customer ID, payment status, etc.) transaction flow, bill and log data from multiple business systems such as ERP, financial system, payment gateway, etc. through API interface, database synchronization, ETL tasks and real-time stream processing. After key field matching, data cleaning and standardization preprocessing, reconciliation data containing multiple dimensions is generated. Each reconciliation data corresponds to a label, that is, reconciliation normal or reconciliation abnormal (implementation personnel can also increase the setting of label categories according to the specific implementation situation).
[0031] Specifically, the reconciliation data includes a unique identifier, a transaction amount, a transaction status, a timestamp, and key fields.
[0032] S2: Determine the degree of data dispersion in each dimension.
[0033] It should be noted that data of different dimensions in the reconciliation data have different distribution characteristics, which also makes the data of different dimensions have different discrete degrees. This difference affects the construction of the rule engine under different reconciliation cycles. In order to make the rule engine construction of the reconciliation system more adaptable to the needs of reconciliation work in different periods, this step calculates the discrete degree of each dimension based on the reconciliation data of each dimension.
[0034] For each numerical dimension in the multiple dimensions of the reconciliation data, the data dispersion degree of the numerical dimension is determined according to the standard deviation, interquartile range and maximum value of the reconciliation data of the numerical dimension.
[0035] Specifically, the data discreteness of the numerical dimension satisfies:
[0036] ;
[0037] In the formula, For the The degree of data dispersion in the numerical dimension, For the The standard deviation of the reconciliation data in the numerical dimension, For this The interquartile range of the reconciliation data for the numerical dimension, For the The maximum value of the reconciliation data of the numeric dimension.
[0038] in, The larger the value, the more drastic the fluctuation of the data in this dimension, the more discrete the distribution, and the greater the degree of discreteness of this dimension. The smaller the value, the more stable the data of the dimension is, the more concentrated the distribution is, and the smaller the discrete degree of the dimension is. In order to facilitate calculation, we divide it by The way to Normalization is performed. The larger it is, the wider the distribution range of the regular (non-prominent) data in the data of this dimension, the more discrete the data distribution of this dimension, and the greater the degree of dispersion of this dimension; The smaller it is, the more the data in the data of this dimension is concentrated towards the median of the data of this dimension, the more concentrated the data distribution of this dimension, and the smaller the degree of dispersion of this dimension; here, for the convenience of calculation, it is also divided by in the way of Normalization is performed.
[0039] For each text dimension among multiple dimensions of the reconciliation data, determine the degree of data dispersion of the text dimension according to the degree of data dispersion of the numerical dimension, the information entropy of the reconciliation data of the text dimension, the minimum value, and the cosine similarity between the sentence vectors of the reconciliation data.
[0040] Specifically, the sentence vector is obtained by transforming the reconciliation data of the text dimension using Sentence-BERT.
[0041] Specifically, the degree of data dispersion of the text dimension satisfies:
[0042] ;
[0043] In the formula, is the degree of data dispersion of the th text dimension, is the minimum value among the degrees of data dispersion of all numerical dimensions, is the th information entropy of the reconciliation data of the text dimension, is the mean value of the cosine similarities between the sentence vectors of the th copy of the reconciliation data and the remaining copies of the reconciliation data in the th text dimension, is the quantity of the reconciliation data, is the natural exponential function, is the linear normalization function.
[0044] Among them, since usually in the reconciliation system, the data of the numerical dimension is more discrete than the data of the text dimension, so on the basis of downward adjustment is made for the calculation of the degree of data dispersion of the text dimension, with the purpose of unifying the dimension and making the degree of data dispersion of the text dimension less than that of the numerical dimension. The larger it is, the more diverse the text patterns of the reconciliation data of this text dimension, and the greater the degree of data dispersion of this text dimension; The smaller it is, the more single the text patterns of the reconciliation data of this text dimension, and the smaller the degree of data dispersion of this text dimension. The larger it is, the more similar the reconciliation data of this text dimension is, indicating that the content of the reconciliation data of this text dimension is more unified, and the data dispersion degree of this text dimension is smaller; The smaller it is, the greater the semantic difference of the reconciliation data of this text dimension, the more diverse the content, and the greater the data dispersion degree of this text dimension.
[0045] S3: Determine the anomaly sensitivity of each dimension.
[0046] It should be noted that in the reconciliation data, different dimensions have different sensitivities to identifying abnormal data; this difference in anomaly sensitivity will directly affect the degree of dependence of the rule engine on each dimension when judging whether the data is abnormal; in order to build a more accurate and effective rule engine to enable it to more accurately identify abnormal reconciliation data, this step calculates the anomaly sensitivity of each dimension according to the characteristics of normal and abnormal reconciliation data in different dimension data.
[0047] For each numerical dimension among multiple dimensions of the reconciliation data, determine the anomaly sensitivity of the numerical dimension according to the means of normal and abnormal reconciliation data in the numerical dimension.
[0048] Specifically, the anomaly sensitivity of the numerical dimension satisfies:
[0049] ;
[0050] In the formula, is the anomaly sensitivity of the th numerical dimension, is the mean of normal reconciliation data in the th numerical dimension, is the mean of abnormal reconciliation data in the th numerical dimension, is the maximum value function, is the absolute value symbol.
[0051] Among them, The larger it is, the greater the difference between the mean of normal reconciliation data and the mean of abnormal reconciliation data in the th numerical dimension, which also means that the sensitivity of this dimension to distinguishing normal and abnormal data is higher, and the anomaly sensitivity of this dimension is greater; The smaller it is, the smaller the difference between the means of normal and abnormal data in this numerical dimension, the weaker the ability of this dimension to distinguish normal and abnormal data, and the smaller the anomaly sensitivity of this dimension; by dividing by in this way, is normalized, so that the value of the anomaly sensitivity is within a reasonable range, which is convenient for comparing the anomaly sensitivities between different numerical dimensions.
[0052] For each text dimension among multiple dimensions of reconciliation data, determine the anomaly sensitivity of the text dimension according to the cosine similarity between the mean of the sentence vectors of the normal reconciliation data and the anomaly reconciliation data in the text dimension; assume the number of normal reconciliation data is 2, and the sentence vector of normal reconciliation data 1 in the th text dimension is and the sentence vector of normal reconciliation data 2 in the th text dimension is , then the mean of the sentence vectors of the normal reconciliation data is ; Similarly for the mean of the sentence vectors of the anomaly reconciliation data.
[0053] Specifically, the anomaly sensitivity of the text dimension satisfies:
[0054] ;
[0055] In the formula, is the anomaly sensitivity of the th text dimension, is the cosine similarity between the mean of the sentence vectors of the normal reconciliation data and the anomaly reconciliation data in the th text dimension, is the natural exponential function.
[0056] Among them, The larger it is, the higher the cosine similarity between the mean of the sentence vectors of the normal reconciliation data and the anomaly reconciliation data in the th text dimension, that is, the higher the semantic similarity degree between these two types of data. Then the sensitivity of this text dimension to distinguish normal and abnormal data is lower, and the anomaly sensitivity is smaller; The smaller it is, the lower the cosine similarity between the mean of the sentence vectors of the normal and abnormal data in this text dimension, the greater the semantic difference, the stronger the ability of this text dimension to distinguish normal and abnormal data, and the greater the anomaly sensitivity; Through 's calculation method, the cosine similarity is transformed into a suitable anomaly sensitivity value, so that the anomaly sensitivity of the text dimension can play a reasonable role when constructing a rule engine and better adapt to the characteristics of data in different text dimensions.
[0057] S4: Obtain the importance of each dimension.
[0058] It should be noted that the reconciliation data usually contains multi-dimensional features, but not all features contribute equally to anomaly recognition. Moreover, in the reconciliation under different business cycle scenarios, the importance of dimensions may change dynamically. That is, the data of a certain dimension has a higher degree of importance in the reconciliation task at a certain period, while the importance of the data of this dimension is lower in another period. Treating the reconciliation data of each dimension equally may not be able to well meet the reconciliation requirements. Therefore, in this step, the importance of each dimension of the reconciliation data is calculated according to the dispersion degree and anomaly sensitivity of each dimension.
[0059] Specifically, the importance satisfies:
[0060] ;
[0061] In the formula, is the importance of the th dimension, is the numerical dimension, is the text dimension, and are respectively the dispersion degree and anomaly sensitivity when the th dimension belongs to the numerical dimension, and are respectively the dispersion degree and anomaly sensitivity when the th dimension belongs to the text dimension.
[0062] Among them, or The larger it is, the more differentiated data the th dimension of the reconciliation data has. Through the th dimension of the reconciliation data, the data can be more effectively segmented, improving the classification ability of the decision tree, and thus enabling the rule engine to be better applied to the reconciliation work. Therefore or The larger it is, the greater the importance of the reconciliation data of the th dimension (that is, the importance of the th dimension); or The smaller it is, the smaller the data differentiation degree of the th dimension of the reconciliation data. Through the th dimension of the reconciliation data, the data cannot be effectively segmented. In order to improve the classification ability of the decision tree and enable the rule engine to be better applied to the reconciliation work, the importance of the reconciliation data of the th dimension should be made smaller. or The larger it is, the greater the difference between the reconciliation data of different categories (normal or abnormal) in the th dimension. The The more sensitive the reconciliation data of a dimension is to the progress of the reconciliation work, the higher the importance of the reconciliation data of the th dimension; Or The smaller it is, it indicates that the difference in the reconciliation data of different categories (normal or abnormal) in the th dimension is smaller, and the more insensitive the reconciliation data of the th dimension is to the progress of the reconciliation work, the lower the importance of the reconciliation data of the th dimension.
[0063] S5: Obtain the corrected information gain for data partitioning of each node in the decision tree using each dimension.
[0064] It should be noted that in the construction process of the reconciliation rule engine, the decision tree can be used to extract data splitting rules to meet the reconciliation requirements. However, the traditional decision tree generation process lacks a dynamic adjustment mechanism, that is, it lacks the ability to dynamically adjust the feature importance and is difficult to respond to the changes in the business scenario in real time. By introducing the importance of each dimension to adjust the information gain calculation, the splitting decision can be made more in line with the feature distribution and anomaly recognition requirements in the actual business, and then more accurate and adaptable reconciliation rules can be generated. Therefore, in this step, the information gain calculation is adjusted through the importance of each dimension.
[0065] In the process of constructing the rule engine for reconciliation data using the decision tree, the information gain of each node of the decision tree is corrected according to the importance degree, and the corrected information gain for data partitioning of each node using each dimension is obtained.
[0066] Specifically, the corrected information gain satisfies:
[0067] ;
[0068] In the formula, is the corrected information gain for data partitioning of node using the th dimension, is the information entropy of the reconciliation data within node , is the number of branch nodes obtained by data partitioning of node using the th dimension, is the information entropy of the th branch node obtained by data partitioning of node using the th dimension, is the number of reconciliation data included in node , is for node The quantity of reconciliation data included in the th branch node obtained by partitioning data using the th dimension, is the importance degree of the th dimension, where
[0069] is the calculation method of the existing information gain. The larger this value is, the greater the improvement in the purity of the branch nodes obtained by partitioning the node using the th dimension. The more it should be partitioned by the th dimension; the smaller this value is, the smaller the improvement in the purity of the branch nodes obtained by partitioning the node using the th dimension. The less it should be partitioned by the th dimension. The larger is, the more important the th dimension is for the reconciliation work. The more it should be partitioned by the th dimension to make the construction of the decision tree more adaptable to the requirements of the reconciliation rule engine. Therefore, the larger is, the greater the corrected information gain; the smaller is, the relatively lower the importance degree of the th dimension for the reconciliation work. When partitioning the node, more consideration should be given to the remaining dimensions. Therefore, the smaller
[0070] S6: Generate a decision tree and construct a rule engine for the reconciliation data based on the corrected information gain.
[0071] Specifically, generating a decision tree and constructing a rule engine for the reconciliation data based on the corrected information gain includes:
[0072] Replace the information gain in generating the decision tree with the corrected information gain, and select the dimension corresponding to the maximum value of the corrected information gain for partitioning the branch nodes of the decision tree until all samples on all leaf nodes of the decision tree belong to the same category, completing the generation of the decision tree;
[0073] Take each node of the decision tree as a judgment condition, each branch as a judgment result, and each leaf node as a reconciliation rule. Traverse the decision tree, extract all reconciliation rules and form a rule set to achieve the construction of the rule engine. When new reconciliation data arrives, the rule engine will match the reconciliation data according to the reconciliation rules, thereby realizing automatic reconciliation and completing real-time reconciliation based on the rule engine.
[0074] When performing reconciliation during a new reconciliation cycle, according to the reconciliation data of the new reconciliation cycle, a decision tree is regenerated and a rule engine for the new reconciliation data is constructed.
[0075] An embodiment of the present invention also discloses a real-time reconciliation system based on a rule engine, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a real-time reconciliation method based on a rule engine according to the present invention is implemented.
[0076] The above system further includes other components well-known to those skilled in the art such as a communication bus and a communication interface, and their settings and functions are known in the art, so they will not be described in detail here.
[0077] The above are all preferred embodiments of the present invention. The protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention.
Claims
1. A real-time reconciliation method based on a rule engine, characterized in that: include: Obtain the company's reconciliation data containing multiple dimensions, and determine normal reconciliation data and abnormal reconciliation data based on the results of each reconciliation data; For each numerical dimension in the multiple dimensions of the reconciliation data, determine the data dispersion of the numerical dimension according to the standard deviation, interquartile range and maximum value of the reconciliation data of the numerical dimension; determine the abnormal sensitivity of the numerical dimension according to the mean of the normal reconciliation data and the abnormal reconciliation data in the numerical dimension; For each text dimension in the multiple dimensions of the reconciliation data, the data discreteness of the text dimension is determined according to the data discreteness of the numerical dimension, the information entropy of the reconciliation data in the text dimension, the minimum value, and the cosine similarity between the sentence vectors of the reconciliation data; the abnormal sensitivity of the text dimension is determined according to the cosine similarity between the mean values of the sentence vectors of the normal reconciliation data and the abnormal reconciliation data in the text dimension; The product of the data discreteness and abnormal sensitivity of each dimension is recorded as the importance of each dimension; In the process of constructing a rule engine for reconciliation data using a decision tree, the information gain of each node of the decision tree is corrected according to the importance, so as to obtain a corrected information gain for data division of each node using each dimension; Based on the modified information gain, a decision tree is generated and a rule engine for reconciliation data is constructed.
2. A real-time reconciliation method based on a rule engine according to claim 1, characterized in that: The reconciliation data includes a unique identifier, transaction amount, transaction status, timestamp, and key fields.
3. The real-time reconciliation method based on rule engine according to claim 1, characterized in that: The data discreteness of the numerical dimension satisfies: ; In the formula, For the The degree of data dispersion in the numerical dimension, For the The standard deviation of the reconciliation data in the numerical dimension, For this The interquartile range of the reconciliation data for the numerical dimension, For the The maximum value of the reconciliation data of the numeric dimension.
4. The real-time reconciliation method based on rule engine according to claim 1, characterized in that: The abnormal sensitivity of the numerical dimension satisfies: ; In the formula, For the Abnormal sensitivity of numerical dimensions, For the The average value of normal reconciliation data in the numerical dimension, For the The average of abnormal reconciliation data in the numerical dimension, is the maximum value function, is the absolute value symbol.
5. The real-time reconciliation method based on rule engine according to claim 1, characterized in that: The sentence vector is obtained by transforming the reconciliation data of the text dimension using Sentence-BERT.
6. A real-time reconciliation method based on a rule engine according to claim 5, characterized in that: The data discreteness of the text dimension satisfies: ; In the formula, For the The degree of data dispersion of the text dimension, is the minimum value of the data discreteness of all numerical dimensions, For the The information entropy of the reconciliation data in the text dimension, For the In the text dimension The mean of the cosine similarity between the sentence vectors of the reconciliation data and the rest of the reconciliation data, is the number of reconciliation data, is the natural exponential function, is a linear normalization function.
7. A real-time reconciliation method based on a rule engine according to claim 5, characterized in that: The abnormal sensitivity of the text dimension satisfies: ; In the formula, For the Abnormal sensitivity of the text dimension, For the The cosine similarity between the mean of the sentence vectors of normal reconciliation data and abnormal reconciliation data in the text dimension, is a natural exponential function.
8. The real-time reconciliation method based on rule engine according to claim 1, characterized in that: The modified information gain satisfies: ; In the formula, For the node Use the The modified information gain of data partitioning in dimensions, For Node The information entropy of internal reconciliation data, For the node Use the The number of branch nodes obtained by dividing the data into dimensions, For the node Use the The first The information entropy of the branch nodes is For Node The amount of reconciliation data included, For the node Use the The first The number of reconciliation data contained in each branch node, For the The importance of the dimension, is a natural exponential function.
9. The real-time reconciliation method based on rule engine according to claim 1, characterized in that: The method of generating a decision tree and constructing a rule engine for reconciliation data based on the modified information gain includes: The modified information gain is used to replace the information gain when generating the decision tree, and the dimension corresponding to the maximum value of the modified information gain is selected to divide the branch nodes of the decision tree until all samples on all leaf nodes of the decision tree belong to the same category, thereby completing the generation of the decision tree; Take each node of the decision tree as a judgment condition, each branch as a judgment result, and each leaf node as a reconciliation rule. Traverse the decision tree, extract all reconciliation rules and form a rule set to build a rule engine and complete real-time reconciliation based on the rule engine.
10. A real-time reconciliation system based on a rule engine, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a real-time reconciliation method based on a rule engine according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Full-link reconciliation method, device, equipment and medium based on snowflake algorithm
CN117350880B
Engine rule definition method and device, electronic equipment and storage medium
CN114969123A
Unmanned ship multi-mode perception and decision support method and system based on fuzzy logic and rule engine
CN119294538A