A data tracing analysis system and method based on big data
Through a big data-based data traceability analysis system, the problem of difficulty in traceability of financial statement data in the existing technology is solved, and more accurate rule generation and risk identification is achieved.
Patent Information
- Application Number
- CN202510185852.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-20
AI Technical Summary
It is difficult for existing technology to quickly and accurately trace the source and circulation of financial statement data, especially when new financial businesses and risk models appear, existing rules cannot be updated and adjusted in a timely manner, resulting in insufficient ability to identify potential risks and compliance issues.
A big data traceability analysis system is adopted to collect and preprocess historical data, use convolutional neural networks to build a big data model, extract rules and build special task maps, and combine in-depth priority search and rule matching analysis to achieve traceability and rule optimization of financial statement data.
It realizes in-depth mining of financial data and automatic learning of rules, generates more accurate and comprehensive rules, can quickly and accurately trace the source and circulation of data, and improves the ability to identify potential risks and compliance issues.
Smart Images

Figure CN119722337B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data tracing technology, and in particular to a data tracing analysis system and method based on big data. Background Art
[0002] Traditional rule-making often relies on experience and simple statistical analysis, which makes it difficult to adapt to complex and changing financial business scenarios. For new financial businesses and risk models, existing rules cannot be updated and adjusted in a timely manner, resulting in insufficient ability to identify potential risks and compliance issues. Existing data analysis methods are mostly statistical analysis based on simple indicators, which cannot deeply explore the inherent connections and potential laws between data. When financial statement data is abnormal or in violation of regulations, existing technologies cannot quickly and accurately trace the source and flow of data. Summary of the invention
[0003] The purpose of the present invention is to provide a data tracing and analysis system and method based on big data to solve the problems raised in the prior art.
[0004] To achieve the above object, the present invention provides the following technical solutions:
[0005] In a first aspect, the present invention provides a data provenance analysis method based on big data, the method comprising the following steps:
[0006] Collect historical data, perform preprocessing and feature engineering on the historical data, use convolutional neural networks to train and evaluate based on the extracted features to obtain a big data model; extract rules based on the big data model to obtain a preliminary verification set;
[0007] Determine the special task objectives and parse them into data indicators and business rule requirements; define and configure the special task verification set based on the data indicators, obtain the data elements in the financial statements, and define the data nodes; extract the business logic, build the association edges based on the connection between the business logic and the data nodes, and obtain the special task graph;
[0008] Use depth-first search to search in the special task graph, integrate the searched data nodes and their attributes related to the special task objectives, and obtain a special task data subset;
[0009] Match and analyze the rules in the preliminary verification set with the special task data subset, formulate new rules or optimize existing rules, verify the verification set, and obtain the special task verification set;
[0010] The financial statements input by the user are verified based on the special task verification set to obtain the verification results, determine the starting point of tracing, and backtrack according to the special task diagram to obtain the data tracing results.
[0011] In combination with the first aspect, in a first implementation of the first aspect of the present application, the historical data is collected, preprocessed and feature engineered, and a convolutional neural network is used to train and evaluate based on the extracted features to obtain a big data model, including:
[0012] The historical data is specifically the historical data generated during the verification process of financial statements; determine the collection scope, conduct multi-channel collection, and collect historical data; the preprocessing includes data cleaning, data standardization and outlier processing;
[0013] The financial statements contain asset and liability data, profit and loss data, risk indicator data, liquidity risk data and capital adequacy ratio data, extract features; determine the correlation between features and financial risk assessment results by calculating the Pearson correlation coefficient between each feature; for categorical features, use the chi-square test to evaluate the independence of business type and financial risk assessment results; perform feature conversion; build a convolutional neural network model based on the extracted features; design a network structure, including a convolutional layer, a pooling layer and a fully connected layer;
[0014] The feature-engineered data is divided into training set, validation set and test set. The training set is used to train the convolutional neural network model. During the training process, the validation set is used to verify the model. The trained convolutional neural network model is evaluated using the test set, and the various performance indicators of the model are calculated. Based on the evaluation results, the model is optimized and adjusted to obtain a big data model.
[0015] In combination with the first aspect, in a second implementation of the first aspect of the present application, extracting rules based on a big data model to obtain a preliminary verification set includes:
[0016] A decision tree is used to extract rules, and the method is as follows: a tree structure is used to represent the relationship between different feature combinations and output results in the big data model; starting from the root node of the decision tree, along the branches downward, each internal node represents a feature, the branches represent the value of the feature, and the leaf nodes represent the output results; by traversing the decision tree, the rules are extracted;
[0017] The extracted rules are integrated and formatted to obtain a preliminary verification set.
[0018] In combination with the first aspect, in a third implementation of the first aspect of the present application, determining the special task target and parsing the special task target into data indicators and business rule requirements includes:
[0019] Determine the special task objectives, analyze the data types associated with the special task objectives, select the corresponding data indicators based on the data types, and define and set quantitative standards for the data indicators;
[0020] Collect financial rules and internal business rules, classify and organize them, organize experts to interpret the rules, and convert the interpreted rules into business rules, which represent specific and executable rules; for each business rule, clearly define the violation and the corresponding handling measures; use the rule conflict detection algorithm to check the business rules and analyze the causes of the conflict; adjust the rules according to the conflict detection results and actual business conditions; invite employees who are directly involved in the actual operation process of financial business, have direct contact with customers or business objects and perform specific business tasks to participate in the integrity assessment of the rules to obtain the business rule requirements.
[0021] In combination with the first aspect, in a fourth implementation of the first aspect of the present application, the definition and configuration of a special task verification set based on data indicators, obtaining data elements in a financial statement, and defining data nodes include:
[0022] Based on data indicators, build the framework of the special task verification set, determine the name, scope of application and verification objectives of the verification set; for each data indicator, set verification rules in combination with business rule requirements; configure the defined special task verification set, extract relevant data elements from financial statements, and use the data elements as data nodes; define attributes for the data nodes, and establish connection relationships between data nodes.
[0023] In combination with the first aspect, in a fifth implementation of the first aspect of the present application, extracting the business logic, building an associated edge according to the connection between the business logic and the data node, and obtaining a special task graph includes:
[0024] The extracting business logic specifically includes extracting the business logic from the business rule requirements; matching the business logic with the data nodes to find out the data nodes involved in each business logic; and determining the association direction between the data nodes according to the data flow direction in the business logic;
[0025] According to the matching results of business logic and data nodes and the association direction, Graphviz is used to draw association edges between data nodes, and attributes are defined for each association edge, including association strength, association conditions and business meaning; wherein the association strength sets a value according to the importance of the business logic, the association condition specifies under what circumstances the association is established, and the business meaning describes the business relationship represented by the association edge; the data nodes of the association edge are integrated into a special task graph.
[0026] In combination with the first aspect, in a sixth implementation of the first aspect of the present application, extracting the business logic from the business rule requirements includes:
[0027] Preprocess the text of the business rule requirements, and build a rule template library based on business domain knowledge and business rule types; use the KMP algorithm to match the preprocessed business rule requirements with the templates in the rule template library; find the most suitable template through matching and determine the structural type of the business rule; extract business elements from the business rule requirements using regular expressions based on the matched template;
[0028] The dependency syntax analysis algorithm is used to analyze the semantic dependency relationship between words in the business rule requirements, and a semantic dependency tree is constructed. The grammatical and semantic connections between the components in the sentence are obtained through the semantic dependency tree. The extracted business elements and semantic connections are integrated into the knowledge graph, and the knowledge graph is converted into predicate logic to obtain business logic.
[0029] In combination with the first aspect, in a seventh implementation of the first aspect of the present application, matching and analyzing the rules in the preliminary verification set with the special task data subset, formulating new rules or optimizing existing rules, and verifying the verification set to obtain the special task verification set includes:
[0030] Use the Levenshtein distance algorithm to convert the rules in the preliminary verification set into text strings to obtain the rule text; for the special task data subset, convert the feature description of the data node into text to obtain the data feature text; calculate the edit distance between the condition part in the rule text and the data feature text one by one; for the data features whose matching degree with the rule condition part exceeds the set threshold, determine whether the corresponding business data meets the rule conclusion and obtain the matching analysis result;
[0031] Based on the matching analysis results, when there are business scenarios or data patterns in the special task data subset that cannot be covered by the existing rules, new rules are formulated; when existing rules that are not fully compatible with the special task data subset are retrieved, the existing rules are optimized; the consistency between the newly formulated and optimized rules is checked; part of the data from the special task data subset is divided as test data, and new actual business cases are collected as supplementary test data; the new rules and optimized rules are applied to the test data, and verification operations are performed, and the judgment results of each rule during the verification process are recorded; the judgment results are quantitatively evaluated using accuracy, recall rate and F1 value to obtain the special task verification set.
[0032] In combination with the first aspect, in an eighth implementation of the first aspect of the present application, the financial statements input by the user are verified based on the special task verification set to obtain a verification result, determine the starting point of tracing, and backtrack according to the special task graph to obtain a data tracing result, including:
[0033] Each rule in the special task verification set is split into conditions and conclusions, and the transaction records of customer transfers in the financial statements are sorted into a transaction record set; the transaction record set is compared with the number and amount standards set in the rules to obtain the verification result; based on the verification result, the data that fails the verification is regarded as abnormal data, and the node corresponding to the key indicator in the abnormal data is used as the starting point of tracing;
[0034] Trace back according to the business logic relationship in the special task diagram, obtain the connection between each data node, and obtain the data tracing result.
[0035] In a second aspect, the present invention provides a data tracing and analysis system based on big data, comprising:
[0036] Data processing and model training module: including: data collection unit, data processing unit and model training and evaluation unit; wherein the data collection unit collects historical data, the data processing unit performs preprocessing and feature engineering on the historical data, and the model training and evaluation unit uses a convolutional neural network to perform training and evaluation based on the extracted features to obtain a big data model; rules are extracted based on the big data model to obtain a preliminary verification set;
[0037] Special task graph construction module: including: task target parsing unit, check set definition and configuration unit and graph construction unit; wherein the task target parsing unit determines the special task target and parses the special task target into data indicators and business rule requirements; the check set definition and configuration unit defines and configures the special task check set based on the data indicators, obtains the data elements in the financial report, and defines the data nodes; the graph construction unit extracts the business logic, constructs the associated edges according to the connection between the business logic and the data nodes, and obtains the special task graph;
[0038] Data subset acquisition module: including: a depth-first search unit and a special task data subset integration unit; wherein the depth-first search unit uses a depth-first search to search in the special task graph, and the special task data subset integration unit integrates the searched data nodes and their attributes related to the special task target to obtain the special task data subset;
[0039] Verification set optimization module: including: rule matching and optimization unit and verification set verification unit; wherein the rule matching and optimization unit matches and analyzes the rules in the preliminary verification set with the special task data subset, formulates new rules or optimizes existing rules, and the verification set verification unit verifies the verification set to obtain the special task verification set;
[0040] Report verification and data traceability module: includes: report verification unit, traceability starting point determination unit and data traceability unit; wherein, the report verification unit verifies the financial report input by the user based on the special task verification set to obtain the verification result, the traceability starting point determination unit determines the traceability starting point, and the data traceability unit backtracks according to the special task diagram to obtain the data traceability result.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] 1. The present invention extracts rules based on a big data model, can automatically learn the characteristics and rules of financial data, and generate more accurate and comprehensive rules; and in the rule matching and optimization unit, can flexibly adjust and optimize the rules according to different special task data subsets.
[0043] 2. The present invention uses convolutional neural networks for training and evaluation, which can deeply explore the complex relationships and potential laws between financial data and build a more accurate and reliable big data model. Through the construction and analysis of special task graphs, the internal connections of financial business data can be intuitively displayed, providing more powerful support for financial risk assessment and business decision-making.
[0044] 3. In the report verification and data tracing module, the present invention can quickly and accurately trace the source and flow process of data and clearly determine the root cause of abnormal data through a clear tracing starting point determination method and a backtracking algorithm based on a special task graph. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a schematic diagram of the steps of a data tracing and analysis method based on big data of the present invention;
[0046] Figure 2 It is a system structure diagram of a data tracing and analysis system based on big data of the present invention. DETAILED DESCRIPTION
[0047] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0048] Example: Figure 1-Figure 2 As shown, the present invention provides a technical solution.
[0049] like Figure 1As shown in a schematic diagram of the steps of a data tracing and analysis method based on big data, the present invention provides a data tracing and analysis method based on big data, and the method comprises the following steps:
[0050] Step S100: Collect historical data, perform preprocessing and feature engineering on the historical data, use a convolutional neural network to perform training and evaluation based on the extracted features to obtain a big data model; extract rules based on the big data model to obtain a preliminary verification set;
[0051] Specifically, the historical data is the historical data generated during the verification process of financial statements; determining the collection scope, conducting multi-channel collection, and collecting historical data; the preprocessing includes data cleaning, data standardization, and outlier processing;
[0052] The financial statements contain asset and liability data, profit and loss data, risk indicator data, liquidity risk data and capital adequacy ratio data, extract features; determine the correlation between features and financial risk assessment results by calculating the Pearson correlation coefficient between each feature; for categorical features, use the chi-square test to evaluate the independence of business type and financial risk assessment results; perform feature conversion; build a convolutional neural network model based on the extracted features; design a network structure, including a convolutional layer, a pooling layer and a fully connected layer;
[0053] The feature-engineered data is divided into training set, validation set and test set. The training set is used to train the convolutional neural network model. During the training process, the validation set is used to verify the model. The trained convolutional neural network model is evaluated using the test set, and the various performance indicators of the model are calculated. Based on the evaluation results, the model is optimized and adjusted to obtain a big data model.
[0054] Furthermore, a decision tree is used to extract rules, and the method is as follows: a tree structure is used to represent the relationship between different feature combinations and output results in the big data model; starting from the root node of the decision tree, along the branches downward, each internal node represents a feature, the branches represent the value of the feature, and the leaf nodes represent the output results; by traversing the decision tree, the rules are extracted;
[0055] The extracted rules are integrated and formatted to obtain a preliminary verification set.
[0056] In a specific embodiment, a medium-sized commercial bank was selected as the object, and the historical data generated by its statements in the past five years during the verification process were collected, with a total of about 100,000 records. 20 features including debt-to-asset ratio, current ratio, cost-to-income ratio, etc. were extracted. By calculating the Pearson correlation coefficient, it was found that the correlation between the debt-to-asset ratio and the financial risk assessment results was as high as 0.8, indicating that the debt-to-asset ratio has an important impact on financial risk assessment; while the correlation between the proportion of some non-core business income and the financial risk assessment results was only 0.2, which was a low correlation. For categorical features, such as business type (corporate business, retail business, etc.), the chi-square test results showed that the independence of retail business and financial risk assessment results was weak, indicating that retail business had a certain impact on financial risk. The skewed operating income was logarithmically transformed to make its distribution closer to the normal distribution; the business type was uniquely encoded and converted into a numerical feature.
[0057] A convolutional neural network model with 3 convolutional layers, 2 pooling layers and 2 fully connected layers was designed. The feature-engineered data was divided into a training set (70,000 records), a validation set (15,000 records) and a test set (15,000 records) at a ratio of 70%, 15% and 15%. After 100 iterations of training, the accuracy of the model on the training set gradually increased from the initial 60% to 85%. On the validation set, the accuracy of the model stabilized at around 80%, verifying the effectiveness of the model. The test set was used for evaluation, and the accuracy, recall and F1 values were calculated to be 82%, 80% and 81%. According to the evaluation results, the convolution kernel size and learning rate were adjusted, and the accuracy, recall and F1 values of the optimized model on the test set were increased to 85%, 83% and 84%, respectively, resulting in a big data model with better performance.
[0058] The decision tree is used to extract rules, for example, the rule "if the debt-to-asset ratio is greater than 0.7 and the current ratio is less than 1.5, then the financial risk is high" is obtained. By traversing the decision tree, a total of 50 rules are extracted. The extracted rules are integrated, duplicate and redundant rules are removed, and finally 30 valid rules are retained and formatted, such as unifying the expression method and data format of the rules, to obtain a preliminary verification set.
[0059] Step S200: Determine the special task target, and parse the special task target into data indicators and business rule requirements; define and configure the special task verification set based on the data indicators, obtain the data elements in the financial report, and define the data nodes; extract the business logic, and build the associated edges according to the connection between the business logic and the data nodes to obtain the special task graph;
[0060] Specifically, determine the special task objectives, analyze the data types associated with the special task objectives, select corresponding data indicators based on the data types, and define and set quantitative standards for the data indicators;
[0061] Collect financial rules and internal business rules, classify and organize them, organize experts to interpret the rules, and convert the interpreted rules into business rules, which represent specific and executable rules; for each business rule, clearly define the violation and the corresponding handling measures; use the rule conflict detection algorithm to check the business rules and analyze the causes of the conflict; adjust the rules according to the conflict detection results and actual business conditions; invite employees who are directly involved in the actual operation process of financial business, have direct contact with customers or business objects and perform specific business tasks to participate in the integrity assessment of the rules to obtain the business rule requirements.
[0062] Furthermore, based on the data indicators, a framework of a special task verification set is built to determine the name, scope of application and verification objectives of the verification set; for each data indicator, verification rules are set in combination with business rule requirements; the defined special task verification set is configured, and relevant data elements are extracted from financial statements, and the data elements are used as data nodes; attributes are defined for the data nodes, and connection relationships between data nodes are established.
[0063] Furthermore, the extracting of the business logic specifically includes extracting the business logic from the business rule requirements; matching the business logic with the data nodes to find out the data nodes involved in each business logic; and determining the association direction between the data nodes according to the data flow direction in the business logic;
[0064] According to the matching results of business logic and data nodes and the association direction, Graphviz is used to draw association edges between data nodes, and attributes are defined for each association edge, including association strength, association conditions and business meaning; wherein the association strength sets a value according to the importance of the business logic, the association condition specifies under what circumstances the association is established, and the business meaning describes the business relationship represented by the association edge; the data nodes of the association edge are integrated into a special task graph.
[0065] In a specific embodiment, the collection and sorting function of financial rules and internal business rules in the invention is used to collect business rules on fund operations within the bank, such as fund allocation process specifications, asset allocation strategies, etc., and classify and sort them into categories such as fund raising rules and fund utilization rules. Fund operation experts and business backbones are organized to interpret the rules according to the rule interpretation ideas in the invention and convert them into specific executable business rules. For example, "optimize fund allocation" is converted into "evaluate the yield and risk of various assets every week. When the yield of a certain type of asset is lower than expected for two consecutive weeks and the risk increases, reduce the asset allocation ratio and increase the allocation of high-yield and low-risk assets." For each business rule, the definition and handling measures of violations are clearly defined. For example, if the asset allocation is not adjusted in a timely manner according to regulations, it is defined as a violation.
[0066] By using the rule conflict detection algorithm in the invention, it was found that in the fund allocation rules, there were conflicts in the fund use priority and quota allocation of different business departments. After analysis, it was found that the original rules were not updated in time due to changes in fund demand caused by business expansion. According to the actual business situation, the fund use priority and quota allocation rules were re-established to ensure efficient fund allocation.
[0067] We invited frontline fund operators and risk management specialists to participate in the integrity assessment of the rules. They pointed out that when the market fluctuates greatly, it is more difficult to manage fund liquidity. Based on this feedback, we added corresponding business rules, such as "when market fluctuations exceed a certain threshold, activate the emergency plan, give priority to ensuring fund liquidity, and suspend some high-risk investment businesses", and finally obtained the perfect business rules requirements.
[0068] Based on data indicators, the framework building function in the invention is used to build a framework called "Special Task Verification Set for Improving Fund Operation Efficiency", and the scope of application is clearly defined as all fund operation businesses of the bank, and the verification goal is to evaluate and improve fund operation efficiency. For the idle fund rate indicator, the verification rule is set as "If the idle fund rate is higher than 5%, it is necessary to analyze the causes of idle funds and formulate a fund utilization plan"; for the fund turnover rate indicator, the rule is set as "If the fund turnover rate is lower than the industry average, it is necessary to optimize the fund flow process and shorten the fund recovery cycle". Configure the defined verification set, and extract relevant data elements such as idle fund rate, fund turnover rate, asset return rate, and the proportion of various assets from the report and fund operation system as data nodes. Define attributes for each data node, such as the data type of idle fund rate is numeric, with a value range of 0-100%; the data type of the proportion of various assets is numeric, with a value range of 0-1. Establish a connection relationship between data nodes, for example, the idle fund rate and the asset return rate are associated through the fund utilization efficiency, indicating the impact of the idle fund rate on the asset return rate.
[0069] Extract business logic from business rule requirements, such as "If the idle rate of funds is too high and the asset yield is too low, it may lead to low efficiency of fund operation, and it is necessary to optimize fund allocation". Match this business logic with the data node, and use the matching algorithm in the invention to find out the data nodes involved, such as idle rate of funds, asset yield, and proportion of various assets. According to the data flow in the business logic, determine the association direction between data nodes. For example, the data flow direction of idle rate of funds and asset yield points to the proportion of various assets, because the idle rate of funds and the asset yield will affect the asset allocation decision. According to the matching results and association direction of business logic and data nodes, use Graphviz to draw association edges between data nodes. Define attributes for each association edge. For the "idle rate of funds-proportion of various assets" association edge, the association strength is set to 8 (out of 10 points, indicating that the business logic has a greater impact on asset allocation), the association condition is "idle rate of funds exceeds the threshold and the asset yield is lower than the standard", and the business meaning is "associate with the proportion of various assets through abnormal idle rate of funds, which is used to adjust asset allocation to improve fund operation efficiency". Integrate all related edges and data nodes into a special task graph to intuitively display the data relationships and business logic in the fund operation business.
[0070] Specifically, the business rule requirements are preprocessed, and a rule template library is constructed according to business domain knowledge and business rule types; the preprocessed business rule requirements are matched with templates in the rule template library using a KMP algorithm; the most suitable template is found through matching to determine the structural type of the business rule; and business elements are extracted from the business rule requirements using regular expressions according to the matched template;
[0071] The dependency syntax analysis algorithm is used to analyze the semantic dependency relationship between words in the business rule requirements, and a semantic dependency tree is constructed. The grammatical and semantic connections between the components in the sentence are obtained through the semantic dependency tree. The extracted business elements and semantic connections are integrated into the knowledge graph, and the knowledge graph is converted into predicate logic to obtain business logic.
[0072] In a specific embodiment, based on business domain knowledge and business rule types, the following rule templates are constructed: Condition-Action Template: "When [condition], execute [action]". Cycle-Task Template: "[cycle] perform [task] on [object]".
[0073] Match the preprocessed business rules with the templates in the rule template library. For the rule "Evaluate the yield and risk of various types of assets every week. When the yield of a certain type of asset is lower than expected for two consecutive weeks and the risk increases, reduce the asset allocation ratio and increase the allocation of high-yield, low-risk assets", first match it with the "cycle-task template" and find that "weekly" corresponds to "cycle", "the yield and risk of various types of assets" corresponds to "object", and "evaluation" corresponds to "task", and the initial match is successful; then match it with the "condition-action template", "the yield of a certain type of asset is lower than expected for two consecutive weeks and the risk increases" corresponds to "condition", "reduce the asset allocation ratio and increase the allocation of high-yield, low-risk assets" corresponds to "action", and finally determine that the "condition-action template" is the most suitable template, and the structural type of the business rule is conditional trigger action type.
[0074] According to the matched "condition-action template", regular expressions are used to extract business elements: Condition element: "The yield of a certain type of asset has been lower than expected for two consecutive weeks and the risk has increased." Action element: "Reduce the allocation ratio of this asset and increase the allocation of high-yield, low-risk assets."
[0075] Taking "a certain type of asset has a yield lower than expected for two consecutive weeks and the risk has increased" as an example, a semantic dependency tree is constructed. In the semantic dependency tree, "yield" is the core word, "a certain type of asset" is the owner of "yield", "two consecutive weeks" modifies "yield", indicating the time period, "lower than" is the relationship word between "yield" and "expectation", and "risks increase" and "yield lower than expected" are connected by "and", indicating a parallel condition. The semantic dependency tree clearly shows the grammatical and semantic connections between the various components in the sentence.
[0076] In the knowledge graph, nodes include "asset", "return", "risk", "expectation", etc., and edges represent the relationship between them, such as "belonging", "modification", "comparative relationship", "parallel relationship", etc. Convert the knowledge graph into predicate logic.
[0077] Step S300: Use depth-first search to search in the special task graph, integrate the searched data nodes and their attributes related to the special task target, and obtain a special task data subset;
[0078] In a specific embodiment, the "Idle Fund Rate" data node is used as the search starting point, because the idle fund rate is a key indicator for measuring the efficiency of bank fund operation and is directly related to the special task target. The attributes of the "Idle Fund Rate" node include data type (numeric type), value range (0-100%) and current value (8%). Starting from the "Idle Fund Rate" node, find the "Asset Allocation" node connected to it according to the associated edge. The "Asset Allocation" node contains attributes such as the allocation ratio of different assets, such as the stock asset allocation ratio is 30%, the bond asset allocation ratio is 50%, etc. Then access the adjacent "Rate of Return" node from the "Asset Allocation" node, which records the rate of return attributes of various assets, such as the annualized rate of return of stock assets is 8%, and the annualized rate of return of bond assets is 5%. Then from the "Rate of Return" node, follow the associated edge to reach the "Risk Assessment" node. The "Risk Assessment" node attributes include risk level (such as low, medium, and high). The current stock asset risk level is medium, and the bond asset risk level is low. Continue to follow the depth-first search strategy, accessing other adjacent nodes from the "Risk Assessment" node until all reachable nodes related to the special task objectives are traversed.
[0079] During the search process, all the data nodes and their attributes related to the special task objectives that are accessed are recorded. The integrated data subset includes the "fund idle rate" node and its attributes, the "asset allocation" node and its attributes, the "yield" node and its attributes, and the "risk assessment" node and its attributes.
[0080] Step S400: matching and analyzing the rules in the preliminary verification set with the special task data subset, formulating new rules or optimizing existing rules, verifying the verification set, and obtaining the special task verification set;
[0081] Specifically, the Levenshtein distance algorithm is used to convert the rules in the preliminary verification set into text strings to obtain the rule text; for the special task data subset, the feature description of the data node is converted into text to obtain the data feature text; the edit distance between the condition part in the rule text and the data feature text is calculated one by one; for the data features whose matching degree with the rule condition part exceeds the set threshold, it is determined whether the corresponding business data meets the rule conclusion to obtain the matching analysis result;
[0082] Based on the matching analysis results, when there are business scenarios or data patterns in the special task data subset that cannot be covered by the existing rules, new rules are formulated; when existing rules that are not fully compatible with the special task data subset are retrieved, the existing rules are optimized; the consistency between the newly formulated and optimized rules is checked; part of the data from the special task data subset is divided as test data, and new actual business cases are collected as supplementary test data; the new rules and optimized rules are applied to the test data, and verification operations are performed, and the judgment results of each rule during the verification process are recorded; the judgment results are quantitatively evaluated using accuracy, recall rate and F1 value to obtain the special task verification set.
[0083] In a specific embodiment, the rules in the preliminary verification set are converted into text strings. For the special task data subset, the feature description of the data node is converted into text. Using the Levenshtein distance algorithm, the conditional part in the rule text and the data feature text are calculated one by one for the edit distance. For example, for the rule "Idle fund rate is higher than 5% → adjust asset allocation", the conditional part "Idle fund rate is higher than 5%" and the data feature text "Idle fund rate is 8%" are calculated to have an edit distance, and the threshold is set to 2. The calculated edit distance is 1 (changing "higher than 5%" to "8%" only requires one replacement operation), and the matching degree exceeds the threshold.
[0084] For data features whose matching degree exceeds the threshold, determine whether the corresponding business data meets the rule conclusion. For the rules of "idle capital rate is 8%" and "idle capital rate is higher than 5% → adjust asset allocation", because 8% is higher than 5%, it meets the rule conditions, so determine whether the asset allocation adjustment has been made. If not, the rule will not pass this matching. After matching analysis of all rules and data features, the matching analysis results are obtained.
[0085] Based on the matching analysis results, we found that there are some new business scenarios in the special task data subset, such as "when the risk level of bond assets changes from low to medium and its yield drops by 2 percentage points within a month, the bond assets need to be re-evaluated", which cannot be covered by the existing rules. Therefore, a new rule is formulated: "The risk level of bond assets increases and the yield drops → re-evaluate bond assets".
[0086] The rule "If the stock asset yield is less than 6%, consider reducing the stock asset allocation ratio" was retrieved. In the data subset, it was found that when the stock asset yield is between 5% and 6%, adjusting the asset allocation according to the existing rules may lead to excessive operations. The rule is optimized to "If the stock asset yield is less than 5%, reduce the stock asset allocation ratio; if the stock asset yield is between 5% and 6%, pay close attention to market trends and do not adjust for the time being." Check the consistency between the newly formulated and optimized rules. For example, there is no conflict between the new and optimized rules in the trigger conditions and operations of asset allocation adjustments to ensure the consistency of the rule system.
[0087] 30% of the data is divided from the special task data subset as test data, and 50 new business cases in the recent actual fund operation of the bank are collected as supplementary test data. The new rules and optimized rules are applied to the test data to perform verification operations. For example, for the rule "the risk level of bond assets has increased and the yield has decreased → re-evaluate bond assets", find the bond asset data that meets the conditions in the test data, determine whether a re-evaluation has been performed, and record the judgment results. The accuracy, recall rate and F1 value are used to quantitatively evaluate the judgment results. After calculation, the accuracy rate is 85%, the recall rate is 80%, and the F1 value is 82.5%. According to the evaluation results, it is believed that the accuracy and coverage of the rules meet certain requirements, and the special task verification set is obtained, which can be used for subsequent monitoring and analysis of the efficiency of bank fund operations.
[0088] Step S500: Verify the financial statements input by the user based on the special task verification set, obtain the verification results, determine the starting point of tracing, and backtrack according to the special task diagram to obtain the data tracing results.
[0089] Specifically, each rule in the special task verification set is split into conditions and conclusions, and the transaction records of customer transfers in the financial statements are sorted into a transaction record set; the transaction record set is compared with the number and amount standards set in the rules to obtain a verification result; based on the verification result, the data that fails the verification is regarded as abnormal data, and the node corresponding to the key indicator in the abnormal data is used as the starting point for tracing;
[0090] Trace back according to the business logic relationship in the special task diagram, obtain the connection between each data node, and obtain the data tracing result.
[0091] In a specific embodiment, backtracking is performed according to the business logic relationship in the special task diagram. The "transaction frequency" node and the "customer capital demand" node in the special task diagram are associated through the business logic of "reflecting the customer's capital activity", and the "transaction amount" node and the "customer asset scale" node are associated through the business logic of "reflecting the customer's capital mobilization ability". Backtracking starts from the "transaction frequency" and "transaction amount" tracing starting points to obtain the connection between each data node. After backtracking, it was found that the customer's high transaction frequency and large amount were due to the recent participation in a large investment project, and its capital demand suddenly increased. At the same time, the customer's asset scale is large and has the corresponding capital mobilization ability. Through such backtracking analysis, the data tracing results are obtained, and the causes and related influencing factors of abnormal transactions are clarified.
[0092] like Figure 2 As shown in the system structure diagram of a data source tracing and analysis system based on big data, the present invention provides a data source tracing and analysis system based on big data, including:
[0093] Data processing and model training module: including: data collection unit, data processing unit and model training and evaluation unit; wherein the data collection unit collects historical data, the data processing unit performs preprocessing and feature engineering on the historical data, and the model training and evaluation unit uses a convolutional neural network to perform training and evaluation based on the extracted features to obtain a big data model; rules are extracted based on the big data model to obtain a preliminary verification set;
[0094] Special task graph construction module: including: task target parsing unit, check set definition and configuration unit and graph construction unit; wherein the task target parsing unit determines the special task target and parses the special task target into data indicators and business rule requirements; the check set definition and configuration unit defines and configures the special task check set based on the data indicators, obtains the data elements in the financial report, and defines the data nodes; the graph construction unit extracts the business logic, constructs the associated edges according to the connection between the business logic and the data nodes, and obtains the special task graph;
[0095] Data subset acquisition module: including: a depth-first search unit and a special task data subset integration unit; wherein the depth-first search unit uses a depth-first search to search in the special task graph, and the special task data subset integration unit integrates the searched data nodes and their attributes related to the special task target to obtain the special task data subset;
[0096] Verification set optimization module: including: rule matching and optimization unit and verification set verification unit; wherein the rule matching and optimization unit matches and analyzes the rules in the preliminary verification set with the special task data subset, formulates new rules or optimizes existing rules, and the verification set verification unit verifies the verification set to obtain the special task verification set;
[0097] Report verification and data traceability module: includes: report verification unit, traceability starting point determination unit and data traceability unit; wherein, the report verification unit verifies the financial report input by the user based on the special task verification set to obtain the verification result, the traceability starting point determination unit determines the traceability starting point, and the data traceability unit backtracks according to the special task diagram to obtain the data traceability result.
[0098] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.
Claims
1. A data tracing analysis method based on big data, characterized in that: The method comprises the following steps: Collect historical data, perform preprocessing and feature engineering on the historical data, use convolutional neural networks to train and evaluate based on the extracted features to obtain a big data model; extract rules based on the big data model to obtain a preliminary verification set; The historical data is specifically the historical data generated during the verification process of financial statements; determine the collection scope, conduct multi-channel collection, and collect historical data; the preprocessing includes data cleaning, data standardization and outlier processing; The financial statements contain asset and liability data, profit and loss data, risk indicator data, liquidity risk data and capital adequacy ratio data, extract features; determine the correlation between features and financial risk assessment results by calculating the Pearson correlation coefficient between each feature; for categorical features, use the chi-square test to evaluate the independence of business type and financial risk assessment results; perform feature conversion; build a convolutional neural network model based on the extracted features; design a network structure, including a convolutional layer, a pooling layer and a fully connected layer; The feature-engineered data is divided into training set, validation set and test set. The training set is used to train the convolutional neural network model. During the training process, the validation set is used to validate the model. The trained convolutional neural network model is evaluated using the test set, and the performance indicators of the model are calculated. Based on the evaluation results, the model is optimized and adjusted to obtain a big data model. A decision tree is used to extract rules, and the method is as follows: a tree structure is used to represent the relationship between different feature combinations and output results in the big data model; starting from the root node of the decision tree, along the branches downward, each internal node represents a feature, the branches represent the value of the feature, and the leaf nodes represent the output results; by traversing the decision tree, the rules are extracted; Integrate and format the extracted rules to obtain a preliminary verification set; Determine the special task objectives and parse them into data indicators and business rule requirements; define and configure the special task verification set based on the data indicators, obtain the data elements in the financial statements, and define the data nodes; extract the business logic, build the association edges based on the connection between the business logic and the data nodes, and obtain the special task graph; The extracting business logic specifically involves extracting business logic from business rule requirements; Preprocess the text of the business rule requirements, and build a rule template library based on business domain knowledge and business rule types; use the KMP algorithm to match the preprocessed business rule requirements with the templates in the rule template library; find the most suitable template through matching and determine the structural type of the business rule; extract business elements from the business rule requirements using regular expressions based on the matched template; Analyze the semantic dependency relationship between words in the business rule requirements using the dependency syntax analysis algorithm, construct a semantic dependency tree, and obtain the grammatical and semantic connections between the components in the sentence through the semantic dependency tree; integrate the extracted business elements and semantic connections into the knowledge graph, convert the knowledge graph into predicate logic, and obtain the business logic; Use depth-first search to search in the special task graph, integrate the searched data nodes and their attributes related to the special task objectives, and obtain a special task data subset; Match and analyze the rules in the preliminary verification set with the special task data subset, formulate new rules or optimize existing rules, verify the verification set, and obtain the special task verification set; The Levenshtein distance algorithm is used to convert the rules in the preliminary verification set into text strings to obtain rule texts; for the special task data subset, the feature descriptions of the data nodes are converted into text to obtain data feature texts; The financial statements input by the user are verified based on the special task verification set to obtain the verification results, determine the starting point of tracing, and backtrack according to the special task diagram to obtain the data tracing results.
2. According to the data tracing and analysis method based on big data in claim 1, it is characterized in that: Determining the special task objectives and parsing the special task objectives into data indicators and business rule requirements includes: Determine the special task objectives, analyze the data types associated with the special task objectives, select the corresponding data indicators based on the data types, and define and set quantitative standards for the data indicators; Collect financial rules and internal business rules, classify and organize them, organize experts to interpret the rules, and convert the interpreted rules into business rules, which represent specific and executable rules; for each business rule, clearly define the violation and the corresponding handling measures; use the rule conflict detection algorithm to check the business rules and analyze the causes of the conflict; adjust the rules according to the conflict detection results and actual business conditions; invite employees who are directly involved in the actual operation process of financial business, have direct contact with customers or business objects and perform specific business tasks to participate in the integrity assessment of the rules to obtain the business rule requirements.
3. The data tracing and analysis method based on big data according to claim 1 is characterized in that: The definition and configuration of the special task verification set based on the data indicators, obtaining the data elements in the financial statements, and defining the data nodes include: Based on data indicators, build the framework of the special task verification set, determine the name, scope of application and verification objectives of the verification set; for each data indicator, set verification rules in combination with business rule requirements; configure the defined special task verification set, extract relevant data elements from financial statements, and use the data elements as data nodes; define attributes for the data nodes, and establish connection relationships between data nodes.
4. The data tracing and analysis method based on big data according to claim 1 is characterized in that: The extracting of business logic, building associated edges according to the connection between business logic and data nodes, and obtaining a special task graph include: Match business logic with data nodes to find the data nodes involved in each business logic; determine the association direction between data nodes based on the data flow in the business logic; According to the matching results of business logic and data nodes and the association direction, Graphviz is used to draw association edges between data nodes, and attributes are defined for each association edge, including association strength, association conditions and business meaning; wherein the association strength sets a value according to the importance of the business logic, the association condition specifies under what circumstances the association is established, and the business meaning describes the business relationship represented by the association edge; the data nodes of the association edge are integrated into a special task graph.
5. The data tracing and analysis method based on big data according to claim 1 is characterized in that: The matching analysis of the rules in the preliminary verification set with the special task data subset, formulating new rules or optimizing existing rules, and verifying the verification set to obtain the special task verification set includes: The condition part in the rule text is used to calculate the edit distance with the data feature text one by one; for data features whose matching degree with the rule condition part exceeds the set threshold, it is determined whether the corresponding business data meets the rule conclusion to obtain the matching analysis result; Based on the matching analysis results, when there are business scenarios or data patterns in the special task data subset that cannot be covered by the existing rules, new rules are formulated; when existing rules that are not fully compatible with the special task data subset are retrieved, the existing rules are optimized; the consistency between the newly formulated and optimized rules is checked; part of the data from the special task data subset is divided as test data, and new actual business cases are collected as supplementary test data; the new rules and optimized rules are applied to the test data, and verification operations are performed, and the judgment results of each rule during the verification process are recorded; the judgment results are quantitatively evaluated using accuracy, recall rate and F1 value to obtain the special task verification set.
6. The data tracing and analysis method based on big data according to claim 1 is characterized in that: The financial statements input by the user are verified based on the special task verification set to obtain the verification results, determine the starting point of the traceability, and backtrack according to the special task diagram to obtain the data traceability results, including: Each rule in the special task verification set is split into conditions and conclusions, and the transaction records of customer transfers in the financial statements are sorted into a transaction record set; the transaction record set is compared with the number and amount standards set in the rules to obtain the verification result; based on the verification result, the data that fails the verification is regarded as abnormal data, and the node corresponding to the key indicator in the abnormal data is used as the starting point of tracing; Trace back according to the business logic relationship in the special task diagram, obtain the connection between each data node, and obtain the data tracing result.
7. A data tracing and analysis system based on big data, using a data tracing and analysis method based on big data according to any one of claims 1 to 6, characterized in that: include: Data processing and model training module: including: data collection unit, data processing unit and model training and evaluation unit; wherein the data collection unit collects historical data, the data processing unit performs preprocessing and feature engineering on the historical data, and the model training and evaluation unit uses a convolutional neural network to perform training and evaluation based on the extracted features to obtain a big data model; rules are extracted based on the big data model to obtain a preliminary verification set; Special task graph construction module: including: task target parsing unit, check set definition and configuration unit and graph construction unit; wherein the task target parsing unit determines the special task target and parses the special task target into data indicators and business rule requirements; the check set definition and configuration unit defines and configures the special task check set based on the data indicators, obtains the data elements in the financial report, and defines the data nodes; the graph construction unit extracts the business logic, constructs the associated edges according to the connection between the business logic and the data nodes, and obtains the special task graph; Data subset acquisition module: including: a depth-first search unit and a special task data subset integration unit; wherein the depth-first search unit uses a depth-first search to search in the special task graph, and the special task data subset integration unit integrates the searched data nodes and their attributes related to the special task target to obtain the special task data subset; Verification set optimization module: including: rule matching and optimization unit and verification set verification unit; wherein the rule matching and optimization unit matches and analyzes the rules in the preliminary verification set with the special task data subset, formulates new rules or optimizes existing rules, and the verification set verification unit verifies the verification set to obtain the special task verification set; Report verification and data traceability module: includes: report verification unit, traceability starting point determination unit and data traceability unit; wherein, the report verification unit verifies the financial report input by the user based on the special task verification set to obtain the verification result, the traceability starting point determination unit determines the traceability starting point, and the data traceability unit backtracks according to the special task diagram to obtain the data traceability result.
Citation Information
Patent Citations
Cascading failure evolution path traceability and prediction method and device based on knowledge graph
CN112990551A
Large language model knowledge graph construction method in financial field
CN119378660A