Methods, systems, and computer-readable media for the management of software data assets
By employing multi-dimensional risk clustering analysis, regulatory compliance evaluation, and abnormal transaction behavior analysis, this technology addresses the challenges of static data identification, unscientific value assessment, inability to dynamically adjust security strategies, and cross-institutional sharing in financial software data asset management. It enables intelligent management of data assets throughout their entire lifecycle and collaborative value mining across institutions, thereby improving management efficiency and decision-making accuracy.
Patent Information
- Application Number
- CN202510298365.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-03-13
AI Technical Summary
Existing financial software data asset management technologies suffer from several problems, including static data identification and classification, lack of scientific quantification in data value assessment, inability to dynamically adjust security protection strategies, fragmented data lifecycle management, and conflicts between privacy protection and data sharing in cross-institutional data collaborative analysis. These issues limit the efficiency of data value mining and utilization.
A data asset catalog is generated through multi-dimensional risk clustering analysis, and the data risk value coefficient is calculated by combining a regulatory compliance evaluation model. A hierarchical security protection mechanism is established by applying anomaly analysis algorithms for transaction behavior, monitoring data timeliness and adjusting storage strategies, generating risk control value insight data by using anti-fraud scenario identification algorithms, and achieving cross-institutional data sharing through federated learning technology. A compliant exchange platform is established by adopting financial-grade differential privacy protection.
It enables intelligent management of data assets throughout their entire lifecycle, improves the integrity and accuracy of data identification, dynamically adjusts security policies to enhance the accuracy of abnormal behavior detection, optimizes storage efficiency, ensures data quality, and enables cross-institutional data sharing while protecting privacy, thereby improving management efficiency and decision-making accuracy.
Smart Images

Figure CN120181985B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method, system, and computer-readable medium for managing software data assets. Background Technology
[0002] With the rapid development of fintech, the amount of data generated by financial software systems has exploded, and this data has become a core asset of financial institutions. Existing financial data asset management technologies mainly focus on three aspects: data storage, data retrieval, and data security. In terms of data storage, traditional technologies combine relational databases with distributed file systems to store structured and unstructured data. Regarding data retrieval, existing technologies primarily rely on SQL query languages and NoSQL interfaces to provide data access services. In terms of data security, encrypted storage, access control, and firewall isolation are the main protection measures. In addition, some financial institutions also use data governance frameworks to classify and manage the quality of data assets; however, this management is often limited to within a single institution and lacks cross-institutional collaborative sharing mechanisms.
[0003] However, existing financial software data asset management technologies have significant shortcomings. First, the identification and classification of data assets are too static, failing to adapt to the dynamic nature of data value changes over time. Second, data value assessment lacks scientific quantitative methods, leading to a mismatch between resource allocation and the actual value of the data. Third, data security protection strategies are usually based on preset, fixed rules, unable to dynamically adjust protection levels according to data usage behavior. Fourth, data lifecycle management is fragmented, making it difficult to achieve intelligent management of the entire process from creation to destruction. Fifth, data value mining is often limited to within a single institution, and cross-institutional collaborative data analysis faces the conflict between data privacy protection and data sharing. These problems severely limit the efficiency of data asset value mining and utilization in financial software. Summary of the Invention
[0004] This application provides a method, system, and computer-readable medium for managing software data assets, which enables intelligent management of the entire lifecycle of financial software data assets and collaborative value mining across institutions while ensuring data security and privacy.
[0005] Firstly, this application provides a method for managing software data assets. The method includes: scanning a financial transaction software system, extracting user behavior data and transaction flow information, performing multi-dimensional risk clustering analysis on the extracted financial data, and generating a financial software data asset catalog; loading the financial software data asset catalog, calculating the data risk value coefficient through a regulatory compliance evaluation model, and generating a financial data asset risk heatmap; inputting the financial data asset risk heatmap into a hierarchical processing module, executing a transaction behavior anomaly analysis algorithm, and establishing a hierarchical security protection mechanism for transaction data; connecting to the hierarchical security protection mechanism for transaction data, monitoring the timeliness parameters of financial data, calculating transaction data health indicators, adjusting cold and hot data storage strategies, and forming a financial data asset lifecycle management system; importing the transaction data from the financial data asset lifecycle management system into a value mining engine, applying an anti-fraud scenario identification algorithm, and generating financial risk control value insight data; based on the financial risk control value insight data, applying federated learning technology to process cross-institutional data sharing requests, executing financial differential privacy protection, and establishing a compliant financial data asset exchange platform.
[0006] Secondly, this application provides a management system for software data assets, the management system for software data assets comprising:
[0007] The extraction module is used to scan financial transaction software systems, extract user behavior data and transaction flow information, perform multi-dimensional risk clustering analysis on the extracted financial data, and generate a financial software data asset catalog.
[0008] The calculation module is used to load the financial software data asset catalog, calculate the data risk value coefficient through the regulatory compliance evaluation model, and generate a financial data asset risk heat map.
[0009] The input module is used to input the financial data asset risk heat map into the hierarchical processing module, execute the transaction behavior anomaly analysis algorithm, and establish a hierarchical security protection mechanism for transaction data.
[0010] The adjustment module is used to connect to the transaction data hierarchical security protection mechanism, monitor the timeliness parameters of financial data, calculate the health index of transaction data, adjust the cold and hot data storage strategy, and form a financial data asset lifecycle management system.
[0011] The import module is used to import transaction data from the financial data asset lifecycle management system into the value mining engine, apply anti-fraud scenario identification algorithms, and generate financial risk control value insight data.
[0012] A module is established to process cross-institutional data sharing requests based on the aforementioned financial risk control value insight data, apply federated learning technology, implement financial-grade differential privacy protection, and establish a compliant financial data asset exchange platform.
[0013] A third aspect of the present invention provides a computer device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the computer device to perform the above-described method for managing software data assets.
[0014] A fourth aspect of the present invention provides a computer-readable medium storing instructions that, when executed on a computer, cause the computer to perform the above-described method for managing software data assets.
[0015] The technical solution provided in this application, by scanning the technical characteristics of financial transaction software systems, can comprehensively and automatically identify potential data assets, avoiding omissions and subjectivity in manual identification, and improving the completeness and accuracy of data asset identification. Utilizing the technical characteristics of multi-dimensional risk clustering analysis, data is scientifically classified according to business value, security level, and usage characteristics, forming a structured catalog of financial software data assets, laying the foundation for subsequent management. Through a regulatory compliance evaluation model, the value of data assets is assessed, realizing the quantitative measurement of data value, shifting data asset management decisions from subjective judgment to data-driven approaches. A visualized financial data asset risk heatmap intuitively displays the distribution of data asset value, facilitating managers to quickly identify high-value data areas. The application of anomaly analysis algorithms for transaction behavior transforms security protection from static rules to dynamic behavior analysis, automatically adjusting security strategies based on changes in user behavior patterns, significantly improving the accuracy of anomaly detection. A tiered security protection mechanism enables differentiated allocation of security resources, avoiding the polarized problems of resource waste and insufficient protection. By monitoring the timeliness parameters of financial data and calculating the health indicators of transaction data, combined with intelligent adjustments to hot and cold data storage strategies, a lifecycle management system for financial data assets is formed. This system achieves automated management of the entire lifecycle from data creation to destruction, optimizing storage efficiency and ensuring data quality. The application of anti-fraud scenario identification algorithms significantly improves the ability to identify fraudulent behavior, and the resulting financial risk control insights are directly transformed into business value. In particular, the innovative application of federated learning technology and financial differential privacy protection resolves the contradiction between data sharing and privacy protection, enabling institutions to achieve collaborative model training without exposing raw data. A compliant financial data asset exchange platform breaks down data silos, enabling cross-institutional sharing of data value. The extensive application of artificial intelligence algorithms in this solution, such as K-means clustering in data identification and classification, deep learning in behavioral anomaly detection, and federated learning in privacy-preserving data sharing, fully considers the matching of algorithm characteristics with business scenarios. Through reasonable parameter design and model selection, the applicability and effectiveness of the algorithms in financial data asset management are ensured, significantly improving management efficiency and decision-making accuracy. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of one embodiment of the method for managing software data assets in this application.
[0018] Figure 2 This is a schematic diagram of an embodiment of a management system for software data assets in this application.
[0019] Figure 3 This is a schematic block diagram of the structure of the computer device in an embodiment of the present invention. Detailed Implementation
[0020] This application provides a method, system, and computer-readable medium for managing software data assets. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0021] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the method for managing software data assets in this application includes:
[0022] Step S101: Scan the financial transaction software system, extract user behavior data and transaction flow information, perform multi-dimensional risk clustering analysis on the extracted financial data, and generate a financial software data asset catalog.
[0023] Step S102: Load the financial software data asset catalog, calculate the data risk value coefficient through the regulatory compliance evaluation model, and generate a financial data asset risk heat map;
[0024] Step S103: Input the financial data asset risk heat map into the hierarchical processing module, execute the transaction behavior anomaly analysis algorithm, and establish a hierarchical security protection mechanism for transaction data;
[0025] Step S104: Connect to the transaction data hierarchical security protection mechanism, monitor the timeliness parameters of financial data, calculate the health index of transaction data, adjust the cold and hot data storage strategy, and form a financial data asset lifecycle management system.
[0026] Step S105: Import the transaction data from the financial data asset lifecycle management system into the value mining engine, apply the anti-fraud scenario identification algorithm, and generate financial risk control value insight data.
[0027] Step S106: Based on financial risk control value insight data, apply federated learning technology to process cross-institutional data sharing requests, implement financial-grade differential privacy protection, and establish a compliant financial data asset exchange platform.
[0028] It is understood that the executing entity of this application can be a management system for software data assets, or it can be a terminal or a server; no specific limitation is made here. This application's embodiment uses a server as an example for illustration.
[0029] Specifically, adaptive web crawling technology is used to comprehensively scan financial trading software systems. Adaptive web crawling is a technique that automatically adjusts its crawling strategy based on the system structure, connecting to various database interfaces within the financial trading software system to collect user login frequency data, transaction operation trajectory data, and fund flow records. The collected user behavior data includes detailed behavioral indicators such as user login time, operation sequence, and dwell time; transaction flow information includes key elements such as transaction time, amount, counterparty, and transaction type. This raw data is then input into a K-means clustering algorithm for multi-dimensional risk clustering analysis. This algorithm groups similar data into the same cluster by calculating the Euclidean distance between data points. Multi-dimensional risk clustering analysis of financial data involves comprehensive processing of user behavior stratification, transaction pattern identification, and fund flow analysis, ultimately forming a financial software data asset catalog that includes data classification, correlation relationships, and risk levels.
[0030] After loading the financial software data asset catalog, a multi-level fuzzy comprehensive evaluation model is used to calculate the data risk value coefficient. This model first extracts data sensitivity level identifiers, data update cycle values, and data access frequency indicators from the catalog to construct a three-dimensional evaluation parameter matrix. Then, the analytic hierarchy process (AHP) is applied to assign weights to the matrix elements, resulting in a weight system reflecting the importance of each factor. The model establishes scoring criteria based on national financial regulatory provisions and industry compliance standards, assigning a compliance score to each data asset. Fuzzy set theory is used to perform fuzzy comprehensive calculations of the weights and scores to obtain the data risk value coefficient. This coefficient is then analyzed over time using a Markov decision process to predict risk change trends and form a risk evolution trajectory. Finally, the risk coefficient and evolution trajectory are mapped onto a color gradient spectrum to generate a heatmap of financial data asset risk that visually displays the risk distribution.
[0031] The financial data asset risk heatmap is input into the hierarchical processing module to extract risk level identifiers, generate a risk data mapping table, and divide it into high, medium, and low risk zones. Simultaneously, time-series analysis is performed on historical transaction records to construct a benchmark model of transaction behavior, calculating transaction frequency curves and transaction amount distributions. Deep learning methods are used to compare the benchmark model with current transaction behavior, identifying abnormal behavior features that deviate from normal patterns, forming an abnormal transaction behavior feature library. Based on this feature library and high-risk zone data, the Isolation Forest algorithm is applied to detect abnormal transaction behavior points and generate anomaly risk scores. The Isolation Forest algorithm calculates the degree of isolation of sample points by constructing random decision trees; anomalies are often more easily isolated. Security levels are determined based on the anomaly risk scores, and targeted access control, encryption, and auditing strategies are formulated. These security strategy sets are combined with blockchain verification mechanisms to construct a hierarchical security protection mechanism for transaction data. Connecting to the hierarchical security protection mechanism for transaction data, a data lifecycle state transition table is established, dividing data states into five stages: creation period, active period, low-frequency period, archiving period, and destruction period. The access frequency of statistical data is used to calculate the access heat of transaction data, forming an access heat curve to identify peak and trough periods. The system performs multi-dimensional checks on the integrity, consistency, and timeliness of transaction data, calculates data quality scores, and generates transaction data health indicators based on business rule compliance. Health thresholds are set to trigger data repair processes, restoring and correcting damaged data. Based on access frequency curves and lifecycle state transition tables, the system dynamically adjusts hot and cold data storage strategies, allocating high-frequency access data to high-performance storage layers and migrating low-frequency access data to lower-cost storage layers. These elements are integrated and correlated to construct a financial data asset lifecycle management system.
[0032] Active transaction data is selected from the financial data asset lifecycle management system to construct a transaction behavior chain diagram, marking transaction entity nodes and fund flow paths. Topological structure analysis is performed on this chain diagram to identify transaction loops and fund convergence points, generating a set of suspicious transaction patterns. Frequent itemset mining is conducted through association rule analysis to calculate the support and confidence of different transaction patterns, forming a fraud risk feature library. This feature library is matched with financial risk control scenario templates to extract key identification features for typical scenarios such as credit card fraud, telecommunications fraud, and account theft, generating scenario-based risk assessment rules. Transfer learning algorithms are used to apply the identified fraud feature patterns to new transaction data to identify potential risky transactions and calculate risk scoring curves. The scenario-based risk assessment rules and risk scoring curves are integrated into a risk control decision framework, generating visualized risk reports and an early warning indicator system, forming valuable insights into financial risk control.
[0033] Based on financial risk control value insights, the data structure is standardized and converted into a common data format for inter-institutional use, generating a value tag index table. A distributed node network is built using a microservice architecture, with federated agent programs deployed at each participating institution to establish secure communication channels and data protocols. The value tag index table is anonymized, removing direct identifiers and retaining only the feature fields required for analysis, forming a shareable data template. Based on this template, a federated learning training process is constructed, allowing each institution to train model parameters locally, transmitting only gradient information rather than raw data, generating a global risk control model. A differential privacy algorithm is used to add precisely calculated random noise to the query results, balancing data sensitivity and privacy budget, providing financial-grade differential privacy protection. Blockchain smart contracts are used to record data sharing operations and contribution calculations, enabling traceable management of data usage permissions and building a compliant financial data asset exchange platform.
[0034] Taking credit card transaction data management as an example, a financial institution scanned its credit card transaction system to extract user spending behavior and transaction data, and generated a data asset catalog through multi-dimensional risk clustering analysis. Based on this catalog, a risk value coefficient was calculated, and a risk heatmap was generated to display high-risk transaction areas. The system identified an abnormal spending pattern of multiple small transactions in the early morning hours and established corresponding security protection mechanisms. By monitoring data health, it was found that the integrity of some historical transaction data was compromised, triggering a repair process. The value mining engine identified a typical fraud pattern of "short-term spending in multiple locations" and generated risk control insight data. The institution shared these insights with other financial institutions through federated learning. Each institution only exchanged model parameters, not raw data, ultimately building a cross-institutional data asset exchange platform that protects user privacy.
[0035] In one specific embodiment, the process of performing step S101 may specifically include the following steps:
[0036] (1) Connect to the database interface of the financial trading software system through adaptive crawling technology to collect user login frequency data, transaction operation trajectory data and fund flow records;
[0037] (2) Calculate user activity index based on collected user login frequency data, extract operation time sequence features from transaction operation trajectory data, and generate transaction association network based on fund flow records;
[0038] (3) Statistical analysis of user activity indicators was conducted to divide users into high-frequency, medium-frequency and low-frequency groups, forming a hierarchical structure of user behavior.
[0039] (4) Input the operation time sequence features into the sliding window algorithm, calculate the operation time sequence similarity matrix, identify typical transaction behavior patterns, and form a transaction behavior feature library;
[0040] (5) Construct a financial data flow diagram based on the transaction association network, mark the dense and sparse nodes of funds, and generate the distribution of fund flow hotspots;
[0041] (6) The hierarchical structure of user behavior, the feature library of transaction behavior and the distribution of hot spots of capital flow are multidimensionally integrated by K-means clustering algorithm to generate a financial software data asset catalog.
[0042] Specifically, adaptive web crawling technology is an intelligent crawling technology that can automatically adjust its crawling strategy according to the target system architecture. It connects to the database interface of financial trading software systems through various methods such as API connections, direct database connections, and front-end page crawling. The core of adaptive web crawling technology lies in its dynamic configuration capability, which can adjust request parameters and frequency according to the interface response characteristics, automatically identify changes in data structure, and adjust the parsing logic accordingly. After the crawler connects to the financial trading system, it collects three types of key data: user login frequency data (including login time, duration, and operating device information), transaction operation trajectory data (including operation type, operation sequence, and operation time interval), and fund flow records (including transaction amount, counterparty, transaction type, and transaction time).
[0043] Subsequently, in-depth processing is performed on the collected raw data. The user activity metric is calculated using a weighted frequency formula:
[0044]
[0045] in, This represents the activity metric of user x. This represents the frequency of user logins within period p. The time weighting factor represents the period p (the weighting is higher for recent periods). This represents the diversity factor of operating devices within period p. The extraction of operation time sequence features from transaction operation trajectory data employs sequence pattern analysis, recording the timestamp, interval, and type of each operation to construct an operation sequence vector. ,in This represents the code for the q-th operation type. This represents the time interval between the previous operation and the current operation. Fund flow records are transformed into a transaction network using graph theory, with each account as a network node and transactions as directed edges. The weight of each edge represents the combined value of the transaction amount and frequency.
[0046] When statistically analyzing the distribution of user activity metrics, the quartile method is used to determine the dividing points. First, the distribution characteristics of all user activity metrics are calculated, yielding the first quartile (Q1), median (Q2), and third quartile (Q3). The high-frequency user group is defined as the set of users with activity metrics greater than Q3; the mid-frequency user group is the set of users with activity metrics between Q1 and Q3; and the low-frequency user group is the set of users with activity metrics less than Q1. This stratification forms a hierarchical structure of user behavior, recording the typical behavioral characteristics and risk preferences of users at each level.
[0047] The operation sequence features are further processed using a sliding window algorithm. This algorithm sets a window size W (e.g., 7 consecutive operations) and slides it across the operation sequence to extract the operation patterns within each window. The similarity is calculated for every two operation windows, forming an operation sequence similarity matrix SM. The similarity calculation comprehensively considers the matching degree of operation type and the similarity of time interval. After threshold filtering and cluster analysis, the similarity matrix identifies typical transaction behavior patterns, such as rapid continuous transfer patterns, dispersed small-amount cash withdrawal patterns, and cross-regional abnormal login patterns. These patterns are stored in a transaction behavior feature database.
[0048] The transaction association network constructs a financial data flow graph through graphical processing, applying the PageRank algorithm and degree centrality analysis to calculate the importance score of each node. Fund-intensive nodes are defined as nodes with in-degree and total transaction amount significantly higher than the average, while fund-sparse nodes are the opposite. Heatmap technology visualizes node importance, generating a distribution of fund flow hotspots and visually displaying concentrated and sparse areas of fund flows. The K-means clustering algorithm, as an unsupervised learning method, is used to multidimensionally integrate the hierarchical structure of user behavior, the feature library of transaction behavior, and the distribution of fund flow hotspots. First, the three types of data are converted into feature vectors. An appropriate distance metric function is defined, and the number of cluster centers, k, is set. An iterative optimization algorithm assigns data points to the nearest cluster centers, and the cluster centers are recalculated until convergence. The clustering results form a financial software data asset catalog, where each cluster represents a type of data asset with similar characteristics. The catalog records key attributes such as asset type, risk level, data volume, and update frequency.
[0049] Taking the transaction data management of an online payment platform as an example, an adaptive web crawler was used to connect to its transaction database, collecting user login records, payment operation sequences, and fund transaction records from the past three months. When calculating user activity metrics, the weight of login frequency and the recent week was set to 1.5 times, and the diversity factor for multi-device login was set to 1.2 times, resulting in a user activity distribution. Using the quartile method, users were divided into three groups: high-frequency merchant users (more than 10 operations per day), medium-frequency ordinary users (2-10 operations per day), and low-frequency dormant users (less than 2 operations per day). A sliding window algorithm was used, setting the window size to 5 consecutive operations, to extract the typical suspicious operation pattern of "multiple small-amount rapid transfers" and store it in the feature database. Based on the transaction association network, several fund-intensive nodes were identified, one of which received more than 500 small-amount transfer transactions daily. The K-means algorithm, with k=8, is used for clustering. The above data is then integrated into a financial software data asset catalog, clearly distinguishing different categories such as high-risk transaction data, abnormal user behavior data, and compliance audit data, laying the foundation for subsequent data asset risk assessment.
[0050] In one specific embodiment, the process of performing step S102 may specifically include the following steps:
[0051] (1) Extract data sensitivity level identifiers, data update cycle values and data access frequency indicators from the financial software data asset catalog, and construct a three-dimensional evaluation parameter matrix;
[0052] (2) The elements in the three-dimensional evaluation parameter matrix are weighted by the analytic hierarchy process to obtain the data sensitivity weight, timeliness weight and use value weight;
[0053] (3) Establish regulatory compliance scoring standards in accordance with national financial regulatory provisions and industry compliance standards, and score each data asset in the financial software data asset catalog for compliance.
[0054] (4) Using fuzzy set theory, the data sensitivity weight, timeliness weight, and use value weight are combined with the compliance score to perform fuzzy comprehensive calculation to obtain the data risk value coefficient;
[0055] (5) By conducting time series analysis on the risk value coefficient of data through Markov decision process, the changing trend of risk coefficient of each data asset is predicted, and a risk evolution trajectory is formed;
[0056] (6) Map the data risk value coefficient and risk evolution trajectory onto a color gradient spectrum, and make spatial arrangements according to data type and business area to draw a heat map of financial data asset risk.
[0057] Specifically, three key parameters are extracted from the financial software data asset catalog: data sensitivity level identifier, data update cycle value, and data access frequency index. The data sensitivity level identifier is a tiered label based on the privacy and importance of the data content, typically categorized into four levels: public data, internal data, sensitive data, and core data. The data update cycle value reflects the timeliness of the data asset, recording the data generation frequency and update interval, such as real-time data, daily updated data, weekly updated data, and historical archived data. The data access frequency index quantifies the frequency and intensity of access to various types of data, derived from data access log statistics. These three parameters are organized into a three-dimensional evaluation parameter matrix, where each element represents the evaluation parameter value of a data asset across the three dimensions. Subsequently, the Analytic Hierarchy Process (AHP) is used to weight the elements in the three-dimensional evaluation parameter matrix. The AHP is a decision-making method that decomposes complex decision problems into multi-level structures. First, a judgment matrix is established, comparing the importance of each of the three parameters pairwise and assigning comparison values (integers from 1 to 9 and their reciprocals), thus constructing judgment matrix A. The judgment matrix's largest eigenvalue and corresponding eigenvector are calculated, and after normalization, data sensitivity weight, timeliness weight, and use value weight are obtained. The consistency ratio of the judgment matrix must be less than 0.1 to ensure the rationality of weight allocation. This process requires the participation of financial experts in the evaluation to ensure that the weight settings conform to the actual situation in the industry. A regulatory compliance scoring standard is established based on national financial regulatory provisions and industry compliance standards. This standard integrates the requirements of regulatory documents such as the "Financial Data Security Protection Regulations" and the "Guidelines for Data Management of Banking Financial Institutions," as well as best practice guidelines for data governance formulated by industry associations. Based on these standards, each data asset in the financial software data asset catalog is scored for compliance. The scoring content includes multiple aspects such as the legality of data collection, storage security, authorized use, standardized sharing, and timely destruction. Each aspect is scored out of 100 points, and the weighted average is taken as the total compliance score for the data asset.
[0058] Subsequently, fuzzy set theory is used to perform fuzzy comprehensive calculations on the weights and compliance scores to obtain the data risk value coefficient. Fuzzy set theory is an effective tool for handling uncertainty and fuzziness, quantifying fuzzy concepts by defining membership functions. First, the scores of the three dimensions of data sensitivity, timeliness, and use value are converted into fuzzy sets, establishing a set of evaluation factors and a set of evaluation levels, and constructing a fuzzy relation matrix. Then, the weight vector and the fuzzy relation matrix are subjected to fuzzy synthesis operations, and a weighted average operator is used to calculate the data risk value coefficient. This coefficient comprehensively reflects the combined characteristics of data assets in terms of both risk and value. Next, a time-series analysis of the data risk value coefficient is conducted using a Markov decision process. A Markov decision process is a stochastic dynamic system model that considers state transition probabilities and decision rewards. It establishes a state transition matrix by analyzing the changes in historical data risk coefficients. States are defined as different intervals of the data risk value coefficient, and the transition probabilities between different states are calculated based on historical data. Through matrix iterative calculation, the changing trends of the risk coefficients of each data asset at future points in time are predicted, forming a risk evolution trajectory. This dynamic analysis helps to anticipate potential risk hotspots and prepare for risk management in advance.
[0059] A risk heatmap of financial data assets is created by mapping data risk value coefficients and risk evolution trajectories onto a color gradient spectrum. The color gradient spectrum uses a gradient scheme from green (low risk) to yellow (medium risk) and then to red (high risk), with risk coefficient values displayed intuitively through color coding. The spatial layout of the heatmap is organized according to data type and business domain. The horizontal axis represents different business domains (such as payments, credit, and investment), and the vertical axis represents different data types (such as customer data, transaction data, and behavioral data). Each data asset is represented by a rectangle in the graph, with the shade of the rectangle representing the level of risk and the size of the rectangle representing the scale of the data asset. The risk evolution trajectory is indicated by arrows or dynamic color changes, indicating the direction of risk trends.
[0060] Taking personal loan data from a commercial bank as an example, parameters such as the sensitivity level identifier (highest level 4) of customer identity information, the update cycle value of transaction flow (real-time update), and the access frequency index of the credit scoring model (200 visits per day) were extracted from its financial software data asset catalog to construct an evaluation matrix. The three-dimensional parameters were weighted using the analytic hierarchy process (AHP), with sensitivity weighting at 0.5, timeliness weighting at 0.3, and usability weighting at 0.2. Compliance scoring of customer identity information was performed in accordance with regulatory requirements, achieving a total compliance score of 85 points in areas such as data anonymization and storage encryption. Using fuzzy set theory to combine the weights with the compliance score, the risk value coefficient of customer identity information was calculated to be 0.75 (high risk, high value). Markov decision process analysis examined the risk coefficient of this type of data over the past six months, predicting that the risk coefficient will rise to 0.82 within the next three months, forming an upward risk evolution trajectory. Ultimately, on the heatmap, the customer identity information data was marked as a dark orange area, and an arrow extending towards the red area indicated its rising risk trend, clearly demonstrating the risk status of the data asset and providing an intuitive basis for banks to formulate targeted security protection strategies.
[0061] In one specific embodiment, the process of executing step S103 may specifically include the following steps:
[0062] (1) Extract risk level identifiers from the financial data asset risk heat map, generate a risk data mapping table, and divide the risk data mapping table into high-risk areas, medium-risk areas and low-risk areas;
[0063] (2) Conduct time series analysis on historical transaction records, construct a benchmark model of transaction behavior, and calculate the transaction frequency curve and transaction amount distribution;
[0064] (3) By using deep learning methods to compare and analyze the benchmark model of trading behavior and the current trading behavior, abnormal behavior feature points that deviate from the normal pattern are identified, and an abnormal trading behavior feature library is formed.
[0065] (4) Based on the high-risk area data in the abnormal transaction behavior feature library and risk data mapping table, the isolated forest algorithm is used to detect abnormal points in transaction behavior and generate an abnormal risk score;
[0066] (5) Classify the transaction data into security levels based on the abnormal risk score, formulate access control policies, encryption policies and audit policies for different security levels, and generate a security policy set;
[0067] (6) Combine the security policy set with the blockchain verification mechanism to record and verify the entire transaction data operation and build a hierarchical security protection mechanism for transaction data.
[0068] Specifically, risk level identifiers are extracted from the financial data asset risk heatmap. This process is achieved through a color analysis algorithm, extracting the risk level information represented by different colors in the heatmap. The specific operation involves scanning each data point in the heatmap, reading its RGB color value, mapping the color value to the corresponding risk level (a real number in the 0-1 range), and generating a risk data mapping table containing the data asset ID and risk level value. Subsequently, based on the risk level threshold, the mapping table is divided into three regions: data with a risk level value greater than 0.7 is classified as high-risk, data with a risk level value between 0.3 and 0.7 is classified as medium-risk, and data with a risk level value less than 0.3 is classified as low-risk. This three-zone division method allows for the rational allocation of security resources according to risk levels. Time series analysis is performed on historical transaction records to construct a benchmark model of transaction behavior. Time series analysis is a statistical method for studying the patterns of data change over time. By extracting key elements such as timestamps, transaction types, and transaction amounts from historical transaction data and arranging them chronologically, a time series is formed. The time series data is processed using the exponentially weighted moving average method, giving higher weight to recent data. The trading frequency within each time window is calculated, and a trading frequency curve is plotted. Simultaneously, the kernel density estimation method is used to calculate the probability distribution of trading amounts, generating a trading amount distribution curve. These two curves together constitute a benchmark model of trading behavior, reflecting the statistical characteristics of normal trading behavior.
[0069] A comparative analysis of a benchmark model of trading behavior and current trading behavior is conducted using deep learning methods. The deep learning method employs a Long Short-Term Memory (LSTM) network, a special type of recurrent neural network capable of learning long-term dependencies in sequential data. The LSTM network uses historical trading sequences as training data to learn the temporal patterns of normal trading behavior, including intraday fluctuations in trading frequency, periodic changes, and the distribution characteristics of trading amounts. After training, current trading behavior data is input into the model to calculate the deviation. Transactions with deviations exceeding a preset threshold are marked as outliers. Through cluster analysis, these outliers are grouped according to feature similarity to form an abnormal trading behavior feature library, with each outlier pattern having a clear description and discrimination criteria. Based on the abnormal trading behavior feature library and high-risk area data in a risk data mapping table, an isolated forest algorithm is used to detect abnormal trading behavior. The isolated forest algorithm is an unsupervised anomaly detection method based on the decision tree principle, identifying outliers by constructing isolated trees in a random subspace. The isolated forest algorithm randomly partitions each attribute of the trading data and calculates the average path length required for each data point to be isolated. Normal data points require more partitioning in the tree to be isolated, while outliers, due to their unique characteristics, can often be isolated within a shorter path. An anomaly score is generated for each data point. The anomaly score ranges from 0 to 1, with values closer to 1 indicating a higher probability of an anomaly.
[0070] Transaction data is categorized into security levels based on anomaly risk scores. Specifically, data with anomaly risk scores greater than 0.8 is classified as the highest security level, requiring the most stringent protection measures; data with scores between 0.5 and 0.8 is classified as high security level; data with scores between 0.3 and 0.5 is classified as medium security level; and data with scores less than 0.3 is classified as general security level. Differentiated security strategies are implemented for each security level: highest security level data employs multi-factor authentication, fine-grained access control, and end-to-end operation auditing; high security level data employs strong encrypted storage, tiered access permissions, and regular security scans; medium security level data implements routine encryption and regular backups; and general security level data implements basic access control. These strategies are integrated into a security policy set, containing specific technical implementation parameters and operational procedures. This security policy set is combined with a blockchain verification mechanism to record and verify all transaction data operations. The blockchain verification mechanism uses a consortium blockchain architecture, jointly maintained by multiple nodes within the financial institution. Each transaction data operation generates an operation record, including the operation type, operator ID, operation time, operation content, and authorization credentials. Operation records are uniquely identified using a hash algorithm, packaged into blocks, and added to the blockchain after verification by all nodes. The immutability of the blockchain ensures the authenticity and integrity of the operation records; any unauthorized access to the data will be recorded and cannot be deleted. In this way, a tiered security protection mechanism for transaction data covering the entire data lifecycle is constructed.
[0071] Taking a securities trading platform as an example, risk level identifiers for user transaction data were extracted from its risk heatmap. High-frequency trading data with a risk value of 0.85 was classified as high-risk. Time series analysis of three months of historical transaction records revealed two peaks in normal trading frequency at 9 AM and 1 PM on trading days, while the transaction amount exhibited a log-normal distribution. Through LSTM network comparative analysis, a group of large transactions conducted during off-peak trading hours were identified, deviating significantly from the normal pattern. These transactions were recorded in the abnormal behavior feature database. Combining the high-risk zone data, the Isolation Forest algorithm detected a batch of transaction behaviors originating from the same IP address but operating multiple accounts, with an anomaly score as high as 0.92. Based on this score, these transactions were classified as the highest security level, and strict protection strategies, including two-factor authentication, full operation recording, and transaction limit control, were implemented. All protective measures and transaction operations were recorded on the blockchain. Once a suspicious transaction was detected, a risk warning was immediately triggered, effectively ensuring the security of transaction data.
[0072] In one specific embodiment, the process of executing step S104 may specifically include the following steps:
[0073] (1) Extract data status information from the transaction data hierarchical security protection mechanism, establish a data life cycle status transition table, and divide the data status into five stages: creation period, active period, low frequency period, archiving period and destruction period;
[0074] (2) Calculate the access popularity of transaction data by statistical analysis of data access frequency, form a data access popularity curve, and identify the peak and trough periods of data access;
[0075] (3) Conduct multi-dimensional testing on the integrity, consistency and timeliness of transaction data, calculate data quality scores, and generate transaction data health indicators in combination with business rule compliance.
[0076] (4) Set a health threshold based on the transaction data health index. When the transaction data health index is lower than the threshold, trigger the data repair process to restore and correct the damaged data.
[0077] (5) Based on the data access heat curve and the data life cycle state transition table, allocate high-frequency access data to the high-performance storage layer and migrate low-frequency access data to the lower-cost storage layer to formulate a cold and hot data storage strategy.
[0078] (6) Integrate and link the data lifecycle status transformation table, transaction data health indicators and cold and hot data storage strategies to construct a financial data asset lifecycle management system.
[0079] Specifically, data status information is extracted from the transaction data hierarchical security protection mechanism. This information includes key indicators such as data creation time, most recent access time, access frequency, and data volume changes. By analyzing this status information, a data lifecycle status transition table is established. This table records the transition conditions and rules between different data lifecycle states. Data status is divided into five stages: the creation period refers to the initial generation and storage of data, characterized by frequent writes and rapid growth in data volume; the active period refers to the stage where data is frequently accessed and used, characterized by high read frequency and significant business value; the low-frequency period refers to the stage where data access frequency decreases significantly, but still has some business reference value; the archived period refers to the stage where data is basically no longer accessed by business systems and is mainly used for historical queries and compliance audits; and the destruction period refers to the stage where data has exceeded the statutory retention period and needs to be destroyed according to the prescribed procedures. This five-stage division method can accurately reflect the entire process of data from generation to destruction, providing a basic framework for subsequent management.
[0080] Access frequency is used to calculate the access popularity of transaction data. Data access popularity is an important indicator of how frequently data is used. The calculation method involves counting the number of times each dataset is accessed within a specific time window, and weighting the access frequency based on its recentity. Specifically, a sliding time window (e.g., 7 days, 30 days) is set, and the number of times each dataset is accessed and the last access time are recorded within the window, with more recent accesses given higher weight. By continuously monitoring the changes in access popularity at different time points, a data access popularity curve is plotted, which shows the trend of data access frequency over time. The curve can identify peak periods (times when access popularity increases significantly) and trough periods (times when access popularity decreases significantly), which helps to optimize data storage strategies and resource allocation.
[0081] Multi-dimensional testing is conducted on the completeness, consistency, and timeliness of transaction data. Completeness testing focuses on whether data is missing, truncated, or corrupted, quantified by metrics such as the completeness rate of data records and the fill rate of required fields. Consistency testing focuses on the consistency of internal logical relationships within the data, assessed by verifying cross-table relationships and compliance with business rules. Timeliness testing focuses on the timely updating of data, measured by calculating metrics such as data update delay and real-time deviation. The results of these three dimensions are used to calculate a data quality score using a weighted average. Simultaneously, compliance checks specific to the financial industry, such as anti-money laundering compliance and financial risk control requirements, are incorporated to comprehensively assess the transaction data health index. This index, with a score between 0 and 100, directly reflects the overall health of the data.
[0082] A health threshold is set based on the transaction data health index, typically with 70 points as the warning line and 60 points as the danger line. When the transaction data health index falls below the threshold, a data repair process is automatically triggered. The data repair process includes several key steps: First, problem diagnosis is performed to accurately locate the specific problem points causing the decline in health; then, an appropriate repair strategy is selected based on the problem type, such as completing missing data, correcting data inconsistencies, and updating outdated data; subsequently, repair operations are performed, which may involve data rollback, incremental updates, or complete reconstruction; finally, the repair effect is verified to confirm whether the health index has recovered above the threshold. This automated health monitoring and repair mechanism ensures the continuous reliability of data quality.
[0083] Based on data access popularity curves and data lifecycle state transition tables, a hot and cold data storage strategy is implemented. The core of this strategy is to allocate data with different access levels and lifecycle stages to appropriate storage tiers. Specifically, high-frequency access data during the creation and active phases is allocated to high-performance storage tiers, such as in-memory databases or high-speed SSD storage, to ensure rapid response; low-frequency data is gradually migrated to lower-cost storage tiers, such as ordinary hard disk arrays; archived data is migrated to archive storage systems, such as object storage or tape libraries; and data in the destruction phase is securely erased according to compliance requirements. Data migration between different storage tiers is automatically triggered by changes in data popularity and lifecycle state transitions, requiring no manual intervention and significantly improving storage resource utilization efficiency. The data lifecycle state transition table, transaction data health indicators, and hot and cold data storage strategies are integrated and correlated to construct a financial data asset lifecycle management system. During the integration process, a unified metadata management layer is established to record and track the state changes, health status, and storage location of data assets at each stage of their lifecycle. Simultaneously, a strategy linkage mechanism was implemented. For example, when data transitions from an active period to a low-frequency period, the system automatically adjusts its storage strategy and correspondingly reduces the health monitoring frequency; when health indicators decline, the monitoring frequency is increased and resources are prepared for repair. This integrated management system enables refined management of data assets throughout their entire lifecycle, optimizes storage resource allocation, and ensures the continuous reliability of data quality.
[0084] Taking a bank's credit card transaction data as an example, after extracting status information from its tiered security protection mechanism, a lifecycle status transition table was established, clearly defining five lifecycle stages for the data. Statistical analysis revealed that transaction data was accessed most frequently within 30 days of its generation, forming a significant peak in access activity; access frequency gradually decreased from 30 to 90 days, entering a low-frequency period; and after 180 days, it was essentially no longer accessed by business systems, entering the archiving period. Multidimensional quality inspection found that some historical transaction data had inconsistent field formats due to system upgrades, causing the health index to drop to 65 points, triggering an automatic repair process. Through field format standardization, the health index was restored to 85 points. Based on access activity analysis, the bank stored transaction data from the past 30 days in a high-performance in-memory database, data from 30 to 90 days on ordinary SSDs, data from 90 to 180 days migrated to ordinary hard disk arrays, and historical data older than 180 days was moved to an archive storage system. This complete lifecycle management system not only optimized storage resource allocation but also ensured the reliability of transaction data throughout its entire lifecycle, providing efficient and stable data support for business systems.
[0085] In one specific embodiment, the process of executing step S105 may specifically include the following steps:
[0086] (1) Select active period transaction data from the financial data asset lifecycle management system, construct a transaction behavior link diagram, and mark transaction entity nodes and fund flow paths;
[0087] (2) Perform topological analysis on the transaction behavior chain diagram, identify transaction loops and fund convergence points, and generate a set of suspicious transaction patterns;
[0088] (3) Frequent itemset mining is performed on the suspicious transaction pattern set through association rule analysis, and the support and confidence of different transaction patterns are calculated to form a fraud risk feature library;
[0089] (4) Match and analyze the fraud risk feature library with the financial risk control scenario template, extract key identification features for typical scenarios such as credit card fraud, telecommunications network fraud and account theft, and generate scenario-based risk judgment rules;
[0090] (5) Apply the identified fraud feature patterns to new transaction data through transfer learning algorithms to identify potential risky transactions and calculate risk score curves;
[0091] (6) Integrate scenario-based risk assessment rules and risk scoring curves into a risk control decision framework, generate visualized risk reports and early warning indicator systems, and form financial risk control value insight data.
[0092] Specifically, active transaction data is screened from the financial data asset lifecycle management system. This screening is based on the data's lifecycle status identifier and access frequency indicators. Active transaction data refers to data that has been frequently accessed recently and has high business value, typically including transaction records from the past 90 days. This data is used to construct a transaction behavior link graph, a directed graph structure representing the relationships between transaction entities. Nodes represent transaction entities (such as user accounts and merchants), edges represent fund flow paths, the direction of the edges indicates the flow of funds, and the weight of the edges represents the transaction amount. During the construction process, each transaction record is parsed into a "payer-payee-amount-time" quadruple, and multiple transaction records are combined to form a complete transaction network topology. Topological structure analysis of the transaction behavior link graph aims to identify abnormal transaction patterns. Topological structure analysis employs various network analysis algorithms, including centrality analysis, community detection, and anomaly subgraph detection. Centrality analysis calculates the degree centrality, betweenness centrality, and eigenvector centrality of nodes to identify key nodes in the network; community detection algorithms divide the network into multiple tightly connected subgroups to discover potential transaction groups; and anomaly subgraph detection searches for structurally abnormal parts of the network. These analyses identified two main types of anomalous patterns: transaction loops, where funds pass through multiple accounts and then return to the initial account, forming a closed loop structure, commonly seen in money laundering activities; and fund convergence points, where funds from multiple transactions ultimately flow to the same account. These nodes exhibit an in-degree much greater than out-degree in the network, commonly seen in fund aggregation in fraudulent activities. These identified anomalous patterns were collected to form a set of suspicious transaction patterns.
[0093] Association rule analysis is a data mining technique that uses frequent itemset mining to discover relationships between data items in a set of suspicious transaction patterns. Association rule analysis defines each element in a transaction pattern as an item (such as transaction time, amount, frequency, etc.), treats each suspicious transaction as a transaction itemset, and then calculates the frequent combinations of these itemsets. In the specific calculation process, support and confidence are two core metrics.
[0094]
[0095]
[0096] in, and Represents the set of transaction feature items. This indicates the proportion of transactions that include both X and Y in the total transactions. This indicates the proportion of transactions containing X that also contain Y. This indicates the number of transactions that include both X and Y. Indicates the total number of transactions. This represents the number of transactions containing X. A lift metric was also introduced to evaluate the effectiveness of the rule.
[0097]
[0098] A lift greater than 1 indicates a rule The probability of its occurrence is greater than the random expectation, which is of practical significance. By setting minimum support and minimum confidence thresholds, meaningful association rules are selected, and fraud risk features are extracted based on these rules to form a fraud risk feature library.
[0099] A fraud risk feature database is matched and analyzed with financial risk control scenario templates. These templates are predefined sets of descriptions of typical fraud scenarios, including behavioral characteristics and discrimination criteria for various fraudulent activities. For different fraud scenarios, such as credit card fraud, telecommunications fraud, and account theft, relevant features are extracted from the fraud risk feature database. The matching process combines semantic similarity calculation and rule matching, scoring the similarity between features in the feature database and features in the scenario templates, selecting features with high similarity as key identification features for that scenario. These key identification features are organized into judgment rules, such as an "if-then" structure rule set. Each rule includes triggering conditions and a corresponding risk level judgment, forming scenario-based risk judgment rules.
[0100] This approach utilizes transfer learning algorithms to apply identified fraud patterns to novel transaction data. Transfer learning is a machine learning method that applies knowledge learned in one domain to another related domain. In this scheme, transfer learning is implemented through feature mapping and model fine-tuning. Feature mapping transforms the feature space of the source domain (known fraud patterns) into the feature space of the target domain (novel transactions), maintaining semantic consistency. Model fine-tuning adjusts the parameters of a model trained in the source domain using a small amount of data from the target domain. Through transfer learning, even without a large number of labeled samples of novel transactions, potential risky transactions can be effectively identified. A risk score is calculated for each transaction, and a risk score curve is plotted over time. This curve shows the temporal evolution trend of transaction risk, helping to identify peak risk periods and abnormal fluctuations.
[0101] This framework integrates scenario-based risk assessment rules with risk scoring curves into a multi-layered decision-making structure, combining the advantages of rule engines and scoring models. The top layer provides an overall risk overview, the middle layer assesses risks in different scenarios, and the bottom layer determines the detailed risks of specific transactions. Based on this framework, a visualized risk report is generated, including risk summaries, risk distribution, lists of abnormal transactions, and risk trends. Simultaneously, an early warning indicator system is constructed, encompassing risk indicator definitions, threshold settings, and alarm trigger conditions. These elements collectively form valuable financial risk control insights, providing decision support for financial institutions' risk management.
[0102] Taking a certain internet finance platform as an example, nearly 60 days of active transaction data were selected from its data asset lifecycle management system to construct a transaction behavior link diagram containing 5,000 user nodes and 15,000 transaction paths. Through topology analysis, 23 transaction loops and 17 fund convergence points were identified, forming 40 suspicious transaction patterns. Association rule analysis calculated the support and confidence of each pattern. For example, the support of "multiple small transfers at night → large cash withdrawal the next day" was 0.02 (meaning this pattern exists in 2% of all transactions), the confidence was 0.75 (meaning that after multiple small transfers at night, 75% will result in a large cash withdrawal the next day), and the lift was 15 (significantly greater than 1, indicating high correlation). This pattern highly matches the credit card cash-out scenario and was extracted as a key identification feature for this scenario. Applying transfer learning algorithms to analyze recently added transactions, 35 high-risk transactions were identified, and the risk score curve showed a significant increase in risky transactions over the weekend. The risk control insights generated by integrating judgment rules and scoring curves successfully helped the platform intercept multiple fraud attempts, effectively safeguarding the platform's funds.
[0103] In one specific embodiment, the process of executing step S106 may specifically include the following steps:
[0104] (1) Standardize the data structure of financial risk control value insights, convert them into a common data format between institutions, and generate a value tag index table;
[0105] (2) Build a distributed node network through a microservice architecture, deploy federated agent programs in each participating institution, and establish secure communication channels and data protocols;
[0106] (3) Desensitize the value tag index table, remove direct identifiers, retain the feature fields required for analysis, and form a shareable data template;
[0107] (4) Based on the shareable data template, a federated learning training process is constructed, enabling each institution to train model parameters locally, transmitting only gradient information rather than raw data, and generating a global risk control model;
[0108] (5) By adding precisely calculated random noise to the query results through differential privacy algorithm, the balance between data sensitivity and privacy budget is controlled, providing financial-grade differential privacy protection;
[0109] (6) Use blockchain smart contracts to record data sharing operations and contribution calculations to achieve traceable management of data usage rights and build a compliant financial data asset exchange platform.
[0110] Specifically, standardizing the data structure of financial risk control value insights is fundamental to achieving data interoperability between institutions. Data structure standardization refers to converting risk control value insight data from different sources and in different formats into a unified, standardized data format that conforms to common financial industry standards. This process includes standardized field naming, unified data types, and consistent data encoding. For example, various time formats are unified to the ISO 8601 standard format, different transaction type codes are mapped to a unified encoding system, and risk scores are standardized to a range of 0-1. The standardized data is organized into a value tag index table, a structured data organization where each row represents a risk control insight point, and columns contain key information such as risk type, risk level, related transaction characteristics, trigger conditions, and recommended actions. The index table uses a hierarchical key-value structure, supporting efficient retrieval and update operations while maintaining the complete semantics of the data.
[0111] A distributed node network is built using a microservices architecture. This architecture breaks down system functionality into multiple independently deployed services, each responsible for a specific function and collaborating through lightweight communication mechanisms. A federated agent is deployed across participating financial institutions. This agent is a special type of middleware responsible for coordinating local data processing and cross-institutional communication. The agent includes a data preprocessing module, a model training module, a secure communication module, and a permission management module. Secure communication channels are established using TLS / SSL encryption protocols to achieve end-to-end encrypted data transmission, while a two-way authentication mechanism ensures the trustworthiness of both communicating parties. The data protocol defines the format specifications, interaction flow, and exception handling mechanisms for data exchange, using lightweight data exchange formats such as JSON or Protobuf, and features a detailed message header structure containing metadata such as version information, timestamps, sender ID, receiver ID, and message type. De-identifying the value tag index table is a necessary step in protecting data privacy. De-identification first identifies and categorizes sensitive fields, including direct identifiers (such as account numbers, ID card numbers, and mobile phone numbers) and indirect identifiers (such as age, occupation, and place of residence, which may lead to identity inference). Direct identifiers are completely removed or replaced with random hash values, while indirect identifiers are handled using techniques such as generalization (e.g., replacing precise age with age ranges) or perturbation (adding random noise). Simultaneously, essential feature fields for analysis, such as transaction behavior patterns, risk triggering conditions, and risk scores, are retained. These fields are crucial for training the risk control model but do not directly expose personal identities. By balancing privacy protection and data availability, a shareable data template is created. This template includes field definitions, data types, value ranges, and descriptions of business implications, facilitating understanding and use of shared data by different institutions.
[0112] Federated learning, a technical framework for collaborative modeling while protecting data privacy, is built upon shared data templates. The training process begins by initializing global model parameters on a central server, then distributing the initial model to participating institutions. Each institution trains its model using its local dataset and the initial model parameters; the training process is entirely local, with the raw data remaining within the institution's internal network. After training, each institution only sends updated gradient information (changes in model parameters) back to the central server, without transmitting any raw data. The central server aggregates the gradient information from each institution, updates the global model parameters, and redistributes the updated model to the institutions for the next round of training. This iterative process continues until the model converges, ultimately generating a global risk control model. To prevent gradient information from leaking privacy, secure aggregation and gradient encryption technologies are employed to protect gradients during transmission.
[0113] Differential privacy (DPP) adds precisely calculated random noise to query results, providing a robust data privacy protection mechanism. The core idea of DPP is to ensure that query results do not differ significantly due to the presence or absence of any single data record, thus preventing the leakage of individual information. The implementation first calculates the sensitivity of the query function, i.e., the maximum possible difference in query results between any two datasets that differ by only one record. Then, based on the sensitivity and a preset privacy budget ε (a parameter controlling the strength of privacy protection), the amount of noise to be added is calculated. The noise typically follows a Laplace or Gaussian distribution and is added to the query results to form the protected output. The privacy budget needs precise control; too small a budget will result in excessive noise affecting the usability of the results, while too large a budget will reduce the strength of privacy protection. Through this mechanism, even if an attacker obtains a large number of query results, they cannot infer information about any specific individual, providing financial-grade differential privacy protection.
[0114] By utilizing blockchain smart contracts to record data sharing operations and calculate contributions, blockchain technology, due to its immutability and distributed consensus mechanism, is well-suited as a trusted recording system for data sharing. Each data sharing operation is recorded as a transaction on the blockchain, containing key information such as the data provider, user, purpose of use, authorized scope, and usage period. Smart contracts are pre-programmed, automatically executed protocols stored on the blockchain, responsible for controlling the data sharing process and verifying permissions. Smart contracts also implement the logic for calculating data contributions, evaluating the contributions of each participating institution based on dimensions such as data quality, data volume, and model improvement effects, and distributing rewards according to contribution. This mechanism ensures the entire process of data use is traceable; any unauthorized data access or use beyond the authorized scope is recorded and cannot be tampered with, thus constructing a compliant financial data asset exchange platform.
[0115] Taking a financial consortium as an example, this consortium comprises multiple banks, insurance companies, and securities firms that have jointly established a financial data asset exchange platform. In practical application, the risk control value insight data generated by a bank's anti-fraud system is first standardized, converting the original JSON format into a unified XML format across the consortium to generate an index table containing various fraud risk tags. This data is securely exchanged through a federated agent program deployed across the institutions, which ensures encrypted data transmission and identity authentication across institutions. Before the exchange, the system anonymizes the risk tag index table, replacing user accounts with hash values, removing direct identifiers such as geographical locations, and retaining transaction pattern characteristics. Based on the anonymized data template, each institution trains its risk control model locally, transmitting only the model gradient to the central server. After multiple iterations, a global risk control model capable of identifying various fraud patterns is formed. To protect query privacy, when other institutions query the distribution of a certain type of fraud risk, the system automatically adds random noise, ensuring that individual information cannot be inferred without sacrificing statistical validity. All data sharing activities are recorded through blockchain smart contracts, including information such as the data provider, user, purpose, and time, achieving completely transparent data usage tracking. This platform effectively solves the problem of data silos among financial institutions, protecting the data privacy of all parties while enhancing the risk prevention and control capabilities of the entire financial system.
[0116] The above describes the management method for software data assets in the embodiments of this application. The following describes the management system for software data assets in the embodiments of this application. Please refer to [link / reference]. Figure 2 One embodiment of the management system for software data assets in this application includes:
[0117] The extraction module is used to scan financial transaction software systems, extract user behavior data and transaction flow information, perform multi-dimensional risk clustering analysis on the extracted financial data, and generate a financial software data asset catalog.
[0118] The calculation module is used to load the financial software data asset catalog, calculate the data risk value coefficient through the regulatory compliance evaluation model, and generate a risk heat map of financial data assets.
[0119] The input module is used to input the financial data asset risk heat map into the hierarchical processing module, execute the transaction behavior anomaly analysis algorithm, and establish a hierarchical security protection mechanism for transaction data.
[0120] The adjustment module is used to connect to the transaction data hierarchical security protection mechanism, monitor the timeliness parameters of financial data, calculate the health index of transaction data, adjust the cold and hot data storage strategy, and form a financial data asset lifecycle management system.
[0121] The import module is used to import transaction data from the financial data asset lifecycle management system into the value mining engine, apply anti-fraud scenario identification algorithms, and generate financial risk control value insight data.
[0122] A module is established to process cross-institutional data sharing requests based on the aforementioned financial risk control value insight data, apply federated learning technology, implement financial-grade differential privacy protection, and establish a compliant financial data asset exchange platform.
[0123] Through the collaborative efforts of the aforementioned components, and by scanning the technical characteristics of financial transaction software systems, this solution can comprehensively and automatically identify potential data assets, avoiding omissions and subjectivity inherent in manual identification, and improving the completeness and accuracy of data asset identification. Utilizing the technical features of multi-dimensional risk clustering analysis, data is scientifically classified according to business value, security level, and usage characteristics, forming a structured catalog of financial software data assets, laying the foundation for subsequent management. By using a regulatory compliance evaluation model to assess the value of data assets, quantitative measurement of data value is achieved, shifting data asset management decisions from subjective judgment to data-driven approaches. A visualized financial data asset risk heatmap intuitively displays the distribution of data asset value, facilitating managers to quickly identify high-value data areas. The application of anomaly analysis algorithms for transaction behavior transforms security protection from static rules to dynamic behavior analysis, automatically adjusting security strategies based on changes in user behavior patterns, significantly improving the accuracy of anomaly detection. A tiered security protection mechanism enables differentiated allocation of security resources, avoiding the extremes of resource waste and insufficient protection. By monitoring the timeliness parameters of financial data and calculating the health indicators of transaction data, combined with intelligent adjustments to hot and cold data storage strategies, a lifecycle management system for financial data assets is formed. This system achieves automated management of the entire lifecycle from data creation to destruction, optimizing storage efficiency and ensuring data quality. The application of anti-fraud scenario identification algorithms significantly improves the ability to identify fraudulent behavior, and the resulting financial risk control insights are directly transformed into business value. In particular, the innovative application of federated learning technology and financial differential privacy protection resolves the contradiction between data sharing and privacy protection, enabling institutions to achieve collaborative model training without exposing raw data. A compliant financial data asset exchange platform breaks down data silos, enabling cross-institutional sharing of data value. The extensive application of artificial intelligence algorithms in this solution, such as K-means clustering in data identification and classification, deep learning in behavioral anomaly detection, and federated learning in privacy-preserving data sharing, fully considers the matching of algorithm characteristics with business scenarios. Through reasonable parameter design and model selection, the applicability and effectiveness of the algorithms in financial data asset management are ensured, significantly improving management efficiency and decision-making accuracy.
[0124] Reference Figure 3This invention also provides a computer device, which can be a server, and its internal structure can be as follows: Figure 3 As shown, the computer device includes a processor, memory, display screen, input device, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores the data corresponding to this embodiment. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the above-described method.
[0125] Those skilled in the art will understand that Figure 3 The structures shown are merely block diagrams of some structures related to the present invention and do not constitute a limitation on the computer devices on which the present invention is applied.
[0126] An embodiment of the present invention also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. It is understood that the computer-readable medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0127] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the present invention and embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.
[0128] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0129] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0130] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for managing software data assets, characterized in that, The method for managing software data assets includes: Scan financial transaction software systems, extract user behavior data and transaction flow information, perform multi-dimensional risk clustering analysis on the extracted financial data, and generate a financial software data asset catalog. Load the financial software data asset catalog, calculate the data risk value coefficient through the regulatory compliance evaluation model, and generate a financial data asset risk heat map; The financial data asset risk heatmap is input into the hierarchical processing module, and an abnormal transaction behavior analysis algorithm is executed to establish a hierarchical security protection mechanism for transaction data. This includes: extracting risk level identifiers from the financial data asset risk heatmap, generating a risk data mapping table, and dividing the risk data mapping table into high-risk, medium-risk, and low-risk areas; performing time-series analysis on historical transaction behavior records to construct a transaction behavior benchmark model, calculating transaction frequency curves and transaction amount distributions; comparing and analyzing the transaction behavior benchmark model and current transaction behavior using deep learning methods to identify abnormal behavior feature points deviating from normal patterns, forming an abnormal transaction behavior feature library; based on the abnormal transaction behavior feature library and high-risk area data in the risk data mapping table, detecting abnormal points in transaction behavior using the isolated forest algorithm to generate an abnormal risk score; classifying transaction data into security levels according to the abnormal risk score, formulating access control policies, encryption policies, and audit policies for different security levels, and generating a security policy set; combining the security policy set with a blockchain verification mechanism to record and verify the entire transaction data operation process, thus constructing the hierarchical security protection mechanism for transaction data. By connecting to the aforementioned tiered security protection mechanism for transaction data, monitoring the timeliness parameters of financial data, calculating the health index of transaction data, adjusting the cold and hot data storage strategy, and forming a lifecycle management system for financial data assets; The transaction data in the financial data asset lifecycle management system is imported into the value mining engine, and anti-fraud scenario identification algorithms are applied to generate financial risk control value insight data. Based on the aforementioned financial risk control value insights data, federated learning technology is applied to process cross-institutional data sharing requests, financial-grade differential privacy protection is implemented, and a compliant financial data asset exchange platform is established.
2. The method for managing software data assets according to claim 1, characterized in that, The scanning financial transaction software system extracts user behavior data and transaction flow information, performs multi-dimensional risk clustering analysis on the extracted financial data, and generates a financial software data asset catalog, including: By using adaptive web crawling technology to connect to the database interface of the financial trading software system, data such as user login frequency, transaction operation trajectory data, and fund flow records are collected. The user activity index is calculated based on the collected user login frequency data, the operation time sequence features are extracted from the transaction operation trajectory data, and a transaction association network is generated based on the fund flow record. The distribution statistics of user activity indicators are analyzed to divide users into high-frequency, medium-frequency and low-frequency groups, forming a hierarchical structure of user behavior. Input the operation time series features into the sliding window algorithm, calculate the operation time series similarity matrix, identify typical transaction behavior patterns, and form a transaction behavior feature library; Construct a financial data flow graph based on the transaction association network, mark dense and sparse nodes of funds, and generate the distribution of hot spots of fund flow. The user behavior hierarchy, transaction behavior feature library, and fund flow hotspot distribution are multidimensionally fused using the K-means clustering algorithm to generate the financial software data asset catalog.
3. The method for managing software data assets according to claim 1, characterized in that, The process of loading the financial software data asset catalog, calculating the data risk value coefficient through a regulatory compliance evaluation model, and generating a financial data asset risk heatmap includes: Extract data sensitivity level identifiers, data update cycle values, and data access frequency indicators from the financial software data asset catalog to construct a three-dimensional evaluation parameter matrix; The elements in the three-dimensional evaluation parameter matrix are weighted using the analytic hierarchy process (AHP) to obtain the weights for data sensitivity, timeliness, and use value. Based on national financial regulatory provisions and industry compliance standards, a regulatory compliance scoring standard is established to score each data asset in the financial software data asset catalog for compliance. Fuzzy set theory is used to perform fuzzy comprehensive calculation of data sensitivity weight, timeliness weight, and use value weight with compliance score to obtain the data risk value coefficient; By conducting time-series analysis of the risk value coefficient of data through Markov decision processes, the changing trend of the risk coefficient of each data asset is predicted, and a risk evolution trajectory is formed. The risk value coefficient and risk evolution trajectory of the data are mapped onto a color gradient spectrum, and the spatial layout is carried out according to the data type and business field to draw the risk heat map of the financial data assets.
4. The method for managing software data assets according to claim 1, characterized in that, The aforementioned hierarchical security protection mechanism for transaction data monitors financial data timeliness parameters, calculates transaction data health indicators, adjusts cold and hot data storage strategies, and forms a financial data asset lifecycle management system, including: Data status information is extracted from the transaction data hierarchical security protection mechanism, and a data lifecycle status transformation table is established to divide the data status into five stages: creation period, active period, low frequency period, archiving period, and destruction period. By statistically analyzing the frequency of data access, the access popularity of transaction data is calculated, and a data access popularity curve is generated to identify peak and trough periods of data access. The integrity, consistency, and timeliness of transaction data are tested in multiple dimensions, a data quality score is calculated, and the transaction data health index is generated by combining the compliance with business rules. A health threshold is set based on the transaction data health index. When the transaction data health index is lower than the threshold, a data repair process is triggered to restore and correct the damaged data. Based on the data access popularity curve and the data lifecycle state transition table, high-frequency access data is allocated to the high-performance storage layer, and low-frequency access data is migrated to the lower-cost storage layer to formulate the cold and hot data storage strategy. By integrating and associating the data lifecycle status transformation table, transaction data health indicators, and hot and cold data storage strategies, the financial data asset lifecycle management system is constructed.
5. The method for managing software data assets according to claim 1, characterized in that, The process of importing transaction data from the financial data asset lifecycle management system into a value mining engine, applying anti-fraud scenario identification algorithms, and generating financial risk control value insight data includes: Active period transaction data is selected from the aforementioned financial data asset lifecycle management system to construct a transaction behavior chain diagram, marking transaction entity nodes and fund flow paths; Perform topology analysis on the transaction behavior chain diagram to identify transaction loops and fund convergence points, and generate a set of suspicious transaction patterns; By analyzing association rules, frequent itemset mining is performed on suspicious transaction pattern sets to calculate the support and confidence of different transaction patterns, thus forming a fraud risk feature library. The fraud risk feature database is matched and analyzed with financial risk control scenario templates. Key identification features are extracted for typical scenarios such as credit card fraud, telecommunications network fraud and account theft, and scenario-based risk judgment rules are generated. By applying the identified fraud feature patterns to new transaction data through transfer learning algorithms, potential risky transactions can be identified and risk score curves can be calculated. By integrating scenario-based risk assessment rules and risk scoring curves into a risk control decision framework, a visualized risk report and early warning indicator system are generated, forming the aforementioned financial risk control value insight data.
6. The method for managing software data assets according to claim 1, characterized in that, The process of using federated learning technology to process cross-institutional data sharing requests based on the aforementioned financial risk control value insight data, implementing financial-grade differential privacy protection, and establishing a compliant financial data asset exchange platform includes: The financial risk control value insights are processed through data structure standardization, converted into a common data format between institutions, and a value tag index table is generated. A distributed node network is built through a microservice architecture, and a federated agent program is deployed in each participating institution to establish a secure communication channel and data protocol. The value tag index table is anonymized by removing direct identifiers and retaining the feature fields required for analysis, thus forming a shareable data template. Based on the shared data template, a federated learning training process is constructed, enabling each institution to train model parameters locally, transmitting only gradient information rather than raw data, and generating a global risk control model. By adding precisely calculated random noise to the query results using a differential privacy algorithm, the balance between data sensitivity and privacy budget is controlled, thus providing the financial-grade differential privacy protection. By utilizing blockchain smart contracts to record data sharing operations and contribution calculations, traceable management of data usage permissions is achieved, thereby constructing a compliant financial data asset exchange platform.
7. A management system for software data assets, used to implement the management method for software data assets as described in any one of claims 1-6, characterized in that, The management system for software data assets includes: The extraction module is used to scan financial transaction software systems, extract user behavior data and transaction flow information, perform multi-dimensional risk clustering analysis on the extracted financial data, and generate a financial software data asset catalog. The calculation module is used to load the financial software data asset catalog, calculate the data risk value coefficient through the regulatory compliance evaluation model, and generate a financial data asset risk heat map. The input module is used to input the financial data asset risk heatmap into the hierarchical processing module, execute the transaction behavior anomaly analysis algorithm, and establish a hierarchical security protection mechanism for transaction data. This includes: extracting risk level identifiers from the financial data asset risk heatmap, generating a risk data mapping table, and dividing the risk data mapping table into high-risk, medium-risk, and low-risk areas; performing time-series analysis on historical transaction behavior records, constructing a transaction behavior benchmark model, and calculating the transaction frequency curve and transaction amount distribution; comparing and analyzing the transaction behavior benchmark model and current transaction behavior using deep learning methods to identify abnormal behavior feature points deviating from the normal pattern, forming an abnormal transaction behavior feature library; based on the abnormal transaction behavior feature library and high-risk area data in the risk data mapping table, detecting abnormal points in transaction behavior using the isolated forest algorithm, and generating an abnormal risk score; classifying transaction data into security levels according to the abnormal risk score, formulating access control policies, encryption policies, and audit policies for different security levels, and generating a security policy set; combining the security policy set with a blockchain verification mechanism to record and verify the entire transaction data operation process, thus constructing the hierarchical security protection mechanism for transaction data. The adjustment module is used to connect to the transaction data hierarchical security protection mechanism, monitor the timeliness parameters of financial data, calculate the health index of transaction data, adjust the cold and hot data storage strategy, and form a financial data asset lifecycle management system. The import module is used to import transaction data from the financial data asset lifecycle management system into the value mining engine, apply anti-fraud scenario identification algorithms, and generate financial risk control value insight data. A module is established to process cross-institutional data sharing requests based on the aforementioned financial risk control value insight data, apply federated learning technology, implement financial-grade differential privacy protection, and establish a compliant financial data asset exchange platform.
8. A computer device, characterized in that, The system includes a memory and a processor, the memory storing a computer program executable on the processor, characterized in that the processor, when executing the computer program, implements the management method for software data assets as described in any one of claims 1 to 6.
9. A computer-readable medium having a computer program stored thereon, the computer program causing a processor, when executed by a processor, to perform the management method for software data assets as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Financial asset management system based on data management
CN116777633A
Network transaction supervision platform
CN118982357A