Management method and system for software data assets and computer readable medium

Through the multi-dimensional risk clustering analysis of the financial transaction software system and the application of the regulatory compliance evaluation model, a financial software data asset catalog and risk heat map are generated, and combined with transaction behavior abnormality analysis and federal learning technology, a financial data asset life cycle management system and a compliant financial data asset exchange platform are established, which solves the problems of data identification, value evaluation, security protection and cross-institutional collaborative analysis in the existing technology, and realizes the full life cycle intelligent management of financial software data assets and cross-institutional value collaborative mining.

CN120181985AActive Publication Date: 2025-06-20北京集联软件科技有限公司

Patent Information

Application Number
CN202510298365.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-20
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The existing financial software data asset management technology has the contradiction between data identification and classification, the lack of scientific quantitative methods for data value assessment, the inability to dynamically adjust data security protection strategies, the separation of data life cycle management, and the collaborative analysis of cross-institutional data faces contradictions between data privacy protection and data sharing.

Method used

By scanning the financial transaction software system, user behavior data and transaction flow information are extracted, multi-dimensional risk clustering analysis is carried out, and the financial software data asset catalog is generated. Then, the data risk value coefficient is calculated using the regulatory compliance evaluation model to generate a risk heat map. Combined with the trading behavior abnormality analysis algorithm, establish a hierarchical security protection mechanism for transaction data, monitor the timeliness parameters of financial data, calculate the health indicators of transaction data, adjust the hot and cold data storage strategies, and form a financial data asset life cycle management system. At the same time, a compliant financial data asset exchange platform is established using federal learning technology and financial-level differential privacy protection.

Benefits of technology

It realizes intelligent management of financial software data assets throughout the life cycle and collaborative mining of cross-institutional value, improves the integrity and accuracy of data asset identification, realizes quantitative measurement and dynamic security protection of data value, optimizes storage efficiency and ensures data quality, and solves the contradiction between data sharing and privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181985A_ABST
    Figure CN120181985A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a management method and system for software data assets and a computer readable medium. The method comprises the following steps: scanning transaction software to extract user behaviors and transaction information, and performing risk analysis to generate a data directory; calculating a risk coefficient through a compliance model to form a thermodynamic diagram; executing exception analysis and establishing a hierarchical protection mechanism; monitoring a data timeliness optimization storage strategy to form life cycle management; an anti-fraud algorithm is applied to generate risk control insights; and compliant cross-mechanism data exchange is realized based on insight data. According to the method, full-life-cycle intelligent management and cross-mechanism value collaborative mining of financial software data assets are realized on the premise of ensuring data security and privacy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a management method, system and computer-readable medium for software data assets. Background Art

[0002] With the rapid development of fintech, the amount of data generated by financial software systems has grown explosively, and this data has become the core assets of financial institutions. Existing financial data asset management technologies mainly focus on three aspects: data storage, data query, and data security. In terms of data storage, traditional technologies use a combination of relational databases and distributed file systems to achieve the storage of structured and unstructured data; in terms of data query, existing technologies mainly rely on SQL query language and NoSQL interfaces to provide data access services; in terms of data security, encrypted storage, access control, and firewall isolation are the main protection means. In addition, some financial institutions also use data governance frameworks to classify and manage the quality of data assets, but such management is often limited within a single institution and lacks a cross-institutional collaborative sharing mechanism.

[0003] However, there are obvious deficiencies in existing financial software data asset management technologies. First, the identification and classification of data assets are too static and cannot adapt to the characteristics of the dynamic change of data value over time; second, the evaluation of data value lacks scientific quantification methods, resulting in a mismatch between resource allocation and the actual value of data; third, data security protection strategies are usually preset fixed rules and cannot dynamically adjust the protection level according to data usage behaviors; fourth, data life cycle management is fragmented, making it difficult to achieve full-process intelligent management from creation to destruction; fifth, data value mining is often limited within a single institution, and cross-institutional data collaborative analysis faces the contradiction between data privacy protection and data sharing. These problems seriously limit the mining and utilization efficiency of the value of financial software data assets. Summary of the Invention

[0004] This application provides a management method, system and computer-readable medium for software data assets, which are used to achieve the full-life cycle intelligent management and cross-institutional value collaborative mining of financial software data assets on the premise of ensuring data security and privacy.

[0005] In a first aspect, the present application provides a management method for software data assets. The management method for software data assets includes: scanning a financial transaction software system, extracting user behavior data and transaction flow information, performing multi-dimensional risk clustering analysis on the extracted financial data to generate a financial software data asset catalog; loading the financial software data asset catalog, calculating a data risk value coefficient through a regulatory compliance evaluation model to generate a financial data asset risk heat map; inputting the financial data asset risk heat map into a grading processing module, executing a transaction behavior anomaly analysis algorithm to establish a hierarchical security protection mechanism for transaction data; docking with the hierarchical security protection mechanism for transaction data, monitoring financial data timeliness parameters, calculating a transaction data health index, adjusting the cold and hot data storage strategy to form a financial data asset lifecycle management system; importing the transaction data in the financial data asset lifecycle management system into a value mining engine, applying an anti-fraud scenario recognition algorithm to generate financial risk control value insight data; and based on the financial risk control value insight data, applying federated learning technology to process cross-institutional data sharing requests, performing financial-level differential privacy protection, and establishing a compliant financial data asset exchange platform.

[0006] In a second aspect, the present application provides a management system for software data assets. The management system for software data assets includes: An extraction module, configured to scan a financial transaction software system, extract user behavior data and transaction flow information, perform multi-dimensional risk clustering analysis on the extracted financial data to generate a financial software data asset catalog; A calculation module, configured to load the financial software data asset catalog, calculate a data risk value coefficient through a regulatory compliance evaluation model to generate a financial data asset risk heat map; An input module, configured to input the financial data asset risk heat map into a grading processing module, execute a transaction behavior anomaly analysis algorithm to establish a hierarchical security protection mechanism for transaction data; An adjustment module, configured to dock with the hierarchical security protection mechanism for transaction data, monitor financial data timeliness parameters, calculate a transaction data health index, adjust the cold and hot data storage strategy to form a financial data asset lifecycle management system; An import module, configured to import the transaction data in the financial data asset lifecycle management system into a value mining engine, apply an anti-fraud scenario recognition algorithm to generate financial risk control value insight data; A establishment module, configured to based on the financial risk control value insight data, apply federated learning technology to process cross-institutional data sharing requests, perform financial-level differential privacy protection, and establish a compliant financial data asset exchange platform.

[0007] In a third aspect of the present invention, a computer device is provided, comprising: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor invokes the instructions in the memory to cause the computer device to execute the above-mentioned management method for software data assets.

[0008] In a fourth aspect of the present invention, a computer-readable medium is provided, wherein instructions are stored in the computer-readable medium, and when it runs on a computer, it causes the computer to execute the above-mentioned management method for software data assets.

[0009] In the technical solution provided by this application, by scanning the technical characteristics of the financial transaction software system, this solution can comprehensively and automatically identify potential data assets, avoiding omissions and subjectivity in manual identification, and improving the integrity and accuracy of data asset identification; using the technical characteristics of multi-dimensional risk clustering analysis, data is scientifically classified according to business value, security level, and usage characteristics to form a structured financial software data asset catalog, laying a foundation for subsequent management. By using a regulatory compliance evaluation model to evaluate the value of data assets, the quantitative measurement of data value is realized, enabling data asset management decisions to shift from subjective judgment to data-driven. The visual financial data asset risk heat map intuitively shows the data asset value distribution, facilitating managers to quickly identify high-value data areas. The application of the transaction behavior anomaly analysis algorithm enables the security protection to shift from static rules to dynamic behavior analysis, and can automatically adjust security policies according to changes in user behavior patterns, significantly improving the detection accuracy of abnormal behaviors. The hierarchical security protection mechanism realizes differential security resource allocation, avoiding the two polar problems of resource waste and insufficient protection. By monitoring the timeliness parameters of financial data and calculating the health indicators of transaction data, combined with the intelligent adjustment of the cold and hot data storage strategy, a financial data asset life cycle management system is formed, realizing the full life cycle automated management from data creation to destruction, optimizing the storage efficiency and ensuring the data quality. The application of the anti-fraud scenario recognition algorithm greatly improves the ability to identify fraud behaviors, and the generated financial risk control value insights are directly converted into business value. In particular, the combined application of federated learning technology and financial-level differential privacy protection innovatively solves the contradiction between data sharing and privacy protection, enabling institutions to achieve collaborative model training without exposing the original data. The compliant financial data asset exchange platform breaks data islands and realizes cross-institutional sharing of data value. The wide application of artificial intelligence algorithms in this solution, such as the application of the K-means clustering algorithm in data identification and classification, the application of deep learning in behavior anomaly detection, and the application of federated learning in privacy-protected data sharing, all fully consider the matching of algorithm characteristics and business scenarios. Through reasonable parameter design and model selection, the applicability and effectiveness of the algorithms in financial data asset management are ensured, greatly improving the management efficiency and decision-making accuracy. Brief Description of the Drawings

[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0011] Figure 1 It is a schematic diagram of an embodiment of the management method for software data assets in the embodiments of the present application; Figure 2 It is a schematic diagram of an embodiment of the management system for software data assets in the embodiments of the present application; Figure 3 It is a schematic block diagram of the structure of a computer device in the embodiments of the present invention. Detailed Embodiments

[0012] The embodiments of the present application provide a management method, system and computer-readable medium for software data assets. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0013] For ease of understanding, the following describes the specific process of the embodiments of the present application. Please refer to Figure 1 An embodiment of the management method for software data assets in the embodiments of the present application includes: Step S101: Scan the financial transaction software system, extract user behavior data and transaction flow information, perform multi-dimensional risk clustering analysis on the extracted financial data, and generate a financial software data asset catalog; Step S102: Load the financial software data asset catalog, calculate the data risk value coefficient through the regulatory compliance evaluation model, and generate a financial data asset risk heat map; Step S103: Input the financial data asset risk heat map into the hierarchical processing module, execute the transaction behavior anomaly analysis algorithm, and establish a hierarchical security protection mechanism for transaction data; Step S104: Interface with the hierarchical security protection mechanism for transaction data, monitor the timeliness parameters of financial data, calculate the health index of transaction data, adjust the cold and hot data storage strategy, and form a financial data asset lifecycle management system; Step S105: Import the transaction data in the financial data asset lifecycle management system into the value mining engine, apply the anti-fraud scenario recognition algorithm, and generate financial risk control value insight data; Step S106: Based on the financial risk control value insight data, apply federated learning technology to process cross-institutional data sharing requests, implement financial-level differential privacy protection, and establish a compliant financial data asset exchange platform.

[0014] It can be understood that the execution subject of this application can be a management system for software data assets, or a terminal or a server. Specifically, it is not limited here. This application example is described by taking the server as the execution subject.

[0015] Specifically, a comprehensive scan of the financial trading software system is achieved through the adaptive crawler technology. The adaptive crawler technology is a technology that can automatically adjust the crawling strategy according to the system structure. It connects to various database interfaces within the financial trading software system and collects user login frequency data, transaction operation track data, and fund flow records. The collected user behavior data includes detailed behavior indicators such as user login time, operation sequence, and stay duration; the transaction flow information includes key elements such as transaction time, amount, counterparty, and transaction type. These raw data are then input into the K-means clustering algorithm for multi-dimensional risk clustering analysis. This algorithm divides similar data into the same cluster by calculating the Euclidean distance between data points. The multi-dimensional risk clustering analysis of financial data involves the comprehensive processing of user behavior stratification, transaction pattern recognition, and fund flow analysis, and finally forms a financial software data asset catalog containing data classification, association relationships, and risk levels.

[0016] After loading the financial software data asset catalog, calculate the data risk value coefficient through the multi-level fuzzy comprehensive evaluation model. This model first extracts the data sensitivity level identifier, data update cycle value, and data access frequency index from the catalog to construct a three-dimensional evaluation parameter matrix. Subsequently, the analytic hierarchy process is applied to assign weights to the matrix elements to obtain a weight system reflecting the importance of each factor. The model establishes a scoring standard in combination with national financial supervision regulations and industry compliance standards, and scores each data asset for compliance. Through the fuzzy set theory, the weights and scores are subjected to fuzzy comprehensive calculation to obtain the data risk value coefficient. This coefficient is subjected to time series analysis through the Markov decision process to predict the risk change trend and form a risk evolution trajectory. Finally, the risk coefficient and the evolution trajectory are mapped onto a color gradient spectrum to generate a financial data asset risk heat map that intuitively displays the risk distribution.

[0017] Input the financial data asset risk heat map into the hierarchical processing module, extract the risk level identifiers from it, generate a risk data mapping table, and divide it into three risk zones: high, medium, and low. At the same time, conduct time series analysis on historical transaction behavior records, construct a benchmark model of transaction behavior, calculate the transaction frequency curve and transaction amount distribution. Compare the benchmark model with the current transaction behavior through deep learning methods, identify the abnormal behavior feature points deviating from the normal mode, and form an abnormal transaction behavior feature library. Based on this feature library and the data in the high-risk zone, apply the isolation forest algorithm to detect abnormal transaction points and generate an abnormal risk score. The isolation forest algorithm calculates the isolation degree of sample points by constructing random decision trees, and abnormal points are more likely to be isolated. Divide the security level according to the abnormal risk score, formulate targeted access control policies, encryption policies, and audit policies, combine these security policy sets with the blockchain verification mechanism, and construct a hierarchical security protection mechanism for transaction data. Connect to the hierarchical security protection mechanism for transaction data, establish a data life cycle state transition table, and divide the data state into five stages: creation period, active period, low-frequency period, archival period, and destruction period. Calculate the access heat of transaction data by counting the data access frequency, form an access heat curve, and identify the peak and trough periods of access. Conduct multi-dimensional detection on the integrity, consistency, and timeliness of transaction data, calculate the data quality score, and generate a transaction data health index in combination with the compliance degree of business rules. Set a health threshold to trigger the data repair process, and recover and correct the damaged data. Based on the access heat curve and the life cycle state transition table, realize the dynamic adjustment of the hot and cold data storage strategy, allocate high-frequency access data to the high-performance storage layer, and migrate low-frequency access data to the storage layer with lower costs. Integrate and associate these elements to construct a financial data asset life cycle management system.

[0018] Screen the active period transaction data from the financial data asset life cycle management system, construct a transaction behavior link graph, and mark the transaction entity nodes and the fund flow path. Conduct topological structure analysis on this link graph, identify transaction loops and fund convergence points, and generate a set of suspicious transaction patterns. Conduct frequent item set mining through association rule analysis, calculate the support degree and confidence degree of different transaction patterns, and form a fraud risk feature library. Match and analyze this feature library with the financial risk control scenario template, extract key identification features for typical scenarios such as credit card fraud, telecom network fraud, and account theft, and generate scenario-based risk judgment rules. Apply the identified fraud feature patterns to new transaction data through the transfer learning algorithm, identify potential risk transactions, and calculate the risk score curve. Integrate the scenario-based risk judgment rules and the risk score curve into a risk control decision framework, generate a visual risk report and an early warning indicator system, and form financial risk control value insight data.

[0019] Based on the financial risk control value insight data, standardize its data structure, convert it into a common data format among institutions, and generate a value label index table. Build a distributed node network through a microservices architecture, deploy a federated proxy program in each participating institution, and establish a secure communication channel and data protocol. Perform desensitization processing on the value label index table, remove direct identifiers, and retain the feature fields required for analysis to form a shareable data template. Based on this template, construct a federated learning training process, enabling each institution to train model parameters locally and only transmit gradient information instead of raw data to generate a global risk control model. Add precisely calculated random noise to the query results through the differential privacy algorithm to control the balance between data sensitivity and privacy budget, providing financial-level differential privacy protection. Use blockchain smart contracts to record data sharing operations and contribution calculations, realize traceable management of data usage permissions, and build a compliant financial data asset exchange platform.

[0020] Taking the management of credit card transaction data as an example, a financial institution scans its credit card transaction system, extracts user consumption behavior and transaction data, and generates a data asset catalog through multi-dimensional risk clustering analysis. Calculate the risk value coefficient based on this catalog and generate a risk heat map to display high-risk transaction areas. The system identifies an abnormal consumption pattern of multiple small transactions in the early morning hours and establishes a corresponding security protection mechanism. By monitoring the data health, it is found that the integrity of some historical transaction data is damaged, triggering a repair process. The value mining engine identifies a typical fraud pattern of "consuming in multiple places in a short time" from it and generates risk control insight data. This institution shares these insights with other financial institutions through federated learning, and each institution only exchanges model parameters instead of raw data, and finally constructs a cross-institutional data asset exchange platform that protects user privacy.

[0021] In a specific embodiment, the process of executing step S101 may specifically include the following steps: (1) Connect to the database interface in the financial trading software system through an adaptive crawler technology to collect user login frequency data, transaction operation track data, and fund flow records; (2) Calculate the user activity index based on the collected user login frequency data, extract operation timing features from the transaction operation track data, and generate a transaction association network based on the fund flow records; (3) Conduct distribution statistics on the user activity index, divide it into high-frequency user groups, medium-frequency user groups, and low-frequency user groups, and form a user behavior hierarchical structure; (4) Input the operation timing features into the sliding window algorithm, calculate the operation timing similarity matrix, identify typical transaction behavior patterns, and form a transaction behavior feature library; (5) Construct a financial data flow map based on the transaction association network, mark the capital-intensive nodes and sparse nodes, and generate the distribution of capital flow hotspots; (6) Multidimensionally integrate the user behavior hierarchical structure, transaction behavior feature library, and capital flow hotspot distribution through the K-means clustering algorithm to generate a financial software data asset catalog.

[0022] Specifically, the adaptive crawler technology is an intelligent crawling technology that can automatically adjust the crawling strategy according to the target system architecture. It connects to the database interface of the financial trading software system through various methods such as API connection, direct database connection, and front-end page crawling. The core of the adaptive crawler technology lies in its dynamic configuration ability, which can adjust request parameters and frequencies according to the interface response characteristics, automatically identify changes in data structures, and accordingly adjust the parsing logic. After the crawler connects to the financial trading system, it collects three types of key data: user login frequency data (including login time, duration, and operating device information), transaction operation track data (including operation type, operation sequence, and operation time interval), and capital flow records (including transaction amount, counterparty, transaction type, and transaction time).

[0023] Subsequently, in-depth processing is performed on the collected raw data. The calculation of the user activity index uses a weighted frequency formula:

[0024] where, represents the activity index of user x, represents the login frequency of the user within the period p, represents the time weight factor of the period p (the weight of the recent period is higher), represents the operating device diversity factor within the period p. The extraction of the operation timing characteristics in the transaction operation track data uses sequence pattern analysis, records the timestamp, operation interval, and operation type of each operation, and constructs an operation sequence vector where represents the qth operation type code, represents the time interval from the previous operation. The capital flow records are transformed into a transaction association network through graph theory methods. Each account serves as a network node, and the transaction serves as a directed edge. The weight of the edge represents the comprehensive value of the transaction amount and frequency.

[0025] When conducting distribution statistics on the user activity index, the quartile method is used to determine the demarcation point. First, calculate the distribution characteristics of all user activity indexes to obtain the first quartile value Q1, median Q2, and third quartile value Q3 of the index. The high-frequency user group is defined as the set of users whose activity index is greater than Q3, the medium-frequency user group is the set of users whose activity index is between Q1 and Q3, and the low-frequency user group is the set of users whose activity index is less than Q1. Through this stratification, a user behavior hierarchical structure is formed, and the typical behavior characteristics and risk preferences of each layer of users are recorded.

[0026] The operation timing characteristics are further processed through a sliding window algorithm. This algorithm sets the window size W (such as 7 consecutive operations) and slides it over the operation sequence to extract the operation patterns within each window. The similarity between every two operation windows is calculated to form an operation timing similarity matrix SM. The similarity calculation comprehensively considers the operation type matching degree and the time interval similarity. The similarity matrix undergoes threshold filtering and clustering analysis to identify typical transaction behavior patterns, such as the pattern of rapid consecutive transfers, the pattern of scattered small withdrawals, the pattern of abnormal logins across regions, etc. These patterns are stored in the transaction behavior feature library.

[0027] The transaction association network constructs a financial data flow diagram through graphical processing, and applies the PageRank algorithm and degree centrality analysis to calculate the importance scores of each node. The capital-intensive nodes are defined as the nodes whose in-degree and the total transaction amount are significantly higher than the average level, while the capital-sparse nodes are the opposite. The node importance is visualized through heat mapping technology to generate the hot distribution of capital flow, intuitively showing the concentrated and sparse areas of the capital flow. The K-means clustering algorithm, as an unsupervised learning method, is used to perform multi-dimensional fusion on the user behavior hierarchical structure, the transaction behavior feature library, and the hot distribution of capital flow. First, the three types of data are converted into feature vectors, an appropriate distance metric function is defined, the number of clustering centers k is set, and the data points are assigned to the nearest clustering center through an iterative optimization algorithm, and the clustering centers are recalculated until convergence. The clustering results form a financial software data asset catalog. Each cluster represents a type of data asset with similar characteristics, and the catalog records key attributes such as asset type, risk level, data volume, update frequency, etc.

[0028] Taking the transaction data management of an online payment platform as an example, an adaptive crawler is used to connect to its transaction database to collect user login records, payment operation sequences, and capital transaction records within three months. When calculating the user activity index, the weight of the login frequency in the recent week is set to 1.5 times, and the diversity factor of multi-device logins is set to 1.2 times to obtain the user activity distribution. Through the quartile method, users are divided into three groups: high-frequency merchant users (with more than 10 operations per day), medium-frequency ordinary users (with 2 - 10 operations per day), and low-frequency dormant users (with less than 2 operations per day). The sliding window algorithm sets the window size to 5 consecutive operations, extracts the typical suspicious operation pattern of "multiple small and rapid transfers" and stores it in the feature library. Based on the transaction association network, several capital-intensive nodes are identified, and one of the nodes receives more than 500 small incoming transfer transactions per day. The K-means algorithm sets k = 8 for clustering, fuses the above data into a financial software data asset catalog, clearly divides different categories such as high-risk transaction data, abnormal user behavior data, and compliance audit data, laying a foundation for subsequent data asset risk assessment.

[0029] In a specific embodiment, the process of executing step S102 may specifically include the following steps: (1) Extract the data sensitivity level identifier, data update cycle value, and data access frequency indicator from the financial software data asset catalog to construct a three-dimensional evaluation parameter matrix; (2) Allocate weights to the elements in the three-dimensional evaluation parameter matrix through the analytic hierarchy process to obtain the data sensitivity weight, timeliness weight, and usage value weight; (3) Establish a regulatory compliance scoring standard according to national financial supervision regulations and industry compliance standards, and score each data asset in the financial software data asset catalog for compliance; (4) Use fuzzy set theory to perform fuzzy comprehensive calculations on the data sensitivity weight, timeliness weight, usage value weight, and compliance score to obtain the data risk value coefficient; (5) Perform time series analysis on the data risk value coefficient through the Markov decision process to predict the change trend of the risk coefficient of each data asset and form a risk evolution trajectory; (6) Map the data risk value coefficient and the risk evolution trajectory to a color gradient spectrum, and perform spatial layout according to data types and business areas to draw a financial data asset risk heat map.

[0030] Specifically, three types of key parameters are extracted from the financial software data asset catalog: data sensitivity level identifier, data update cycle value, and data access frequency indicator. The data sensitivity level identifier is a grading mark divided according to the privacy degree and importance of the data content, usually divided into four levels: public data, internal data, sensitive data, and core data. The data update cycle value reflects the timeliness characteristics of the data asset, recording the generation frequency and update interval of the data, such as real-time data, daily updated data, weekly updated data, and historical archived data, etc. The data access frequency indicator quantifies the frequency and usage intensity of various types of data being accessed, which is obtained by statistical analysis of the data access logs. These three types of parameters are organized into a three-dimensional evaluation parameter matrix, and each element in the matrix represents the evaluation parameter values of a data asset in three dimensions. Subsequently, the analytic hierarchy process (AHP) is used to assign weights to the elements in the three-dimensional evaluation parameter matrix. The AHP is a decision-making method that decomposes complex decision-making problems into multi-level structures. First, a judgment matrix is established, the importance of the three parameters is compared pairwise, and a comparison value (an integer from 1 to 9 and its reciprocal) is assigned to construct the judgment matrix A. Calculate the maximum eigenvalue and the corresponding eigenvector of the judgment matrix, and after normalization, obtain the data sensitivity weight, timeliness weight, and usage value weight. The consistency ratio of the judgment matrix needs to be less than 0.1 to ensure the rationality of the weight assignment. This process requires the participation of financial domain experts to ensure that the weight settings conform to the actual situation of the industry. According to national financial supervision regulations and industry compliance standards, a regulatory compliance scoring standard is established. This standard integrates the requirements of regulatory documents such as the "Regulations on the Security Protection of Financial Data" and the "Guidelines for Data Management of Banking Financial Institutions", as well as the best practice guidelines for data governance formulated by industry associations. According to these standards, each data asset in the financial software data asset catalog is scored for compliance, and the scoring content includes multiple aspects such as the legality of data collection, storage security, usage authorization, sharing standardization, and destruction timeliness. Each aspect is scored on a 100-point scale, and finally, the weighted average is taken as the compliance total score of the data asset.

[0031] Subsequently, the fuzzy comprehensive calculation of weights and compliance scores is carried out using fuzzy set theory to obtain the data risk value coefficient. Fuzzy set theory is an effective tool for dealing with uncertainty and ambiguity problems, and it quantifies fuzzy concepts by defining membership functions. First, the scores of the three dimensions of data sensitivity, timeliness, and usage value are converted into fuzzy sets, an evaluation factor set and an evaluation grade set are established, and a fuzzy relation matrix is constructed. Then, the weight vector and the fuzzy relation matrix are subjected to a fuzzy composition operation, and a weighted average operator is used to calculate the data risk value coefficient. This coefficient comprehensively reflects the comprehensive characteristics of data assets in terms of both risk and value. Next, the time series analysis of the data risk value coefficient is carried out through the Markov decision process. The Markov decision process is a stochastic dynamic system model that considers state transition probabilities and decision rewards. It analyzes the changes in historical data risk coefficients to establish a state transition matrix. The state is defined as different intervals of the data risk value coefficient, and the transition probabilities between different states are calculated based on historical data. Through matrix iterative calculation, the change trend of the risk coefficients of each data asset at future time points is predicted to form a risk evolution trajectory. This dynamic analysis helps to foresee potential risk hotspots and take risk control measures in advance.

[0032] The data risk value coefficient and the risk evolution trajectory are mapped onto a color gradient spectrum to draw a heat map of financial data asset risks. The color gradient spectrum adopts a gradient scheme from green (low risk) to yellow (medium risk) and then to red (high risk), and the risk coefficient values are intuitively displayed through color coding. The spatial layout of the heat map is organized according to data types and business domains. The horizontal axis represents different business domains (such as payment, credit, investment, etc.), and the vertical axis represents different data types (such as customer data, transaction data, behavior data, etc.). Each data asset is represented by a rectangular block in the figure. The shade of the rectangle represents the level of the risk coefficient, and the size of the rectangle represents the scale of the data asset. The risk evolution trajectory is represented by arrows or dynamic color changes, indicating the direction of the risk trend.

[0033] Taking the personal loan business data of a certain commercial bank as an example, parameters such as the sensitive level identifier of customer identity information (the highest level is 4), the update cycle value of transaction records (real-time update), and the access frequency index of the credit scoring model (200 times per day on average) were extracted from its financial software data asset catalog to construct an evaluation matrix. The weights of the three-dimensional parameters were assigned through the analytic hierarchy process, and the sensitivity weight was determined to be 0.5, the timeliness weight was 0.3, and the value weight was 0.2. The customer identity information was scored for compliance with reference to regulatory requirements, and a total compliance score of 85 was obtained in terms of data desensitization processing, storage encryption, etc. The weight and the compliance score were combined using fuzzy set theory, and the risk value coefficient of the customer identity information was calculated to be 0.75 (high risk and high value). The Markov decision process analyzed the risk coefficient of this type of data in the past six months and predicted that the risk coefficient would rise to 0.82 within the next three months, forming an upward risk evolution trajectory. Finally, on the heat map, the customer identity information data was marked in the dark orange area, and its rising risk trend was shown by an arrow extending towards the red area, clearly demonstrating the risk status of the data assets and providing an intuitive basis for the bank to formulate targeted security protection strategies.

[0034] In a specific embodiment, the process of executing step S103 may specifically include the following steps: (1) Extract the risk level identifier from the financial data asset risk heat map, generate a risk data mapping table, and divide the risk data mapping table into a high-risk area, a medium-risk area, and a low-risk area; (2) Conduct time series analysis on historical transaction behavior records, construct a transaction behavior benchmark model, and calculate the transaction frequency curve and transaction amount distribution; (3) Through deep learning methods, compare and analyze the transaction behavior benchmark model with the current transaction behavior, identify abnormal behavior feature points deviating from the normal mode, and form an abnormal transaction behavior feature library; (4) Based on the abnormal transaction behavior feature library and the high-risk area data in the risk data mapping table, detect abnormal transaction behavior points through the isolation forest algorithm and generate an abnormal risk score; (5) Divide the transaction data into security levels according to the abnormal risk score, formulate access control policies, encryption policies, and audit policies for different security levels, and generate a security policy set; (6) Combine the security policy set with the blockchain verification mechanism to record and verify the entire process of transaction data operations, and construct a hierarchical security protection mechanism for transaction data.

[0035] Specifically, risk level identifiers are extracted from the financial data asset risk heat map. This process is achieved through a color analysis algorithm to extract the risk level information represented by different colors in the heat map. The specific operation is as follows: Scan each data point in the heat map, read its RGB color value, map the color value to the corresponding risk level (a real number within the range of 0 - 1), and generate a risk data mapping table containing the data asset ID and the risk level value. Subsequently, the mapping table is divided into three regions according to the risk level threshold: data with a risk level value greater than 0.7 is classified into the high - risk area, data with a risk level value between 0.3 and 0.7 is classified into the medium - risk area, and data with a risk level value less than 0.3 is classified into the low - risk area. This three - region division method enables the reasonable allocation of security resources according to the risk degree. Perform time - series analysis on historical transaction behavior records to construct a transaction behavior benchmark model. Time - series analysis is a statistical method for studying the law of data change over time. By extracting key elements such as timestamps, transaction types, and transaction amounts from historical transaction data and arranging them in chronological order to form a time series. The exponentially weighted moving average method is used to process the time - series data, giving higher weights to recent data, calculating the transaction frequency within each time window, and plotting the transaction frequency curve. At the same time, the kernel density estimation method is used to calculate the probability distribution of transaction amounts and generate a transaction amount distribution curve. These two curves together constitute the transaction behavior benchmark model, reflecting the statistical characteristics of normal transaction behavior.

[0036] The trading behavior benchmark model and the current trading behavior are compared and analyzed through deep learning methods. The deep learning method uses a long short-term memory network (LSTM), which is a special type of recurrent neural network that can learn long-term dependencies in sequential data. The LSTM network uses historical trading sequences as training data to learn the temporal patterns of normal trading behavior, including intraday fluctuations in trading frequency, periodic changes, and distribution characteristics of trading amounts. After training, the current trading behavior data is input into the model to calculate the behavior deviation. Transactions with a deviation exceeding a preset threshold are marked as outliers. Through cluster analysis, these outliers are grouped according to feature similarity to form a library of abnormal trading behavior characteristics, and each abnormal pattern has a clear description and discrimination criterion. Based on the library of abnormal trading behavior characteristics and the data in the high-risk area of the risk data mapping table, the isolation forest algorithm is used to detect outliers in trading behavior. The isolation forest algorithm is an unsupervised outlier detection method based on the decision tree principle, which identifies outliers by constructing isolation trees in random subspaces. The isolation forest algorithm randomly splits each attribute of the trading data and calculates the average path length required to isolate each data point. Normal data points require more splits in the tree to be isolated, while outliers, due to their special nature, can often be isolated within a shorter path length. By calculating the anomaly score for each data point, an anomaly risk score is generated. The value range of the anomaly risk score is from 0 to 1, and the closer the value is to 1, the higher the likelihood of anomaly.

[0037] The trading data is classified into security levels according to the abnormal risk score. The specific classification method is as follows: The data with an abnormal risk score greater than 0.8 is classified into the highest security level and requires the implementation of the strictest protection measures; the data with a score between 0.5 and 0.8 is of high security level; the data with a score between 0.3 and 0.5 is of medium security level; the data with a score less than 0.3 is of general security level. For different security levels, differential security policies are formulated: The data of the highest security level adopts multi-factor authentication, fine-grained permission control, and full-process operation auditing; the data of high security level adopts strong encryption storage, access permission grading, and regular security scanning; the data of medium security level implements conventional encryption and regular backup; the data of general security level performs basic access control. These policies are integrated into a security policy set, which includes specific technical implementation parameters and operation processes. The security policy set is combined with the blockchain verification mechanism to record and verify the trading data operations throughout the process. The blockchain verification mechanism adopts a consortium blockchain architecture and is jointly maintained by multiple nodes within the financial institution. Each trading data operation generates an operation record, which includes the operation type, operator ID, operation time, operation content, and permission certificate. The operation record generates a unique identifier through the hash algorithm and is packaged into a block, which is added to the chain after being jointly verified by each node. The immutable feature of the blockchain ensures the authenticity and integrity of the operation record, and any unauthorized access to the data will be recorded and cannot be deleted. In this way, a hierarchical security protection mechanism for trading data covering the entire data life cycle is constructed.

[0038] Taking a certain securities trading platform as an example, the risk level identifier of the user trading data is extracted from its risk heat map, and the high-frequency trading data with a risk value of 0.85 is classified into the high-risk area. Through time series analysis of the historical trading records for three months, it is found that there are two peaks in the normal trading frequency at 9 am and 1 pm on trading days, while the trading amount shows a lognormal distribution. Through comparative analysis using the LSTM network, a group of large transactions conducted during non-trading peak periods is identified, which deviate significantly from the normal pattern, and these transactions are recorded in the abnormal behavior feature library. Combining with the data in the high-risk area, the isolation forest algorithm detects a batch of trading behaviors from the same IP but operating multiple accounts, with an abnormal score as high as 0.92. According to this score, it is classified into the highest security level, and strict protection policies including two-factor authentication, full-process operation recording, and limit control are implemented. All protection measures and trading operations are recorded on the chain, and once a suspicious transaction is found, a risk warning is immediately triggered, effectively ensuring the security of trading data.

[0039] In a specific embodiment, the process of executing step S104 may specifically include the following steps: (1)Extract data status information from the hierarchical security protection mechanism for transaction data, establish a data life cycle status conversion table, and divide the data status into five stages: creation period, active period, low-frequency period, archival period, and destruction period. (2)Calculate the access heat of transaction data through statistical analysis of data access frequency, form a data access heat curve, and identify the peak and trough periods of data access. (3)Conduct multi-dimensional detection on the integrity, consistency, and timeliness of transaction data, calculate the data quality score, and generate a transaction data health index in combination with the compliance degree of business rules. (4)Set a health threshold according to the transaction data health index. When the transaction data health index is lower than the threshold, trigger the data repair process to recover and correct the damaged data. (5)Based on the data access heat curve and the data life cycle status conversion table, allocate high-frequency access data to the high-performance storage layer, migrate low-frequency access data to the storage layer with lower costs, and formulate a cold and hot data storage strategy. (6)Integrate and correlate the data life cycle status conversion table, the transaction data health index, and the cold and hot data storage strategy to construct a financial data asset life cycle management system.

[0040] Specifically, extract data status information from the hierarchical security protection mechanism for transaction data. This information includes key indicators such as the creation time of the data, the most recent access time, the access frequency, and the change in data volume. By analyzing this status information, establish a data life cycle status conversion table, which records the conversion conditions and rules between different life cycle states of the data. The data status is divided into five stages: The creation period refers to the stage of initial data generation and warehousing, characterized by frequent writing and rapid growth of data volume; the active period refers to the stage when the data is frequently accessed and used, characterized by high read frequency and significant business value; the low-frequency period refers to the stage when the data access frequency significantly decreases, but still has certain business reference value; the archival period refers to the stage when the data is basically no longer accessed by the business system and is mainly used for historical queries and compliance audits; the destruction period refers to the stage when the data has exceeded the legal preservation period and needs to be destroyed according to the prescribed procedures. This five-stage division method can accurately reflect the whole process of data from generation to destruction and provide an infrastructure for subsequent management.

[0041] Calculate the access popularity of transaction data through data access frequency statistics. Data access popularity is an important indicator to measure the frequency of data usage. The calculation method is to count the access times of each data set within a specific time window and assign weights based on the recency of the access time. The specific implementation is as follows: Set a sliding time window (such as 7 days, 30 days, etc.), record the number of times each data set is accessed within the window and the last access time, and give higher weights to recent accesses. By continuously monitoring the changes in access popularity at different time points, draw a data access popularity curve, which shows the changing trend of data access frequency over time. Data access peak periods (periods when access popularity rises significantly) and trough periods (periods when access popularity drops significantly) can be identified from the curve, and this identification helps to optimize data storage strategies and resource allocation.

[0042] Conduct multi-dimensional detection on the integrity, consistency, and timeliness of transaction data. Integrity detection mainly focuses on whether there are missing, truncated, or damaged data, and is quantified by verifying indicators such as the completeness rate of data records and the filling rate of required fields; Consistency detection focuses on whether the internal logical relationships of data are coordinated and is evaluated by verifying cross-table relationships, business rule compliance, etc.; Timeliness detection focuses on the timeliness of data updates and is measured by calculating indicators such as data update latency and real-time deviation. The detection results of these three dimensions are used to calculate the data quality score through a weighted average method. At the same time, combined with the compliance check of business rules unique to the financial industry, such as anti-money laundering rule compliance and financial risk control requirements, a transaction data health indicator is comprehensively evaluated. This indicator is a score between 0 and 100, which intuitively reflects the overall health status of the data.

[0043] Set health thresholds according to the transaction data health indicator. Usually, 70 points are set as the warning line and 60 points are set as the danger line. When the transaction data health indicator is lower than the threshold, the data repair process is automatically triggered. The data repair process includes several key steps: First, conduct problem diagnosis to accurately locate the specific problem points that cause the decline in health; Then select the corresponding repair strategy according to the problem type, such as filling in missing data, correcting inconsistent data, and updating outdated data; Subsequently, perform the repair operation, which may involve data rollback, incremental update, or complete reconstruction; Finally, verify the repair effect to confirm whether the health indicator has risen above the threshold. This automated health monitoring and repair mechanism ensures the continuous reliability of data quality.

[0044] Implement a hot and cold data storage strategy based on the data access heat curve and the data life cycle status transition table. The core of the strategy is to allocate data with different heat levels and life cycle stages to appropriate storage tiers. The specific implementation method is as follows: Frequently accessed data during the creation and active periods is allocated to high-performance storage tiers, such as in-memory databases or high-speed SSD storage, to ensure fast response; data during the low-frequency period is gradually migrated to lower-cost storage tiers, such as ordinary hard disk arrays; data during the archival period is migrated to an archival storage system, such as object storage or tape libraries; data during the destruction period is securely erased in accordance with compliance requirements. The migration of data between different storage tiers is automatically triggered by changes in data heat and life cycle status transitions, without manual intervention, greatly improving the utilization efficiency of storage resources. Integrate and correlate the data life cycle status transition table, transaction data health indicators, and the hot and cold data storage strategy to construct a financial data asset life cycle management system. During the integration process, a unified metadata management layer is established to record and track the status changes, health conditions, and storage locations of data assets at each stage of the life cycle. At the same time, a policy linkage mechanism is implemented. For example, when data transitions from the active period to the low-frequency period, the system automatically adjusts its storage strategy and correspondingly reduces the health monitoring frequency; when the health indicator drops, the monitoring frequency is increased and resources are prepared for repair. This integrated management system realizes the refined management of the entire life cycle of data assets, optimizes the storage resource configuration, and ensures the continuous reliability of data quality.

[0045] Taking the credit card transaction data of a certain bank as an example, after extracting the status information from its hierarchical security protection mechanism, a life cycle status transition table is established, clearly dividing the five life cycle stages of the data. Through statistical analysis, it is found that the access frequency of transaction data is the highest within 30 days after generation, forming an obvious access heat peak; it gradually decreases and enters the low-frequency period from 30 to 90 days; after more than 180 days, it is basically no longer accessed by the business system and enters the archival period. Multidimensional quality detection finds that due to system upgrades, the field formats of some historical transaction data are inconsistent, and the health indicator drops to 65 points, triggering an automatic repair process. The health is restored to 85 points through unified field format processing. According to the access heat analysis, the bank stores transaction data within 30 days in a high-performance in-memory database, data from 30 to 90 days on ordinary SSDs, data from 90 to 180 days is migrated to ordinary hard disk arrays, and historical data over 180 days is moved to an archival storage system. This complete life cycle management system not only optimizes the storage resource configuration but also ensures the reliable quality of transaction data throughout the life cycle, providing efficient and stable data support for the business system.

[0046] In a specific embodiment, the process of executing step S105 may specifically include the following steps: (1)Screen the active transaction data from the financial data asset lifecycle management system, construct a transaction behavior link graph, and mark the transaction entity nodes and the fund flow paths; (2)Conduct a topological structure analysis on the transaction behavior link graph, identify transaction loops and fund convergence points, and generate a set of suspicious transaction patterns; (3)Perform frequent itemset mining on the set of suspicious transaction patterns through association rule analysis, calculate the support and confidence of different transaction patterns, and form a fraud risk feature library; (4)Match and analyze the fraud risk feature library with the financial risk control scenario template, extract key identification features for typical scenarios such as credit card fraud, telecom network fraud, and account theft, and generate scenario-based risk judgment rules; (5)Apply the identified fraud feature patterns to new transaction data through a transfer learning algorithm, identify potential risk transactions, and calculate the risk score curve; (6)Integrate the scenario-based risk judgment rules and the risk score curve into a risk control decision-making framework, generate a visual risk report and an early warning indicator system, and form financial risk control value insight data.

[0047] Specifically, the active transaction data is screened from the financial data asset lifecycle management system, and this screening is based on the lifecycle status identifier and access heat index of the data. The active transaction data refers to the data that has been frequently accessed recently and has high business value, usually including transaction records within the last 90 days. These data are used to construct a transaction behavior link graph, which is a directed graph structure representing the relationships between transaction entities. Among them, the nodes represent transaction entities (such as user accounts, merchants), the edges represent the fund flow paths, the direction of the edges represents the fund flow direction, and the weight of the edges represents the transaction amount. During the construction process, each transaction record is parsed into a four-tuple of "payer - payee - amount - time", and multiple transaction records are combined to form a complete transaction network topology structure. Conducting a topological structure analysis on the transaction behavior link graph aims to identify abnormal transaction patterns. The topological structure analysis uses a variety of network analysis algorithms, including centrality analysis, community discovery, and abnormal subgraph detection. Centrality analysis calculates the degree centrality, betweenness centrality, and eigenvector centrality of the nodes to identify the key nodes in the network; the community discovery algorithm divides the network into multiple tightly connected subgroups to discover potential transaction groups; abnormal subgraph detection looks for the abnormally structured parts in the network. Through these analyses, two main types of abnormal patterns are identified: transaction loops, that is, the funds return to the starting account after passing through multiple accounts, forming a closed-loop structure, which is common in money laundering activities; fund convergence points, that is, the funds of multiple transactions ultimately flow to the same account, and such nodes show a much larger in-degree than out-degree in the network, which is common in the fund collection behavior in fraud activities. These identified abnormal patterns are collected to form a set of suspicious transaction patterns.

[0048] Frequent itemset mining is performed on the set of suspicious transaction patterns through association rule analysis, which is a data mining technique for discovering the association relationships between data items. Association rule analysis defines each element in the transaction pattern as an item (such as transaction time characteristics, amount characteristics, frequency characteristics, etc.), regards each suspicious transaction as a transaction itemset, and then calculates the frequent combination patterns of these item sets. In the specific calculation process, support and confidence are two core indicators:

[0049]

[0050] Among them, and represent the transaction feature itemset, represents the proportion of transactions that contain both X and Y in the total transactions, represents the proportion of transactions that contain Y among the transactions that contain X, represents the number of transactions that contain both X and Y, represents the total number of transactions, represents the number of transactions that contain X. The lift index is also introduced to evaluate the effectiveness of the rule:

[0051] A lift greater than 1 indicates that the probability of the rule appears greater than the random expectation and has practical significance. By setting the minimum support and minimum confidence thresholds, meaningful association rules are screened out, and fraud risk characteristics are extracted based on these rules to form a fraud risk characteristic library.

[0052] The fraud risk characteristic library is matched and analyzed with the financial risk control scenario template, which is a pre-defined set of descriptions of typical fraud scenarios, including the behavior characteristics and discriminant criteria of various fraud activities. For different fraud scenarios, such as credit card fraud, telecom network fraud, and account theft, the relevant feature items are extracted from the fraud risk characteristic library. The matching process uses a combination of semantic similarity calculation and rule matching methods to score the similarity between the features in the characteristic library and the features in the scenario template, and the features with high similarity are selected as the key identification features for this scenario. These key identification features are organized into decision rules, such as a rule set in the "if-then" structure, and each rule contains a trigger condition and a corresponding risk level determination, forming scenario-based risk decision rules.

[0053] Apply the identified fraud feature patterns to new transaction data through a transfer learning algorithm. Transfer learning is a machine learning method that can apply the knowledge learned in one domain to another related domain. In this solution, transfer learning is achieved through two ways: feature mapping and model fine-tuning. Feature mapping transforms the feature space of the source domain (known fraud patterns) into the feature space of the target domain (new transactions), maintaining semantic consistency; model fine-tuning adjusts the parameters using a small amount of target domain data based on the model trained in the source domain. Through transfer learning, even in the absence of a large number of labeled samples of new transactions, potential risky transactions can be effectively identified. Calculate the risk score for each transaction and plot the risk score curve in time series. This curve shows the time evolution trend of transaction risks, helping to discover peak risk periods and abnormal fluctuations.

[0054] Integrate the scenario-based risk determination rules and the risk score curve into a risk control decision-making framework. This framework is a multi-level decision-making structure that combines the advantages of a rule engine and a scoring model. The top layer is the overall risk overview, the middle layer is the risk assessment by scenario, and the bottom layer is the detailed risk determination of specific transactions. Based on this framework, generate a visual risk report including risk summary, risk distribution, list of abnormal transactions, and risk trends, and at the same time construct an early warning index system including risk index definitions, threshold settings, and alarm trigger conditions. These contents together form the financial risk control value insight data, providing decision-making support for the risk management of financial institutions.

[0055] Taking an Internet financial platform as an example, screen the active period transaction data of the past 60 days from its data asset lifecycle management system, and construct a transaction behavior link graph containing 5,000 user nodes and 15,000 transaction paths. Through topological analysis, identify 23 transaction loops and 17 fund convergence points, forming 40 suspicious transaction patterns. Association rule analysis calculates the support and confidence of each pattern. For example, the support of "small multiple transfers at night → large cash withdrawals the next day" is 0.02 (that is, 2% of the total transactions have this pattern), the confidence is 0.75 (that is, after making small multiple transfers at night, 75% will make large cash withdrawals the next day), and the lift is 15 (much greater than 1, with high correlation). This pattern highly matches the credit card cashing scenario and is extracted as the key identification feature of this scenario. Apply the transfer learning algorithm to analyze recently added transactions, identify 35 high-risk transactions, and the risk score curve shows that the risk transactions increase significantly on weekends. The risk control insight data generated by integrating the determination rules and the score curve has successfully helped the platform intercept multiple fraud attempts and effectively protected the platform's capital security.

[0056] In a specific embodiment, the process of executing step S106 may specifically include the following steps: (1)Standardize the data structure of financial risk control value insights, convert them into a common data format among institutions, and generate a value label index table; (2)Build a distributed node network through a microservices architecture, deploy a federated proxy program in each participating institution, and establish a secure communication channel and data protocol; (3)Desensitize the value label index table, remove direct identifiers, and retain the feature fields required for analysis to form a shareable data template; (4)Build a federated learning training process based on the shareable data template, enabling each institution to train model parameters locally and only transmit gradient information instead of raw data to generate a global risk control model; (5)Add precisely calculated random noise to the query results through the differential privacy algorithm to control the balance between data sensitivity and privacy budget and provide financial-level differential privacy protection; (6)Use blockchain smart contracts to record data sharing operations and contribution calculations, implement traceable management of data usage permissions, and build a compliant financial data asset exchange platform.

[0057] Specifically, standardizing the data structure of financial risk control value insights is the basis for realizing data interconnection among institutions. Standardizing the data structure means converting risk control value insight data from different sources and formats into a standardized data format that conforms to the general standards of the financial industry. The specific processing process includes standardizing field names, unifying data types, and making data encodings consistent. For example, unify various time formats into the ISO8601 standard format, map different transaction type codes to a unified coding system, and standardize risk scores to an interval value of 0-1. The standardized data is organized into a value label index table, which is a structured data organization form. Each row represents a risk control insight point, and the columns contain key information such as risk type, risk level, associated transaction characteristics, trigger conditions, and recommended actions. The index table design adopts a hierarchical key-value structure, supporting efficient retrieval and update operations while maintaining the complete semantics of the data.

[0058] Build a distributed node network through a microservices architecture. This architecture pattern splits system functions into multiple independently deployed services. Each service is responsible for a specific function and collaborates through a lightweight communication mechanism. Deploy a federated agent program in each participating financial institution. The federated agent program is a special middleware responsible for coordinating local data processing and cross-institutional communication. The agent program includes a data preprocessing module, a model training module, a secure communication module, and a permission management module. The establishment of the secure communication channel uses the TLS / SSL encryption protocol to achieve end-to-end data transmission encryption. At the same time, a two-way authentication mechanism is introduced to ensure the credibility of the identities of both communication parties. The data protocol defines the format specifications, interaction processes, and exception handling mechanisms for data exchange. It uses lightweight data exchange formats such as JSON or Protobuf and designs a detailed message header structure that includes metadata such as version information, timestamp, sender ID, recipient ID, and message type. Desensitizing the value label index table is a necessary step to protect data privacy. Desensitization first identifies and classifies sensitive fields, including direct identifiers (such as account numbers, ID numbers, mobile phone numbers) and indirect identifiers (such as age, occupation, place of residence, etc., information that may lead to identity inference). For direct identifiers, a strategy of completely removing or replacing them with random hash values is adopted. For indirect identifiers, techniques such as generalization (such as replacing the exact age with an age range) or perturbation (adding random noise) are used. At the same time, feature fields required for analysis are retained, such as transaction behavior patterns, risk trigger conditions, and risk scores. These fields are crucial for risk control model training but do not directly expose personal identities. By balancing privacy protection and data availability, a sharable data template is formed. This template includes field definitions, data types, value ranges, and descriptions of business meanings, facilitating different institutions to understand and use shared data.

[0059] Build a federated learning training process based on the sharable data template. Federated learning is a technical framework for collaborative modeling while protecting data privacy. The training process first initializes the global model parameters on the central server and then distributes the initial model to each participating institution. Each institution uses its local dataset and the initial model parameters to train the model. The training process is completely carried out locally, and the original data does not leave the internal network of the institution. After training is completed, each institution only sends back the gradient information (the change values of the model parameters) of the model update to the central server, without transmitting any original data. The central server aggregates the gradient information from each institution, updates the global model parameters, and distributes the updated model to each institution again for the next round of training. This iterative process continues until the model converges, and finally a global risk control model is generated. To prevent the leakage of privacy through gradient information, technologies such as secure aggregation and gradient encryption are also used to protect the gradients during transmission.

[0060] Adding precisely calculated random noise to the query results through differential privacy algorithms, which is a strict data privacy protection mechanism. The core idea of differential privacy is to ensure that the query results do not vary significantly due to the presence or absence of any single data record, thereby preventing the leakage of individual information. The implementation process first calculates the sensitivity of the query function, that is, the maximum possible difference in the query results for any two data sets that differ by only one record. Then, based on the sensitivity and the preset privacy budget ε (a parameter that controls the intensity of privacy protection), the amount of noise to be added is calculated. The noise usually follows a Laplace distribution or a Gaussian distribution and is added to the query results to form the protected output. The privacy budget needs to be precisely controlled. If it is too small, the noise will be too large and affect the usability of the results; if it is too large, the privacy protection intensity will be reduced. Through this mechanism, even if an attacker obtains a large number of query results, they cannot infer any information about a specific individual, providing financial-level differential privacy protection.

[0061] Using blockchain smart contracts to record data sharing operations and contribution degree calculations. Due to its immutability and distributed consensus mechanism, blockchain technology is very suitable as a trusted record system for data sharing. Each data sharing operation is recorded as a transaction on the blockchain, including key information such as the data provider, user, purpose of use, authorization scope, and usage period. Smart contracts are pre-programmed and stored on the blockchain as automatically executable protocols, responsible for executing data sharing process control and permission verification. Smart contracts also implement the calculation logic for data contribution degree, evaluating the contribution of each participating institution based on dimensions such as data quality, data volume, and model improvement effect, and distributing benefits according to the contribution. This mechanism ensures the traceability of the entire process of data use. Any unauthorized data access or use beyond the authorized scope will be recorded and cannot be tampered with, thus constructing a compliant financial data asset exchange platform.

[0062] Taking a certain financial alliance as an example, this alliance includes multiple banks, insurance, and securities institutions, and jointly establishes a financial data asset exchange platform. In practical applications, the risk control value insight data generated by the anti-fraud system of a certain bank is first standardized, converting the original JSON format into the unified XML format of the alliance, and generating an index table containing various fraud risk tags. These data are securely exchanged through the federated agents deployed in each institution, and the agents ensure cross-institutional data transmission encryption and identity authentication. Before the exchange, the system desensitizes the risk tag index table, replaces the user account with a hash value, removes direct identifiers such as geographical location, and retains the transaction pattern characteristics. Based on the desensitized data template, each institution trains a risk control model locally and only transmits the model gradients to the central server. After multiple rounds of iteration, a global risk control model capable of identifying various fraud patterns is formed. To protect query privacy, when other institutions query the distribution of a certain type of fraud risk, the system automatically adds random noise to ensure that individual information cannot be inferred without losing statistical validity. All data sharing activities are recorded through blockchain smart contracts, including information such as data providing institutions, using institutions, purposes, and times, realizing completely transparent data usage tracking. This platform effectively solves the problem of data silos in financial institutions, and while protecting the data privacy of all parties, improves the risk prevention and control capabilities of the entire financial system.

[0063] The management method for software data assets in the embodiments of the present application has been described above. Next, the management system for software data assets in the embodiments of the present application will be described. Please refer to Figure 2 , an embodiment of the management system for software data assets in the embodiments of the present application includes: An extraction module, configured to scan the financial transaction software system, extract user behavior data and transaction flow information, perform multi-dimensional risk clustering analysis on the extracted financial data, and generate a financial software data asset catalog; A calculation module, configured to load the financial software data asset catalog, calculate the data risk value coefficient through a regulatory compliance evaluation model, and generate a financial data asset risk heat map; An input module, configured to input the financial data asset risk heat map into a hierarchical processing module, execute an abnormal transaction behavior analysis algorithm, and establish a hierarchical security protection mechanism for transaction data; An adjustment module, configured to interface with the hierarchical security protection mechanism for transaction data, monitor the timeliness parameters of financial data, calculate the health index of transaction data, adjust the cold and hot data storage strategy, and form a financial data asset life cycle management system; An import module, configured to import the transaction data in the financial data asset life cycle management system into a value mining engine, apply an anti-fraud scenario recognition algorithm, and generate financial risk control value insight data; A building module, configured to process cross-institutional data sharing requests by applying federated learning technology based on the financial risk control value insight data, perform financial-level differential privacy protection, and build a compliant financial data asset exchange platform.

[0064] Through the collaborative cooperation of the above-mentioned various components, by scanning the technical characteristics of the financial transaction software system, this solution can comprehensively and automatically identify potential data assets, avoiding the omissions and subjectivity of manual identification, and improving the integrity and accuracy of data asset identification; using the technical characteristics of multi-dimensional risk clustering analysis, the data is scientifically classified according to business value, security level, and usage characteristics, forming a structured financial software data asset catalog, laying a foundation for subsequent management. Through the regulatory compliance evaluation model to evaluate the value of data assets, the quantitative measurement of data value is realized, making the data asset management decision-making shift from subjective judgment to data-driven. The visual financial data asset risk heat map intuitively shows the data asset value distribution, facilitating managers to quickly identify high-value data areas. The application of the transaction behavior anomaly analysis algorithm makes the security protection shift from static rules to dynamic behavior analysis, and can automatically adjust the security policy according to the changes in the user behavior pattern, significantly improving the detection accuracy of abnormal behaviors. The hierarchical security protection mechanism realizes the differential security resource allocation, avoiding the two polar problems of resource waste and insufficient protection. Monitoring the timeliness parameters of financial data and calculating the health indicators of transaction data, combined with the intelligent adjustment of the cold and hot data storage strategy, forms a financial data asset life cycle management system, realizing the full life cycle automation management from data creation to destruction, optimizing the storage efficiency and ensuring the data quality. The application of the anti-fraud scenario recognition algorithm greatly improves the recognition ability of fraud behaviors, and the generated financial risk control value insight is directly converted into business value. In particular, the combined application of federated learning technology and financial-level differential privacy protection innovatively solves the contradiction between data sharing and privacy protection, enabling institutions to achieve collaborative model training without exposing the original data. The compliant financial data asset exchange platform breaks the data silos and realizes the cross-institutional sharing of data value. The wide application of artificial intelligence algorithms in this solution, such as the application of the K-means clustering algorithm in data identification and classification, the application of deep learning in behavior anomaly detection, and the application of federated learning in privacy protection data sharing, all fully consider the matching of algorithm characteristics and business scenarios. Through reasonable parameter design and model selection, the applicability and effectiveness of the algorithm in financial data asset management are ensured, greatly improving the management efficiency and decision-making accuracy.

[0065] Refer to Figure 3 , In the embodiment of the present invention, a computer device is further provided. The computer device can be a server, and its internal structure can be as Figure 3As shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected via a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the above method is implemented.

[0066] Those skilled in the art can understand that Figure 3 the structure shown in is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.

[0067] An embodiment of the present invention also provides a computer-readable medium, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0068] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database, or other media provided by the present invention and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.

[0069] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, systems, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0070] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0071] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A method for managing software data assets, characterized in that: The management method for software data assets includes: Scan financial transaction software systems, extract user behavior data and transaction flow information, conduct multi-dimensional risk cluster analysis on the extracted financial data, and generate a financial software data asset catalog; Load the financial software data asset catalog, calculate the data risk value coefficient through the regulatory compliance evaluation model, and generate a financial data asset risk heat map; Input the financial data asset risk heat map into the hierarchical processing module, execute the transaction behavior abnormality analysis algorithm, and establish a hierarchical security protection mechanism for transaction data; Connect to the transaction data hierarchical security protection mechanism, monitor the timeliness parameters of financial data, calculate the health indicators of transaction data, adjust the cold and hot data storage strategies, and form a financial data asset life cycle management system; Importing the transaction data in the financial data asset lifecycle management system into the value mining engine, applying the anti-fraud scenario recognition algorithm, and generating financial risk control value insight data; Based on the financial risk control value insight data, federated learning technology is applied to process cross-institutional data sharing requests, perform financial-grade differential privacy protection, and establish a compliant financial data asset exchange platform.

2. The method for managing software data assets according to claim 1, characterized in that: The scanning of the financial transaction software system, extracting user behavior data and transaction flow information, performing multi-dimensional risk cluster analysis on the extracted financial data, and generating a financial software data asset catalogue includes: Through adaptive crawler technology, we connect to the database interface in the financial trading software system to collect user login frequency data, transaction operation track data and fund flow records; Calculate user activity indicators based on collected user login frequency data, extract operation timing features from transaction operation trajectory data, and generate transaction association networks based on fund flow records; Conduct distribution statistics on user activity indicators, divide them into high-frequency user groups, medium-frequency user groups, and low-frequency user groups, and form a hierarchical structure of user behavior; Input the operation timing features into the sliding window algorithm, calculate the operation timing similarity matrix, identify typical transaction behavior patterns, and form a transaction behavior feature library; Construct a financial data flow graph based on the transaction association network, mark capital-intensive nodes and sparse nodes, and generate capital flow hotspot distribution; The user behavior hierarchical structure, transaction behavior feature library and capital flow hotspot distribution are multi-dimensionally integrated through the K-means clustering algorithm to generate the financial software data asset directory.

3. The method for managing software data assets according to claim 1, characterized in that: The step of loading the financial software data asset catalog, calculating the data risk value coefficient through a regulatory compliance evaluation model, and generating a financial data asset risk heat map includes: Extract data sensitivity level identification, data update cycle value and data access frequency index from the financial software data asset directory to construct a three-dimensional evaluation parameter matrix; The elements in the three-dimensional evaluation parameter matrix are weighted by the analytic hierarchy process to obtain the data sensitivity weight, timeliness weight and use value weight. Establish regulatory compliance scoring standards in accordance with national financial regulatory provisions and industry compliance standards, and assign compliance scores to each data asset in the financial software data asset catalog; Fuzzy set theory is used to perform fuzzy comprehensive calculations on data sensitivity weight, timeliness weight, usage value weight and compliance score to obtain the data risk value coefficient; Through the Markov decision process, the data risk value coefficient is analyzed in time series to predict the changing trend of the risk coefficient of each data asset and form a risk evolution trajectory; The data risk value coefficient and risk evolution trajectory are mapped to a color gradient spectrum, spatially arranged according to data types and business areas, and the financial data asset risk heat map is drawn.

4. The method for managing software data assets according to claim 1, characterized in that: The financial data asset risk heat map is input into the hierarchical processing module, the transaction behavior abnormality analysis algorithm is executed, and a hierarchical security protection mechanism for transaction data is established, including: Extracting risk level identifiers from the financial data asset risk heat map, generating a risk data mapping table, and dividing the risk data mapping table into high-risk areas, medium-risk areas, and low-risk areas; Conduct time series analysis on historical transaction records, build a transaction behavior benchmark model, and calculate transaction frequency curves and transaction amount distribution; Through deep learning methods, we compare and analyze the benchmark model of trading behavior and the current trading behavior, identify the abnormal behavior feature points that deviate from the normal mode, and form an abnormal trading behavior feature library; Based on the abnormal transaction behavior feature library and the high-risk area data in the risk data mapping table, the abnormal points of the transaction behavior are detected by the isolation forest algorithm to generate an abnormal risk score; According to the abnormal risk score, the transaction data is divided into security levels, and access control policies, encryption policies and audit policies for different security levels are formulated to generate a security policy set; The security policy set is combined with the blockchain verification mechanism to record and verify the transaction data operations throughout the process, and to build a hierarchical security protection mechanism for the transaction data.

5. The method for managing software data assets according to claim 1, characterized in that: The connection to the transaction data hierarchical security protection mechanism, monitoring of financial data timeliness parameters, calculation of transaction data health indicators, adjustment of hot and cold data storage strategies, and formation of a financial data asset lifecycle management system include: Extract data status information from the transaction data hierarchical security protection mechanism, establish a data life cycle state transition table, and divide the data status into five stages: creation period, active period, low frequency period, archiving period, and destruction period; Calculate the access heat of transaction data through data access frequency statistics to form a data access heat curve and identify the peak and trough periods of data access; Conduct multi-dimensional testing on the integrity, consistency and timeliness of transaction data, calculate data quality scores, and generate transaction data health indicators in combination with business rule compliance; A health threshold is set according to the transaction data health indicator, and when the transaction data health indicator is lower than the threshold, a data repair process is triggered to recover and correct the damaged data; Based on the data access heat curve and the data life cycle state transition table, high-frequency access data is allocated to the high-performance storage layer, and low-frequency access data is migrated to the storage layer with lower cost, and the hot and cold data storage strategy is formulated; The data lifecycle state transition table, transaction data health index and hot and cold data storage strategies are integrated and associated to construct the financial data asset lifecycle management system.

6. The method for managing software data assets according to claim 1, characterized in that: The importing of transaction data in the financial data asset lifecycle management system into the value mining engine, applying the anti-fraud scenario recognition algorithm, and generating financial risk control value insight data includes: Filter active transaction data from the financial data asset lifecycle management system, construct a transaction behavior link diagram, and mark transaction entity nodes and capital flow paths; Perform topological structure analysis on the transaction behavior link graph, identify transaction loops and capital convergence points, and generate a set of suspicious transaction patterns; Through association rule analysis, frequent item set mining is performed on suspicious transaction pattern sets, and the support and confidence of different transaction patterns are calculated to form a fraud risk feature library; Match and analyze the fraud risk feature library with the financial risk control scenario template, extract key identification features for typical scenarios such as credit card fraud, telecommunications network fraud, and account theft, and generate scenario-based risk determination rules; Apply the identified fraud feature patterns to new transaction data through transfer learning algorithms, identify potential risky transactions, and calculate risk scoring curves; Integrate scenario-based risk determination rules and risk scoring curves into a risk control decision-making framework, generate visual risk reports and early warning indicator systems, and form the financial risk control value insight data.

7. The method for managing software data assets according to claim 1, characterized in that: Based on the financial risk control value insight data, the federated learning technology is applied to process cross-institutional data sharing requests, implement financial-level differential privacy protection, and establish a compliant financial data asset exchange platform, including: Standardize the data structure of the financial risk control value insights, convert them into a common data format among institutions, and generate a value tag index table; Build a distributed node network through a microservices architecture, deploy federal agents in each participating institution, and establish secure communication channels and data protocols; Desensitizing the value tag index table, removing direct identifiers, retaining feature fields required for analysis, and forming a shareable data template; Building a federated learning training process based on the sharable data template, so that each institution can train model parameters locally, only transmit gradient information instead of raw data, and generate a global risk control model; By using the differential privacy algorithm, accurately calculated random noise is added to the query results to control the balance between data sensitivity and privacy budget, thus providing the financial-grade differential privacy protection. Blockchain smart contracts are used to record data sharing operations and contribution calculations, to achieve traceable management of data usage rights, and to build the compliant financial data asset exchange platform.

8. A management system for software data assets, used to implement the management method for software data assets as described in any one of claims 1 to 7, characterized in that: The management system for software data assets includes: The extraction module is used to scan the financial transaction software system, extract user behavior data and transaction flow information, conduct multi-dimensional risk cluster analysis on the extracted financial data, and generate a financial software data asset catalog; A calculation module, used to load the financial software data asset catalog, calculate the data risk value coefficient through the regulatory compliance evaluation model, and generate a financial data asset risk heat map; An input module, used to input the financial data asset risk heat map into the hierarchical processing module, execute the transaction behavior abnormality analysis algorithm, and establish a hierarchical security protection mechanism for transaction data; An adjustment module is used to connect to the transaction data hierarchical security protection mechanism, monitor the timeliness parameters of financial data, calculate the health index of transaction data, adjust the cold and hot data storage strategies, and form a financial data asset life cycle management system; An import module, used to import the transaction data in the financial data asset lifecycle management system into a value mining engine, apply an anti-fraud scenario recognition algorithm, and generate financial risk control value insight data; A module is established to process cross-institutional data sharing requests based on the financial risk control value insight data, apply federated learning technology, perform financial-grade differential privacy protection, and establish a compliant financial data asset exchange platform.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program executable on the processor, and wherein the processor implements the method for managing software data assets as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor is enabled to execute the method for managing software data assets according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • System and method for a global peer to peer retirement savings system

    CA3080272A1

  • Financial asset management system based on data management

    CN116777633A

  • Federal learning-based block chain supply chain financial risk control method and system, storage medium and electronic terminal

    CN116823417A

  • Financial information security protection method and system based on artificial intelligence

    CN118862129A

  • Financial anti-fraud business database construction method and system

    CN118916348A

Cited By

  • Dynamic data directory management and construction method in data weaving environment

    CN121277934A

  • Method for dynamic data catalog management and construction in data weaving environments

    CN121277934B

  • Data security management method, system and device, storage medium and program product

    CN121413032A

  • Data management system, data management method, electronic device, and storage medium

    CN121724622A