Financial compliance risk assessment method and system based on big data, and storage medium

Through the collection and standardization of financial transaction data, a multi-level risk assessment index system is built, and a cascaded integrated network is used for risk assessment, the problem of inefficient risk assessment in traditional methods is solved, and the accurate identification and timely warning of financial compliance risks is achieved.

CN120182006AActive Publication Date: 2025-06-20FOSHAN UNIVERSITY

Patent Information

Application Number
CN202510654978.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

Traditional financial compliance risk assessment methods are inefficient, unable to identify potential risks in a timely manner, lack consistency and objectivity, and difficult to deal with rapidly changing risk patterns.

Method used

A financial compliance risk assessment method based on big data is adopted to build a multi-level financial compliance risk assessment index system through the collection and standardization of financial transaction data, customer information data, regulatory requirements data and market data, and a multi-level financial compliance risk assessment index system is built, and a stacked fusion network of decision trees, neural networks and support vector machines are used for risk assessment.

Benefits of technology

It has achieved efficient identification, accurate assessment and timely warning of financial compliance risks, and improved the compliance risk management automation level and scientific decision-making capabilities of financial institutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182006A_ABST
    Figure CN120182006A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a financial compliance risk assessment method and system based on big data, and a storage medium. The method comprises the following steps: collecting and standardizing multi-source data to form a structured data set; constructing a three-level evaluation system comprising risk fields, subclasses and indexes; training the stacked fusion network to obtain an evaluation model; analyzing and generating a risk report in real time; identifying risk points and performing quantitative evaluation; and an early warning threshold value is set, and an early warning signal is generated and sent to a risk management department when the threshold value is exceeded. According to the invention, efficient identification, accurate evaluation and timely early warning of compliance risks are realized, and the automation level and scientific decision-making ability of compliance risk management of financial institutions are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a financial compliance risk assessment method, system and storage medium based on big data. Background Art

[0002] With the rapid development of the financial market and the continuous emergence of financial innovations, financial institutions are facing increasingly complex compliance challenges. Financial regulators have higher and higher compliance requirements for financial institutions, requiring financial institutions to strictly abide by laws and regulations in business operations and prevent various compliance risks. Traditional financial compliance risk assessment methods mainly rely on manual review and limited data analysis. Financial institutions usually evaluate compliance risks through regular compliance inspections, internal audits, manual sampling inspections, etc. With the expansion of business scale and the increase in complexity, the volume of financial transaction data has grown explosively, and the data dimensions and types have also become increasingly diverse. Regulatory requirements are also constantly changing and improving, and the compliance requirements such as anti-money laundering, counter-terrorism financing, and customer identity verification introduced by regulatory agencies in various countries are becoming increasingly strict. Financial institutions need to cope with more comprehensive and refined compliance challenges.

[0003] However, traditional financial compliance risk assessment methods have many limitations. First, the efficiency of manual review is low, making it difficult to handle massive financial transaction data and complex business scenarios, resulting in a lag in risk assessment and the inability to detect potential risks in a timely manner. Second, traditional methods highly rely on expert experience and subjective judgment, and the judgment criteria of different assessors may vary, resulting in a lack of consistency and objectivity in assessment results. Third, traditional methods are often static and periodic inspections, lacking the ability to monitor the dynamic changes of risks in real time and are difficult to cope with rapidly evolving risk patterns. In addition, traditional methods usually focus on the assessment of a single risk dimension, ignoring the correlation and conduction effects between risks and being unable to comprehensively grasp the systemic characteristics of risks. Finally, traditional methods are relatively weak in risk quantification and early warning mechanisms, making it difficult to achieve accurate measurement and timely warning of risks, resulting in passive and lagged risk disposal. Summary of the Invention

[0004] This application provides a financial compliance risk assessment method, system and storage medium based on big data, which is used to achieve efficient identification, accurate assessment and timely warning of compliance risks, and improve the automation level and scientific decision-making ability of financial institution compliance risk management.

[0005] In a first aspect, this application provides a financial compliance risk assessment method based on big data. The financial compliance risk assessment method based on big data includes: collecting and standardizing financial transaction data, customer information data, regulatory requirement data and market data to obtain a structured data set; Construct a multi-level financial compliance risk assessment index system based on the structured data set. The index system is a three-level structure including risk areas, risk sub-categories, and risk indicators. Use the structured data set and the multi-level financial compliance risk assessment index system to train a cascaded fusion network including decision trees, neural networks, and support vector machines to obtain a financial compliance risk assessment model. Input the structured data set into the financial compliance risk assessment model for multi-dimensional real-time analysis to obtain a compliance risk assessment report. Identify potential compliance risk points according to the compliance risk assessment report, and assign weights to the compliance risk points for quantitative assessment to obtain the quantitative assessment results of risk points. Set a graded warning threshold based on the quantitative assessment results of risk points. When the risk quantification index exceeds the preset threshold, generate a warning signal including a risk point description, a risk level, and a handling suggestion, and send the warning signal to the corresponding risk management department.

[0006] Second, this application provides a financial compliance risk assessment system based on big data. The financial compliance risk assessment system based on big data includes: A collection module for collecting and standardizing financial transaction data, customer information data, regulatory requirement data, and market data to obtain a structured data set; A construction module for constructing a multi-level financial compliance risk assessment index system based on the structured data set. The index system is a three-level structure including risk areas, risk sub-categories, and risk indicators; A training module for using the structured data set and the multi-level financial compliance risk assessment index system to train a cascaded fusion network including decision trees, neural networks, and support vector machines to obtain a financial compliance risk assessment model; An analysis module for inputting the structured data set into the financial compliance risk assessment model for multi-dimensional real-time analysis to obtain a compliance risk assessment report; A quantification module for identifying potential compliance risk points according to the compliance risk assessment report, and assigning weights to the compliance risk points for quantitative assessment to obtain the quantitative assessment results of risk points; A generation module for setting a graded warning threshold based on the quantitative assessment results of risk points. When the risk quantification index exceeds the preset threshold, generate a warning signal including a risk point description, a risk level, and a handling suggestion, and send the warning signal to the corresponding risk management department.

[0007] In a third aspect, there is provided a financial compliance risk assessment device based on big data, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor calls the instructions in the memory so that the financial compliance risk assessment device based on big data executes the above-mentioned financial compliance risk assessment method based on big data.

[0008] In a fourth aspect, there is provided a computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium, and when it runs on a computer, it causes the computer to execute the above-mentioned financial compliance risk assessment method based on big data.

[0009] In the technical solution provided by this application, through the organic combination of technical features such as the collection and standardized processing of multi-source data, the construction of a multi-level risk assessment index system, the training and application of a cascaded fusion network, multi-dimensional real-time analysis, quantitative assessment of risk points, and hierarchical early warning, the comprehensive assessment, accurate identification, and timely early warning of financial compliance risks are realized, and remarkable technical effects are achieved. First, through the comprehensive collection and standardized processing of financial transaction data, customer information data, regulatory requirement data, and market data, a unified structured data set is established, solving the problems of single data source, inconsistent format, and uneven quality in traditional methods, and providing a comprehensive and accurate data basis for subsequent risk analysis; Second, the construction of a multi-level financial compliance risk assessment index system realizes the hierarchical management of risk areas, risk sub-categories, and risk indicators, making risk assessment more systematic and structured, and effectively avoiding the defects of scattered indicators and weak relevance in traditional methods; Third, using the structured data set and multi-level index system to train a cascaded fusion network including decision trees, neural networks, and support vector machines, giving full play to the advantages of different machine learning algorithms. The decision tree model is good at handling risk judgments with clear rules, the neural network model is good at capturing complex non-linear relationships, and the support vector machine model performs well in high-dimensional spaces. By integrating the advantages of each algorithm through a fusion architecture, the accuracy and stability of risk identification are significantly improved; At the same time, the structured data set is input into the financial compliance risk assessment model for multi-dimensional real-time analysis, generating a comprehensive compliance risk assessment report, realizing the dynamic and continuous monitoring of risks, and solving the problem of lagging risk assessment in traditional methods; In addition, potential compliance risk points are identified based on the compliance risk assessment report and quantitatively evaluated. Through scientific weight assignment and calculation methods, qualitative risks are transformed into quantitative indicators, making risk assessment more objective and accurate; Finally, based on the quantitative assessment results of risk points, hierarchical early warning thresholds are set, and early warning signals including risk point descriptions, risk levels, and handling suggestions are generated, realizing the early warning and accurate push of risks, and providing a time window and action guide for risk disposal. The technical solution of the present invention applies artificial intelligence algorithms and models in the specific field of financial compliance risk management, fully considering the contributions of different algorithm characteristics to the solution, integrating the advantages of multiple algorithms through a cascaded fusion network, realizing the intelligentization, precision, and automation of compliance risk assessment, effectively solving the problems of strong manual dependence, large subjectivity, and poor real-time performance in traditional methods, and providing a comprehensive, efficient, and accurate compliance risk management solution for financial institutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0011] Figure 1 FIG. is a schematic diagram of an embodiment of the financial compliance risk assessment method based on big data in the embodiments of the present application; Figure 2 FIG. is a schematic diagram of an embodiment of the financial compliance risk assessment system based on big data in the embodiments of the present application; Figure 3 It is a schematic block diagram of the structure of the financial compliance risk assessment device based on big data in the embodiments of the present invention. Detailed implementation manners

[0012] The embodiments of the present application provide a financial compliance risk assessment method, system and storage medium based on big data. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and the above drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "comprising" or "having" and any deformation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0013] For ease of understanding, the following describes the specific process of the embodiments of the present application. Please refer to Figure 1 , an embodiment of the financial compliance risk assessment method based on big data in the embodiments of the present application includes: Step S101, collect and standardize financial transaction data, customer information data, regulatory requirement data and market data to obtain a structured data set; Step S102, construct a multi-level financial compliance risk assessment index system based on the structured data set, and the index system is a three-level structure including risk areas, risk sub-categories and risk indicators; Step S103, use the structured data set and the multi-level financial compliance risk assessment index system to train a cascaded fusion network including decision trees, neural networks and support vector machines to obtain a financial compliance risk assessment model; Step S104: Input the structured data set into the financial compliance risk assessment model for multi-dimensional real-time analysis to obtain a compliance risk assessment report; Step S105: Identify potential compliance risk points based on the compliance risk assessment report, and assign weights to the compliance risk points for quantitative assessment to obtain the quantitative assessment results of the risk points; Step S106: Set a hierarchical warning threshold based on the quantitative assessment results of the risk points. When the risk quantification index exceeds the preset threshold, generate a warning signal including a risk point description, a risk level, and a handling suggestion, and send the warning signal to the corresponding risk management department.

[0014] It can be understood that the execution entity of this application can be a financial compliance risk assessment system based on big data, or a terminal or a server. Specifically, it is not limited here. In this embodiment of the application, the server is used as the execution entity for illustration.

[0015] Specifically, financial transaction data is obtained from the internal system of financial institutions, including transaction records, fund flows, and account information, which record the basic attributes and behavioral characteristics of transactions; customer information data is obtained from the customer management system, including customer basic information, credit records, and transaction preferences, reflecting the identity background and behavior patterns of customers; regulatory requirement data includes legal and regulatory texts, compliance guidelines, and penalty cases, providing a basis for compliance judgment; market data includes market indicators such as interest rates, exchange rates, and stock prices, reflecting changes in the market environment. The collected raw data is cleaned, removing duplicate records, correcting format error data, and handling missing values, and then the format is converted to unify the data structure. The numerical data is normalized to standardize the data range, and the categorical data is encoded and converted to facilitate model processing. Finally, data annotation is performed to add risk-related tags to form a structured data set. A multi-level financial compliance risk assessment index system is constructed based on the structured data set. The financial compliance risks are divided into risk areas such as anti-money laundering compliance risks, customer identity verification risks, internal control risks, market manipulation risks, conflict of interest risks, tax compliance risks, and cross-border business compliance risks, forming a risk area classification framework. Each risk area is further divided to establish risk subcategories, and the association relationship between the area and the subcategories is constructed. Taking the anti-money laundering compliance risk as an example, it is divided into subcategories such as customer due diligence risk, abnormal transaction monitoring risk, and suspicious transaction reporting risk. Based on the secondary risk classification structure, specific risk indicators are designed for each risk subcategory, such as transaction frequency, transaction amount, customer credit score, and regulatory compliance degree numerical indicators, forming a three-level index system. The relative importance of each indicator is determined through the analytic hierarchy process, and weights are assigned to form a weighted index system. According to historical data analysis, a risk threshold interval is formulated, the normal value range, the attention value range, and the warning value range are divided, and a dynamic adjustment mechanism is established to update the indicator definition and threshold setting according to regulatory policy changes and risk events.

[0016] Train a stacked fusion network using a structured dataset and a multi-level financial compliance risk assessment index system. Extract training features from the structured dataset and divide the training set, validation set, and test set in a ratio of 7:2:1. Use the risk indicator values as input features and the corresponding risk levels of the risk indicators as target labels to train a decision tree model, a five-layer feedforward neural network model, and a radial basis kernel support vector machine model respectively to form three basic models. The decision tree model is good at handling classification problems and identifying obvious risk features; the neural network model can capture complex non-linear relationships and hidden risk patterns; the support vector machine model performs well in dealing with high-dimensional data and establishing clear decision boundaries. Apply the three basic models to the validation set, calculate the F1 scores of each model in different risk areas to form a performance matrix. Based on the performance matrix, construct a stacked fusion architecture, select the model with the best performance in each risk area as the main model, and the rest as auxiliary models, and combine the model output results through weighted voting to form an initial fusion network. Conduct grid search parameter optimization on the validation set, adjust the weight coefficients and decision thresholds of each basic model to obtain a fusion network with optimized parameters. Apply the optimized fusion network to the test set, calculate the average precision and recall rates of each risk area to form a financial compliance risk assessment model.

[0017] Input the structured data set into the financial compliance risk assessment model for multi-dimensional real-time analysis. Group the structured data set according to the customer dimension, transaction dimension, business dimension, and institutional dimension, and convert it into a feature matrix according to the index definitions of the multi-level financial compliance risk assessment index system to form multi-dimensional feature data. Batch process and real-time incrementally process the multi-dimensional feature data, and merge the results of the two types of processing to form a complete feature vector. Input the complete feature vector into the decision tree model, five-layer feedforward neural network model, and radial basis kernel support vector machine model in the financial compliance risk assessment model respectively to obtain three groups of preliminary risk assessment results. Based on the weight configuration of the main model and auxiliary model corresponding to each risk area in the cascaded fusion network, perform weighted fusion calculation on the three groups of preliminary risk assessment results to generate the risk scores and risk probability values of each risk index, and form a comprehensive risk assessment result. According to the index hierarchy relationship in the multi-level financial compliance risk assessment index system, aggregate and calculate from the risk index level to the risk sub-category level and risk area level step by step to form a hierarchical risk score table. Divide the risk levels according to the preset score intervals in the hierarchical risk score table, generate a risk distribution heat map, a key risk index trend map, and a risk correlation network map, and combine them to form a compliance risk assessment report. Identify and quantitatively evaluate compliance risk points according to the compliance risk assessment report. Analyze the compliance risk assessment report, extract the risk indicators whose risk scores exceed the preset threshold, including the risk indicator name, risk score, belonging risk sub-category, and risk area, to form a high-risk indicator set. Analyze the risk feature pattern based on the high-risk indicator set, identify abnormal transaction behaviors and compliance hidden dangers in the business process, and form a preliminary risk point list. Trace the data for each risk point in the preliminary risk point list, extract the associated original transaction records, customer information, and relevant business data to form a risk point evidence chain, and obtain a complete risk point file. According to the three core factors of risk severity, impact range, and occurrence frequency in the risk point file, assign a weight coefficient to each risk point, and determine the weight size through a multi-factor comprehensive calculation method to form a risk point weight table. Use the risk point weight table and the risk indicator scores in the compliance risk assessment report to perform composite calculations to obtain the quantitative evaluation value of each risk point, and comprehensively consider direct risks and associated risks to obtain the risk point quantitative value. According to the risk point quantitative value, divide the risk points into four risk levels: critical, high, medium, and low, and sort the risk points according to the risk level to form a risk point quantitative evaluation result.

[0018] Implement a risk early warning mechanism based on the quantitative assessment results of risk points. Conduct statistical analysis on the quantitative assessment results of risk points to determine the boundary of the quantitative score interval for different risk levels, which serves as the benchmark threshold for risk early warning, and form a benchmark threshold table. According to the benchmark threshold table and historical risk handling records, set differentiated early warning thresholds for different risk types and degrees, establish a hierarchical early warning threshold matrix, and form a complete early warning threshold system. Use the complete early warning threshold system to monitor the quantitative assessment results of risk points. When the risk quantification index exceeds the corresponding early warning threshold, trigger the early warning signal generation process to form an original early warning trigger mark. Based on the original early warning trigger mark and the risk point data in the quantitative assessment results of risk points, generate early warning content including risk point description, risk type, occurrence location, influence scope, specific numerical indicators, and development trend, and form a risk description report. Match and analyze the risk description report with the disposal cases in the historical risk database to generate treatment suggestions for different risk types and levels, and integrate them to form an early warning signal package. According to the risk level and business area information in the early warning signal package, determine the early warning receiving departments and personnel, and send the early warning signal package to the corresponding risk management departments via email, text message, and system notification to complete the early warning distribution.

[0019] In the embodiments of the present application, through the organic combination of technical features such as the collection and standardized processing of multi-source data, the construction of a multi-level risk assessment index system, the training and application of a cascaded fusion network, multi-dimensional real-time analysis, quantitative assessment of risk points, and hierarchical early warning, the comprehensive assessment, accurate identification, and timely early warning of financial compliance risks are realized, and remarkable technical effects are achieved. First, through the comprehensive collection and standardized processing of financial transaction data, customer information data, regulatory requirement data, and market data, a unified structured data set is established, solving the problems of single data source, inconsistent format, and uneven quality in traditional methods, and providing a comprehensive and accurate data basis for subsequent risk analysis; second, the construction of a multi-level financial compliance risk assessment index system realizes the hierarchical management of risk areas, risk sub-categories, and risk indicators, making the risk assessment more systematic and structured, and effectively avoiding the defects of scattered indicators and weak relevance in traditional methods; third, using the structured data set and multi-level index system to train a cascaded fusion network including decision trees, neural networks, and support vector machines, giving full play to the advantages of different machine learning algorithms. The decision tree model is good at dealing with risk judgments with clear rules, the neural network model is good at capturing complex non-linear relationships, and the support vector machine model performs well in high-dimensional spaces. By integrating the advantages of each algorithm through the fusion architecture, the accuracy and stability of risk identification are significantly improved; at the same time, the structured data set is input into the financial compliance risk assessment model for multi-dimensional real-time analysis, generating a comprehensive compliance risk assessment report, realizing the dynamic and continuous monitoring of risks, and solving the problem of lagging risk assessment in traditional methods; in addition, potential compliance risk points are identified based on the compliance risk assessment report and quantitatively evaluated. Through scientific weight assignment and calculation methods, qualitative risks are converted into quantitative indicators, making the risk assessment more objective and accurate; finally, based on the quantitative assessment results of risk points, hierarchical early warning thresholds are set, and early warning signals including risk point descriptions, risk levels, and handling suggestions are generated, realizing the early warning and accurate push of risks, and providing a time window and action guide for risk disposal. The technical solution of the present invention applies artificial intelligence algorithms and models in the specific field of financial compliance risk management, fully considering the contributions of different algorithm features to the solution, integrating the advantages of multiple algorithms through a cascaded fusion network, realizing the intelligent, precise, and automated compliance risk assessment, effectively solving the problems of strong manual dependence, large subjectivity, and poor real-time performance in traditional methods, and providing a comprehensive, efficient, and precise compliance risk management solution for financial institutions.

[0020] In a specific embodiment, the process of executing step S101 may specifically include the following steps: Extract financial transaction data such as transaction records, fund flows, and account information from the internal system of the financial institution, and obtain customer information data such as customer basic information, credit records, and transaction preferences from the customer management system to obtain an original data set; Clean the original dataset to obtain a cleaned dataset, and then perform format conversion on the cleaned dataset to obtain a dataset with unified format; Normalize the numerical data in the dataset with unified format, and perform encoding conversion on the categorical data in the dataset with unified format to obtain a standardized dataset; Perform data annotation on the standardized dataset and add risk-related labels to obtain a labeled dataset; Integrate and store the labeled dataset in a data warehouse, and establish a data index and access control mechanism to form a structured dataset.

[0021] Specifically, extract financial transaction data from the internal systems of financial institutions. Connect to the core business systems, transaction processing systems, and account management systems of financial institutions by establishing data interfaces, and extract transaction record data through methods such as API calls, database queries, or file transfers. The transaction record data includes fields such as transaction serial numbers, transaction times, transaction amounts, transaction types, and counterparty information; extract fund flow data, including fields such as fund sources, fund destinations, fund amounts, and transfer times; extract account information data, including fields such as account opening times, account statuses, account balance changes, and associated account information. At the same time, obtain customer information data from the customer management system, including customer basic information, credit records, and transaction preferences. The data is extracted through a database connector and temporarily stored in a data buffer to form an original dataset.

[0022] The original dataset needs to be cleaned. First, identify and process duplicate data. Identify duplicate records by comparing key fields and retain the latest or most complete records; then identify and correct incorrect data, including format errors, value errors, logical errors, etc.; then process missing data. For missing non-key fields, use mean filling, mode filling, or prediction models for filling. For records with missing key fields, mark them as requiring manual review or directly eliminate them. After completing the data cleaning, perform format conversion on the cleaned dataset to convert data from different sources and different structures into a unified format, including unified field naming, unified data types, and unified encoding methods, to form a dataset with unified format.

[0023] Normalize the numerical data in the dataset with unified format, so that numerical data with different dimensions and ranges is converted into a unified scale. The minimum-maximum normalization method is mainly adopted to linearly transform the original data into the interval [0, 1]. For fields vulnerable to outliers, the Z-Score standardization method is used to convert the data into a distribution with a mean of 0 and a standard deviation of 1. For data with skewed distributions such as transaction amounts and transaction frequencies, logarithmic transformation is first performed and then standardization is carried out to make the data distribution closer to the normal distribution. At the same time, perform encoding conversion on the categorical data in the dataset with unified format, and convert non-numerical categorical data into a numerical form that can be processed by the model. For ordered categorical variables, ordinal encoding is adopted; for unordered categorical variables with fewer values, one-hot encoding is adopted; for unordered categorical variables with more values, embedding encoding or hashing encoding is adopted.

[0024] Perform data annotation on the standardized dataset and add risk-related labels, mainly including marking data records with similar characteristics according to the characteristic patterns of historical violation cases; adding compliance risk labels to data records according to regulatory requirements and compliance requirements; marking potential risk points according to expert experience and industry best practices. In the annotation process, a semi-automated method is adopted. First, preliminary annotation is carried out through a rule engine, and then it is reviewed and confirmed by compliance experts. For complex or ambiguous situations, multiple experts negotiate and decide. The annotation content includes information such as risk type, risk level, and risk basis. Through annotation, a labeled dataset is formed, and an association is established between data records and risk information.

[0025] Integrate and store the labeled dataset in a data warehouse, establish a unified data storage architecture, integrate data of different types and sources into the data warehouse, and form a consistent data view. Establish a data indexing mechanism, create indexes for key fields to improve data query efficiency; establish a data access control mechanism, set different levels of data access permissions to ensure data security and compliance; establish a data update mechanism to achieve regular and real-time updates of data and maintain data timeliness. Through these steps, a structured dataset is formed.

[0026] In a specific embodiment, the process of executing step S102 may specifically include the following steps: Combined with the data characteristics in the structured dataset, divide the financial compliance risks into anti-money laundering compliance risks, customer identity verification risks, internal control risks, market manipulation risks, conflict of interest risks, tax compliance risks, and cross-border business compliance risks to obtain a risk area classification framework; Subdivide each risk area in the risk area classification framework, establish risk subclasses under each risk area, and construct the association relationship between the risk area and the risk subclasses to obtain a secondary risk classification structure; Based on the secondary risk classification structure, design numerical indicators of transaction frequency, transaction amount, customer credit score, and regulatory compliance degree for each risk subcategory to form a three-level indicator system; Allocate weights to the indicators at all levels in the three-level indicator system, and determine the relative importance of each indicator in the overall assessment through the analytic hierarchy process to obtain a weighted indicator system; According to the weighted indicator system and historical data analysis, formulate a risk threshold range for each risk indicator, and divide the normal value range, attention value range, and warning value range to obtain an indicator system with thresholds; Establish a dynamic adjustment mechanism for the indicator system with thresholds. According to changes in regulatory policies and new risk events, update the indicator definitions and threshold settings to form a multi-level financial compliance risk assessment indicator system.

[0027] Specifically, divide the risk areas in combination with the data characteristics in the structured dataset. By analyzing the characteristic patterns of financial transaction data, customer information data, regulatory requirement data, and market data in the structured dataset, identify the main types of financial compliance risks, and divide financial compliance risks into seven core areas: anti-money laundering compliance risk, customer identity verification risk, internal control risk, market manipulation risk, conflict of interest risk, tax compliance risk, and cross-border business compliance risk. Anti-money laundering compliance risk refers to the compliance risk faced by financial institutions when they fail to effectively identify, monitor, and report suspicious transactions or fail to fulfill customer due diligence obligations; customer identity verification risk refers to the risk of failing to accurately identify and verify customer identity information; internal control risk involves risks caused by defects in internal management mechanisms and operation processes; market manipulation risk refers to the risk of influencing financial market prices through improper trading behaviors; conflict of interest risk refers to the risk arising from the inconsistency between the interests of financial institutions or practitioners and the interests of customers; tax compliance risk refers to the risk of failing to comply with tax regulations; cross-border business compliance risk refers to the risk of failing to meet the regulatory requirements of different countries or regions in cross-border financial activities. This division is based on the cluster analysis of risk characteristics in the structured dataset and the sorting of regulatory requirements to form a risk area classification framework.

[0028] Each risk area in the risk area classification framework is further divided to establish risk sub - categories under each risk area, and the association relationship between the risk area and the risk sub - categories is constructed. Taking the anti - money laundering compliance risk as an example, it is divided into sub - categories such as customer due diligence risk, unusual transaction monitoring risk, suspicious transaction reporting risk, and anti - terrorism financing risk; customer identity verification risk is divided into sub - categories such as authenticity risk of identity information, integrity risk of identity information, and identity information update risk; internal control risk is divided into sub - categories such as job responsibility risk, approval process risk, and internal supervision risk. Each risk area is usually divided into 3 - 5 risk sub - categories. By defining the characteristic attributes and boundary conditions of the risk sub - categories, the differences between the sub - categories are clarified. At the same time, the mapping relationship between the risk area and the risk sub - categories is established to form a two - level risk classification structure. This step is completed through the characteristic analysis of risk events in the structured dataset, combined with expert experience and regulatory classification standards.

[0029] Based on the two - level risk classification structure, specific risk indicators are designed for each risk sub - category. The transaction frequency indicator measures the number of transactions within a unit of time to identify unusual transaction behaviors; the transaction amount indicator evaluates risks by analyzing the size and change trend of the transaction amount. Large - value transactions or sudden changes in the amount usually pose higher risks; the customer credit score indicator synthesizes factors such as the customer's credit history, repayment ability, and default records to provide a basis for customer risk assessment; the regulatory compliance degree indicator measures the degree to which a transaction or behavior complies with regulatory requirements and is scored by referring to regulatory regulations. For each risk sub - category, a suitable combination of indicators is selected according to its characteristics, and the calculation method, data source, and update frequency of the indicators are clarified to form a complete three - level indicator system. This step is achieved through the statistical analysis and feature extraction of relevant data in the structured dataset.

[0030] Weight distribution is carried out for each level of indicators in the three - level indicator system. The analytic hierarchy process (AHP) is used to determine the relative importance of each indicator in the overall assessment. The AHP is a multi - criterion decision - making method that decomposes complex problems into a hierarchical structure and determines the relative importance of each factor through pairwise comparisons. The specific steps include: establishing a hierarchical structure model, dividing the risk assessment indicators into a goal layer, a criterion layer, and an indicator layer; constructing a judgment matrix, making pairwise comparisons between each pair of elements in each level to form a judgment matrix; calculating the weight vector, using the eigenvalue method to calculate the eigenvector of the judgment matrix as the relative weight of the elements in this level to the previous level; consistency check, verifying the consistency of the judgment matrix to ensure the rationality of the weight distribution. Through the AHP, the weights of the risk area layer, the risk sub - category layer, and the risk indicator layer are determined to form a complete weighted indicator system.

[0031] Based on the weighted index system and historical data analysis, risk threshold intervals are formulated for each risk indicator. First, collect historical risk event data and analyze the performance characteristics of each risk indicator in risk events; through statistical analysis methods such as quantile analysis and cluster analysis, determine the distribution law of indicator values; based on the data distribution and risk level, divide the indicator values into three intervals: normal value range, attention value range, and warning value range. The normal value range means that the risk indicator value is at a safe level and does not require special attention; the attention value range means that the risk indicator value is abnormal and needs to be closely monitored; the warning value range means that the risk indicator value is significantly abnormal and there is a high risk, which requires immediate handling. By setting these threshold intervals, an indicator system with thresholds is formed, providing a judgment standard for risk monitoring and early warning.

[0032] Finally, establish a dynamic adjustment mechanism for the indicator system with thresholds, enabling the indicator system to adapt to changes in the financial environment and regulatory requirements. The dynamic adjustment mechanism includes a regular evaluation mechanism to regularly evaluate the effectiveness of the indicator system and analyze the contribution of indicators to risk identification; an event trigger mechanism to trigger indicator updates when new regulatory policies or major risk events occur; and a feedback optimization mechanism to continuously optimize indicator definitions and threshold settings based on feedback from actual usage effects. Through these mechanisms, the indicator system is updated in a timely manner to ensure its continuous effectiveness. The formed multi-level financial compliance risk assessment indicator system provides a framework and standard for the identification, assessment, and early warning of financial compliance risks.

[0033] For example, a certain bank applies this method to construct an anti-money laundering compliance risk assessment indicator system. First, according to transaction data and regulatory requirements, divide the anti-money laundering risk into three sub-categories: customer due diligence risk, transaction monitoring risk, and reporting compliance risk. Then design specific indicators: under customer due diligence risk, set indicators such as the proportion of high-risk customers and the integrity rate of identity information; under transaction monitoring risk, set indicators such as the frequency of large-value transactions and the amount of cross-border transactions; under reporting compliance risk, set indicators such as the timeliness of suspicious transaction reports and the reporting quality score. Determine the weights through the analytic hierarchy process. For example, in the overall anti-money laundering risk, the weight of customer due diligence risk is 0.4, the weight of transaction monitoring risk is 0.4, and the weight of reporting compliance risk is 0.2. Set thresholds for each indicator. For example, in the indicator of the frequency of large-value transactions, less than 5 times per month is the normal range, 5 - 10 times is the attention range, and more than 10 times is the warning range. When the regulatory agency issues new anti-money laundering regulations, immediately update the indicator definitions and threshold settings. Through this multi-level indicator system, the bank can comprehensively assess and monitor anti-money laundering compliance risks, effectively identify potential risks, and take corresponding measures in a timely manner.

[0034] In a specific embodiment, the process of executing step S103 may specifically include the following steps: According to the corresponding relationships among the risk areas, risk sub - categories, and risk indicators in the multi - level financial compliance risk assessment index system, extract training features from the structured dataset, and divide the training set, validation set, and test set in a ratio of 7:2:1 to obtain the model training dataset; Use the risk indicator values in the model training dataset as input features, and use the risk levels corresponding to the risk indicators as target labels to train a decision - tree model, a five - layer feed - forward neural network model, and a radial - basis kernel support vector machine model respectively to obtain three basic models; Apply the three basic models to the validation set, calculate the F1 scores of each model for each risk area in the multi - level financial compliance risk assessment index system, and form a performance matrix containing the mapping relationship between the risk area and the model performance; Based on the performance matrix, construct a cascaded fusion architecture. For each risk area, select the model with the best performance as the main model, and the remaining models as auxiliary models. Combine the output results of the three basic models through weighted voting to obtain the initial fusion network; Perform grid - search parameter optimization on the initial fusion network on the validation set, adjust the weight coefficients and decision thresholds of each basic model to obtain the fusion network with optimized parameters; Apply the fusion network with optimized parameters to the test set, calculate the average precision and recall of each risk area, and form a financial compliance risk assessment model. The financial compliance risk assessment model is used to receive the structured dataset and output risk scores and risk levels.

[0035] Specifically, according to the corresponding relationships among the risk areas, risk sub - categories, and risk indicators in the multi - level financial compliance risk assessment index system, extract training features from the structured dataset. This process is actually to screen out the data fields related to risk assessment from the structured dataset according to the definition of the index system, and convert these fields into feature vectors that the model can process. During the data extraction process, for each risk indicator, determine its data source and calculation method, and then query the corresponding data from the structured dataset. For example, for the transaction frequency indicator, count the number of transactions within a unit time from the transaction data table; for the transaction amount indicator, calculate statistical quantities such as the mean, variance, and maximum value of the transaction amount; for the customer credit score indicator, extract the credit score from the customer information data or calculate it through a credit scoring model. The extracted feature data is divided into a training set, a validation set, and a test set in a ratio of 7:2:1. The training set accounts for 70% for model training, the validation set accounts for 20% for model selection and parameter tuning, and the test set accounts for 10% for the final model performance evaluation. The data division process uses the stratified sampling method to ensure that the risk level distribution in each subset is similar to that of the original dataset and avoid sample bias.

[0036] Using the risk indicator values in the model training dataset as input features and the risk levels corresponding to the risk indicators as target labels, three different types of machine learning models are trained respectively. The decision tree model is a classification model based on a tree structure. It divides the data space into different regions through recursive binary division, and each region corresponds to a prediction result. During the training process, information gain or Gini coefficient is selected as the division criterion to find the best feature and threshold for node splitting until the stopping condition is met. The advantage of the decision tree model is its strong interpretability, which can reflect the importance ranking of features and is suitable for dealing with problems with clear rules in financial risk assessment. The five-layer feedforward neural network model is a neural network composed of an input layer, three hidden layers, and an output layer. Each layer of neurons is fully connected to the next layer. The training process uses the backpropagation algorithm to minimize the loss function between the predicted value and the true label through the gradient descent method, and continuously adjusts the network weights and biases. The advantage of the neural network model is that it can capture the complex non-linear relationships between features and is suitable for dealing with hidden patterns in financial risk assessment. The radial basis kernel support vector machine model is a support vector machine that uses the radial basis function as the kernel function. By mapping the data into a high-dimensional space, it finds the decision boundary with the maximum margin. During the training process, the support vectors and model parameters are found by solving a quadratic programming problem to construct a classification model. The advantage of the support vector machine model is its good performance in high-dimensional spaces and strong anti-overfitting ability, which is suitable for dealing with high-dimensional feature data in financial risk assessment.

[0037] Apply the three basic models to the validation set, and calculate the F1 scores of each model for each risk area in the multi-level financial compliance risk assessment index system. The F1 score is the harmonic mean of precision and recall, and the calculation formula is F1 = 2 × (precision × recall) / (precision + recall), where precision is the proportion of the actual positive class among the positive classes predicted by the model, and recall is the proportion of the positive classes predicted by the model among the actual positive classes. The F1 score comprehensively considers precision and recall and is an important indicator for evaluating the performance of classification models, especially suitable for evaluating financial risk prediction models, because in risk assessment, both missed reports (low recall) and false reports (low precision) will have negative impacts. For each risk area (such as anti-money laundering compliance risk, customer identity verification risk, etc.), calculate the F1 scores of the three basic models respectively to form a performance matrix containing the mapping relationship between risk areas and model performance. The rows of the performance matrix represent different risk areas, the columns represent different models, and the matrix elements are the F1 scores of the corresponding models in the corresponding risk areas.

[0038] Construct a cascaded fusion architecture based on the performance matrix. For each risk area, select the model with the best performance as the main model, and the remaining models as auxiliary models. Cascaded fusion is an ensemble learning method that generates more accurate and stable predictions by combining the prediction results of multiple base models. In each risk area, select the model with the highest F1 score as the main model according to the performance matrix and assign it a higher weight; the other two models are used as auxiliary models and assigned lower weights. Combine the output results of the three base models through weighted voting to form an initial fusion network. Specifically, for each prediction sample, the three base models each output the prediction result and its confidence, and then perform a weighted sum according to the preset weights to obtain the final prediction result. The initial fusion network uses a linear weighting method, that is, prediction result = w1 × result of model 1 + w2 × result of model 2 + w3 × result of model 3, where w1, w2, and w3 are the corresponding weight coefficients, and the initial values are set according to the F1 scores of the models in the corresponding risk areas.

[0039] Perform grid search parameter optimization on the initial fusion network on the validation set, and adjust the weight coefficients and decision thresholds of each base model. Grid search is a method for hyperparameter optimization that finds the optimal parameter settings by systematically trying different combinations in the parameter space. The search range of the weight coefficients is set to the interval [0, 1], and the search range of the decision threshold is set to the interval [0.3, 0.7]. The step size is determined according to the computing resources and time requirements. For each set of parameter combinations, calculate the performance metrics (such as F1 score) of the fusion network on the validation set, and select the parameter combination with the best performance as the final parameters. By adjusting the weight coefficients, change the contribution degree of different models to the final prediction; by adjusting the decision threshold, balance the trade-off between precision and recall, so that the fusion network achieves the best performance in a specific risk area. The parameter optimization process is iterated until the parameter combination with the best performance is found to form a parameter-optimized fusion network.

[0040] Apply the parameter-optimized fusion network to the test set, calculate the average precision and recall of each risk area, and evaluate the final performance of the model. The average precision is the average of the precisions at different thresholds, which can more comprehensively evaluate the model performance; the recall reflects the ability of the model to discover all real risk cases. Through the performance evaluation on the test set, verify the generalization ability of the fusion network to ensure that the model can effectively identify financial compliance risks in actual applications. The finally formed financial compliance risk assessment model includes specialized prediction modules for different risk areas, can receive structured data sets as input, and output risk scores and risk levels, providing decision support for the compliance risk management of financial institutions.

[0041] Taking the compliance risk assessment of a financial institution as an example, the institution extracted 20,000 records containing various risk indicators from historical data, and each record was associated with a risk level label verified afterwards. The data features included 50 indicators such as customer transaction frequency, transaction amount, account activity, and customer credit score, and the target label was the risk level (low, medium, and high levels). The data was divided in a ratio of 7:2:1 to obtain 14,000 training data, 4,000 validation data, and 2,000 test data. When training the decision tree model, the Gini coefficient was selected as the splitting criterion, the maximum depth was set to 5, and the minimum number of samples for splitting was 20; when training the five-layer feedforward neural network, the number of neurons in the three hidden layers was set to 128, 64, and 32 respectively, the ReLU activation function was used, the Adam optimizer was adopted, and the learning rate was 0.001; when training the radial basis kernel support vector machine, the kernel parameter γ was set to 0.1 and the regularization parameter C was set to 10. The performances of the three models were evaluated on the validation set, and it was found that in the anti-money laundering risk field, the F1 score of the neural network was 0.83, that of the decision tree was 0.76, and that of the support vector machine was 0.79; while in the customer identity verification risk field, the decision tree performed best with an F1 score of 0.85. Based on these results, a cascaded fusion network was constructed. For the anti-money laundering risk field, the weight of the neural network was set to 0.6, and those of the other two models were each 0.2; for the customer identity verification risk field, the weight of the decision tree was set to 0.6, and the others were adjusted accordingly. By optimizing the weights and thresholds through grid search, the average precision of the final fusion model on the test set reached 0.87, and the recall rate reached 0.84, significantly superior to the single models, proving the effectiveness of the cascaded fusion architecture in financial compliance risk assessment.

[0042] In a specific embodiment, the process of executing step S104 may specifically include the following steps: Group the structured data set according to the customer dimension, transaction dimension, business dimension, and institutional dimension, and extract the features corresponding to the risk indicators according to the indicator definitions of the multi-level financial compliance risk assessment index system to obtain a multi-dimensional feature matrix; Perform data standardization and missing value processing on the multi-dimensional feature matrix to make it meet the input requirements of the financial compliance risk assessment model, and obtain preprocessed feature data; Input the preprocessed feature data into the decision tree model, five-layer feedforward neural network model, and radial basis kernel support vector machine model in the financial compliance risk assessment model respectively to obtain the risk assessment values of the three groups of models; According to the model performance mapping relationship of each risk field in the performance matrix, perform weighted fusion calculation on the risk assessment values of the three groups of models, and use the parameter-optimized fusion network to generate the risk scores of each risk indicator to obtain the comprehensive risk assessment result; According to the comprehensive risk assessment results, based on the hierarchical relationship of the multi-level financial compliance risk assessment index system, the scores of the risk indicator levels are weighted and aggregated to obtain the risk scores of the risk sub-category level and the risk area level in sequence, forming a hierarchical risk score table; Compare the hierarchical risk score table with the risk level classification criteria to generate a risk level distribution map and a risk trend analysis map, and combine them into a compliance risk assessment report.

[0043] Specifically, the structured data set is grouped and feature-extracted according to different dimensions. The structured data set is a unified format data set formed after previous data collection, cleaning, and standardization processing, and contains content such as financial transaction records, customer information, and regulatory data. The data grouping process divides the data according to the customer dimension, transaction dimension, business dimension, and institution dimension. The customer dimension grouping focuses on individual customer characteristics, including customer basic information, behavior patterns, credit status, etc.; the transaction dimension grouping focuses on transaction-level characteristics, including transaction frequency, amount, time distribution, etc.; the business dimension grouping focuses on the characteristics of different business lines, such as retail business, corporate business, investment business, etc.; the institution dimension grouping focuses on the compliance situation at the institution level, such as the compliance performance of branches and the compliance status of departments. Based on the grouping, corresponding features are extracted from the data of each dimension according to the index definitions of the multi-level financial compliance risk assessment index system. For example, features such as customer risk level and transaction activity are extracted from the customer dimension; features such as transaction amount outliers and changes in transaction frequency are extracted from the transaction dimension; features such as business compliance and business risk concentration are extracted from the business dimension; features such as institution risk exposure level and compliance management effectiveness are extracted from the institution dimension. In this way, a multi-dimensional feature matrix is formed, where the rows of the matrix represent the evaluation objects (such as customers, transactions, business lines, etc.), and the columns represent different risk indicator features.

[0044] Perform data standardization and missing value handling on the multi-dimensional feature matrix to make the data meet the model input requirements. Data standardization is the process of converting features with different dimensions and distributions into a unified scale, mainly including two methods: min-max standardization and Z-Score standardization. Min-max standardization maps feature values to the interval [0, 1], which is suitable for features with clear upper and lower bounds in the distribution; Z-Score standardization converts features into a distribution with a mean of 0 and a standard deviation of 1, which is suitable for features with an approximately normal distribution. Select the appropriate standardization method according to different feature types. For example, perform Z-Score standardization on the transaction amount after logarithmic transformation, and use min-max standardization for the risk score index. Missing value handling mainly adopts methods such as mean filling, median filling, mode filling, or model prediction filling. For randomly missing numerical features, use mean or median filling; for missing categorical features, use mode filling; for missing values strongly correlated with other features, use model prediction based on other features for filling. Through standardization and missing value handling, complete, unified, and standardized preprocessed feature data is obtained to ensure that the data quality and format meet the requirements of subsequent model analysis. Input the preprocessed feature data into three basic models in the financial compliance risk assessment model respectively to obtain risk assessment values. The decision tree model classifies risks by constructing a tree structure, divides the feature space into different regions, and each leaf node corresponds to a risk prediction result. The feature data is screened layer by layer through the internal nodes of the decision tree and finally reaches the leaf node to obtain the risk assessment value, including the risk classification result and the confidence level. The five-layer feedforward neural network model includes an input layer, three hidden layers, and an output layer, and information is transmitted and transformed through activation functions and weights. The feature data is first normalized and then input into the network, and then non-linear transformations are performed in each hidden layer, and finally a risk assessment value is generated in the output layer. The radial basis kernel support vector machine model maps features to a high-dimensional space through a kernel function and constructs an optimal separating hyperplane for classification. After the feature data is transformed by the kernel function, the distance between it and the support vector is calculated, and the risk assessment value is generated accordingly. These three models independently process the feature data and obtain three different sets of risk assessment results, including the risk classification results and the corresponding probability values or confidence level scores.

[0045] According to the model performance mapping relationships of each risk area in the performance matrix, weighted fusion calculations are performed on the risk assessment values of the three groups of models. The performance matrix is formed during the model training stage and contains the performance indicators (such as F1 score) of each model in various risk areas. During the fusion calculation process, the weights of each model in different risk areas are determined according to the performance matrix, and the better-performing models are given higher weights. For each risk indicator, the corresponding evaluation values of the three models are extracted, and weighted summation is performed according to the preset weights to obtain the fusion evaluation result. For example, in the anti-money laundering risk area, if the neural network model performs best, a higher weight is given to it in the fusion calculation; in the customer identification risk area, if the decision tree model performs best, its weight is correspondingly increased. Through the fusion network with parameter optimization, the outputs of the three basic models are intelligently combined, integrating the advantages of each model, generating the risk scores of each risk indicator, and forming a comprehensive risk assessment result.

[0046] According to the comprehensive risk assessment result, weighted aggregation calculations are performed on the scores of the risk indicator levels according to the hierarchical relationship of the multi-level financial compliance risk assessment index system. The multi-level financial compliance risk assessment index system includes three levels: risk area, risk subcategory, and risk indicator. Each upper-level node consists of multiple lower-level nodes, and each node has a corresponding weight. The weighted aggregation calculation process is a bottom-up summary process. First, the risk scores of the bottom-level risk indicators are obtained, and then the risk ratings of the risk subcategories are calculated according to the weight settings in the index system. The specific calculation method is to multiply the scores of each risk indicator under the risk subcategory by the corresponding weights and then sum them to obtain the risk subcategory score. Similarly, multiply the scores of each risk subcategory under the risk area by the corresponding weights and then sum them to obtain the risk area score. Through this way of layer-by-layer aggregation, the ratings of each risk area and the overall compliance risk are finally obtained, forming a hierarchical risk rating table, which clearly shows the rating situation and its logical relationship from specific risk indicators to the overall risk.

[0047] Compare the hierarchical risk scoring table with the risk level classification criteria to generate a risk level distribution map and a risk trend analysis map, and combine them into a compliance risk assessment report. The risk level classification criteria are formulated based on historical data analysis and regulatory requirements. Generally, risks are divided into several levels such as low risk, medium risk, and high risk, and each level corresponds to a specific scoring range. Compare the scoring values in the hierarchical risk scoring table with the classification criteria to determine the risk levels of each risk indicator, risk subclass, and risk area. Based on the risk level information, generate a risk level distribution map to visually display the risk level distribution of different risk areas and different risk subclasses; generate a risk trend analysis map to show the change trend of risk scores over time, which is convenient for identifying risk change patterns. Finally, combine the risk scoring table, risk level distribution map, risk trend analysis map, and relevant explanations into a complete compliance risk assessment report to provide comprehensive risk assessment information and decision-making support for financial institutions.

[0048] For example, extract data from structured datasets for grouping and feature extraction. From the customer dimension, extract basic information such as customer age, occupation, and income, as well as behavioral characteristics such as account activity and transaction preferences; from the transaction dimension, extract features such as transaction frequency, amount, and geographical distribution; from the business dimension, extract features such as the transaction volume of each business line and compliance monitoring indicators; from the institutional dimension, extract features such as the risk control measures of each branch and the allocation of compliance personnel. After combining these features into a multi-dimensional feature matrix, perform Z-Score standardization on numerical data such as transaction amount, one-hot encoding on categorical data such as transaction type, and fill in the missing transaction IP address information with the IP of the most recent transaction to form preprocessed feature data. Input the feature data into three models respectively: The decision tree model classifies customers by judging features such as customer transaction frequency and single transaction amount; the neural network model captures the complex relationships of customer transaction patterns through deep learning; the support vector machine model deals with the non-linear relationships between high-dimensional features. According to the performance matrix, the neural network performs best in transaction fraud detection (F1 = 0.86). When calculating in this field, the weight of the neural network is set to 0.6, and 0.2 for other models; the decision tree is the best in customer identity risk assessment, and the weights are adjusted accordingly. Calculate the scores of each customer on each risk indicator through weighted fusion, and then aggregate them according to the hierarchical relationship of the indicator system. For example, a certain customer scores 0.75 on the transaction frequency anomaly indicator and 0.3 on the transaction geographical anomaly indicator, and the score of the transaction monitoring risk subclass is 0.6 after aggregation according to the weight; the scores of this customer in the three risk subclasses are aggregated by weight to obtain the total anti-money laundering risk area score of 0.55. According to the risk classification criteria (above 0.7 is high risk, 0.4 - 0.7 is medium risk, and below 0.4 is low risk), this customer is rated as a medium risk level. The generated risk assessment report includes the risk scoring table, risk level distribution, and historical trend chart of this customer.

[0049] In a specific embodiment, the process of executing step S105 may specifically include the following steps: Parse the compliance risk assessment report, extract the risk indicators whose risk scores exceed the preset threshold, including the risk indicator name, risk score, the risk subcategory to which it belongs, and the risk area, to obtain a high-risk indicator set; Analyze the risk characteristic patterns based on the high-risk indicator set, identify the abnormal transaction behaviors and compliance hidden dangers in the business process, and obtain a preliminary risk point list; Trace the data of each risk point in the preliminary risk point list, extract the associated original transaction records, customer information, and relevant business data, form a risk point evidence chain, and obtain a risk point file; According to the three core factors of risk severity, impact scope, and occurrence frequency in the risk point file, assign a weight coefficient to each risk point to obtain a risk point weight table; Use the risk point weight table and the risk indicator scores in the compliance risk assessment report to perform a composite calculation to obtain the quantitative evaluation value of each risk point, and comprehensively consider the direct risk and associated risk to obtain the risk point quantitative value; According to the risk point quantitative value, divide the risk points into four risk levels: critical level, high level, medium level, and low level, and sort the risk points according to the risk level to form a risk point quantitative evaluation result.

[0050] Specifically, perform parsing processing on the compliance risk assessment report. The compliance risk assessment report is a comprehensive document generated by the fusion model, including the scores of each risk indicator, the risk distribution, and trend analysis. The parsing process uses a structured text analysis method to extract the risk indicator information whose risk scores exceed the preset threshold from the report. The preset threshold is a risk warning line set according to historical risk data analysis and regulatory requirements, and usually different threshold values are set for different risk types. The parsing program compares the score of each risk indicator in the report with the corresponding threshold one by one. When the indicator score exceeds the threshold, extract the complete information of the indicator, including the risk indicator name (such as "large transaction frequency", "cross-border fund flow"), the actual risk score, the risk subcategory to which it belongs (such as "transaction monitoring risk"), and the risk area (such as "anti-money laundering compliance risk"). These high-risk indicators are aggregated into a high-risk indicator set, which serves as the basis for subsequent risk point identification.

[0051] Analyze the risk characteristic patterns based on the set of high-risk indicators to identify abnormal transaction behaviors and compliance risks in business processes. The risk characteristic pattern analysis uses pattern recognition and anomaly detection techniques to match high-risk indicators with a predefined risk pattern library to identify potential abnormal behavior patterns. The risk pattern library is a knowledge base constructed based on historical risk cases, regulatory requirements, and industry best practices, containing the characteristic descriptions of various known risks. The matching process not only considers the situation of individual indicator exceeding the standard, but also analyzes the combined relationship and time-series changes among multiple indicators to discover complex risk patterns. For example, when the indicators of "high frequency of large-value transactions by customers in a short period" and "high counterparty risk rating" both exceed the standard, it may indicate the existence of money laundering risk; when the combination of the indicators of "frequent cross-border fund flows" and "incomplete customer identity information" appears, it may indicate the existence of risks in illegal cross-border business. Through this characteristic pattern analysis, associate the phenomenon of risk indicator exceeding the standard with specific business scenarios and transaction behaviors to form a preliminary list of risk points, and each risk point includes a risk type description, a combination of risk trigger indicators, and a preliminary risk assessment. Conduct data tracing for each risk point in the preliminary risk point list to form a risk point evidence chain. Data tracing is a process of tracing back to the original data source, aiming to collect the original evidence to support the risk judgment. The tracing process first determines the data sources related to the risk point, mainly including transaction databases, customer information systems, business processing systems, etc. Then design query conditions to accurately extract the data records directly related to the risk point. For example, for the risk point of "large-value and frequent transactions", trace all the recent transaction records of this customer, including detailed information such as transaction time, amount, counterparty, transaction type, etc.; for the risk point of "abnormal customer identity", trace the customer's identity verification history, identity document information, address change records, etc. The tracing is not limited to directly related data, but also extends to associated data, such as counterparty information, relevant account activities, similar pattern cases, etc. Through the integration and correlation analysis of multi-source data, establish a complete risk evidence chain to form a detailed risk point file, providing data support and evidence basis for each risk point.

[0052] Based on the information in the risk point file, evaluate the risk severity from three core factors and assign weights. The risk severity refers to the degree of loss or negative impact that a risk event may cause, and is judged by evaluating factors such as potential financial losses, regulatory penalties, and reputation impacts, and is divided into four levels: extremely high, high, medium, and low. The scope of impact refers to the scope of business and the size of the customer group that a risk event may affect, and is judged by evaluating factors such as the number of business lines involved, the number of customers, and the scale of funds, and is also divided into four levels. The occurrence frequency refers to the number of times or frequency of risk behaviors occurring within a certain period, and is evaluated by counting historical occurrences or predicting future occurrence possibilities, and is also divided into four levels. These three core factors constitute a three-dimensional framework for risk assessment. After independent scoring for each dimension, they are combined to form the overall weight coefficient of the risk point. The weight assignment process takes into account the characteristic differences of different types of risks. For example, fraud risks usually pay more attention to severity, and operational risks pay more attention to occurrence frequency. Through this multi-dimensional assessment, an objective and reasonable weight coefficient is assigned to each risk point to form a risk point weight table.

[0053] Using the risk point weight table and the risk indicator scores in the compliance risk assessment report, perform a composite calculation to obtain the quantitative assessment value of the risk point. The composite calculation is a scoring method that comprehensively considers multiple factors, not only considering the indicator scores directly related to the risk point, but also considering indirectly related risk factors. The direct risk refers to the risk degree of the risk point itself, which is obtained by multiplying the indicator score directly related to the risk point by the weight coefficient; the associated risk refers to the additional risk brought by other risk points or risk factors associated with this risk point, which is obtained through the analysis of the risk association map. In the composite calculation process, first calculate the direct risk value, that is, the weighted sum of the risk indicator scores associated with the risk point and the corresponding indicator weights; then analyze the association relationship between risk points, identify the risk transmission path and the degree of impact, and calculate the associated risk value; finally, combine the direct risk value and the associated risk value according to a preset ratio to obtain the final quantitative assessment value of the risk point. This composite calculation method comprehensively considers the multi-dimensional attributes and associated impacts of risks, making the risk assessment results more comprehensive and accurate.

[0054] According to the quantified values of risk points, the risk points are divided into four risk levels and sorted. The risk level division is a risk classification based on the magnitude of the risk quantification value, combined with the institution's risk tolerance and regulatory requirements. Critical-level risks are the risk points with the highest quantification values, indicating that the risks have become apparent or are about to occur, and immediate intervention measures are required; high-level risks indicate that the risks are in a high-incidence state and need to be prioritized; medium-level risks indicate that the risks are near the warning line and need to be monitored regularly; low-level risks indicate that the risks are within an acceptable range and only require routine management. The risk level division not only considers the absolute magnitude of the quantification value but also takes into account the relative position of the risk distribution. Usually, the quartile or natural breakpoint method is used to determine the level boundaries. After determining the risk levels, the risk points are sorted according to the risk levels and the magnitude of the quantification values to form a priority processing order, generating the final quantified risk assessment result of the risk points. This result clearly shows the risk levels, priorities, and interrelationships of each risk point, providing a decision-making basis for subsequent risk response and management.

[0055] In a specific embodiment, the process of executing step S106 may specifically include the following steps: Conduct statistical analysis on the quantified risk assessment result of the risk points to determine the boundary of the quantified score interval for different risk levels as the benchmark threshold for risk warning, obtaining a benchmark threshold table; According to the benchmark threshold table and historical risk handling records, set differentiated warning thresholds for different risk types and risk levels, establish a hierarchical warning threshold matrix, and obtain a warning threshold system; Use the warning threshold system to monitor the quantified risk assessment result of the risk points. When the risk quantification index exceeds the corresponding warning threshold, trigger the warning signal generation process to obtain the original warning trigger mark; Based on the original warning trigger mark and the risk point data in the quantified risk assessment result of the risk points, generate a warning content including risk point description, risk type, occurrence location, influence scope, specific numerical indicators, and development trend, obtaining a risk description report; Match and analyze the risk description report with the handling cases in the historical risk database, generate handling suggestions for different risk types and risk levels, and integrate them into a warning signal package; According to the risk level and business area information in the warning signal package, determine the warning receiving departments and personnel, and send the warning signal package to the corresponding risk management departments via email, text message, and system notification methods to complete the warning distribution.

[0056] Specifically, statistical analysis is performed on the quantitative assessment results of risk points to determine a reasonable warning threshold. The statistical analysis of the quantitative assessment results of risk points is to process the distribution characteristics of the quantitative values of risk points through data analysis techniques and find the natural separation points of the data as the boundaries of risk levels. The specific processing process includes collecting the quantitative values of all risk points to form a data set; analyzing the distribution characteristics of the data, including calculating statistical quantities such as mean, standard deviation, and quartiles; using the quantile method or clustering analysis method to determine the natural separation points of the data. The quantile method determines the boundaries according to the quantile positions of the data, such as using the 25%, 50%, and 75% quantiles as the separation points for low, medium, and high risks; the clustering analysis method is to cluster the risk quantitative values into several categories through clustering algorithms such as K-means, and the boundaries between the categories are the level boundaries. Through this statistical analysis, the quantitative score intervals corresponding to the four risk levels of critical, high, medium, and low are determined, forming a benchmark threshold table. The benchmark threshold table is a comparison table that maps risk levels to quantitative score intervals and serves as the basic basis for warning triggers.

[0057] Differentiated warning thresholds are set for different risk types and levels based on the benchmark threshold table and historical risk disposal records. The benchmark threshold table provides general warning criteria, but the importance and urgency of different risk types vary in different business scenarios, so differentiated warning thresholds need to be set. The process of differentiated setting first analyzes historical risk disposal records, collects historical cases of different risk types, including information such as the quantitative values, disposal processes, and disposal results when the risks occurred; calculates the sensitivity and tolerance of different risk types, and risk types with high sensitivity and low tolerance need to set more stringent warning thresholds; considers business importance and regulatory requirements, and sets more sensitive thresholds for key businesses and areas with high regulatory attention. Through this differentiated adjustment, a hierarchical warning threshold matrix is established. This matrix is a multi-dimensional structure, with the horizontal axis representing different risk types (such as anti-money laundering risk, fraud risk, etc.), the vertical axis representing different risk levels (critical, high, medium, low), and the matrix elements being the specific warning threshold values. The hierarchical warning threshold matrix and the benchmark threshold table together constitute a complete warning threshold system, providing accurate judgment criteria for risk monitoring. The warning threshold system is used to monitor the real-time quantification evaluation results of risk points. When it is detected that the risk quantification index exceeds the corresponding warning threshold, the warning signal generation process is triggered. The monitoring process is a continuous data comparison process, comparing the real-time calculated risk quantification index with the thresholds in the warning threshold system. The specific implementation includes setting the monitoring period, determining the monitoring frequencies of different risk types. High-risk types such as anti-money laundering compliance risks may require daily or even real-time monitoring, while low-risk types such as general operation risks may be monitored once a week or a month; performing regular batch monitoring, checking the quantitative values of all risk points according to the preset period; establishing a real-time trigger mechanism to immediately evaluate newly added risk points or points with obvious changes in risk values. When it is found that the risk quantification index exceeds the threshold set in the warning threshold system, an original warning trigger mark is generated, and the mark content includes information such as the trigger time, trigger risk point ID, trigger threshold, actual value, and degree of excess, providing basic data for subsequent warning content generation.

[0058] Generate detailed early warning content based on the risk point data in the original early warning trigger mark and the quantitative assessment result of risk points. The generation of early warning content is a process of converting the original early warning mark into a risk description report that can be understood by humans. The specific generation steps include extracting the basic information of risk points, obtaining the basic information such as the risk point name, the business to which it belongs, and the customers involved from the risk point file according to the risk point ID; compiling the risk point description, describing the nature, characteristics, and manifestation forms of the risk in clear and concise language; sorting out the risk type information, clarifying the categories and sub-categories to which the risk belongs; clarifying the location where the risk occurs, including spatial location information such as business lines, departments, and regions; evaluating the scope of risk impact, analyzing the business scope, customer groups, capital scale, etc. that the risk may affect; providing specific numerical indicators, including the indicator values triggering the early warning, historical change trends, and other quantitative information; analyzing the risk development trend, predicting the future change trend of the risk through time series analysis. These contents are integrated into a structured risk description report, which has a unified format, comprehensive content, and accurate description, providing a basis for subsequent disposal suggestions.

[0059] Match and analyze the risk description report with the disposal cases in the historical risk database to generate handling suggestions. The historical risk database is a risk disposal experience database accumulated by financial institutions, containing information such as the disposal methods and effect evaluations of various risk events in history. The matching and analysis process includes feature matching, extracting the key features of the current risk point, matching them with the risk events in the historical case database, and finding similar cases; case screening, screening out the successful cases with good disposal effects from the matched cases; experience extraction, extracting the disposal methods, key steps, precautions, and other experiences from the successful cases; suggestion generation, forming targeted handling suggestions based on the extracted experiences and combined with the specific situation of the current risk. The content of the handling suggestions includes suggestions for emergency measures, risk mitigation methods, responsible departments, time limit requirements, etc. The risk description report and the handling suggestions together form a complete early warning signal package, which is a structured document containing all the information that needs to be transmitted to risk managers.

[0060] According to the risk level and business area information in the early warning signal package, determine the early warning recipients and complete the early warning distribution. Early warning distribution is the process of accurately and timely transmitting the early warning signal package to the relevant responsible persons. The specific distribution process includes determining the recipients, determining the receiving departments and personnel of the early warning information according to the risk level, risk type and business attribution, and forming a recipient list; selecting the distribution channel, selecting a suitable notification channel according to the risk urgency and recipient characteristics, such as using instant messaging methods such as text messages and phone calls for high-urgency risks, and using e-mails or in-system messages for regular risks; customizing the distribution content, appropriately screening and formatting the content of the early warning signal package according to the roles and responsibilities of the recipients to ensure that the recipients obtain information relevant to their responsibilities; executing the distribution, sending the customized early warning content to the recipients through the selected channel; confirming the distribution, tracking the delivery status and reading status of the early warning information to ensure the effective transmission of the information. Through this precise distribution, ensure that the risk early warning information can be timely delivered to the departments and personnel responsible for handling, and promote the timely response and handling of risks.

[0061] Taking the credit business of a certain bank as an example, by analyzing the distribution of the historical risk point quantification values, the bank found that there were natural breakpoints at the three positions of 0.3, 0.5, and 0.7 in the data. Based on this, the risk level was divided into four levels, forming a benchmark threshold table: 0-0.3 is low-level risk, 0.3-0.5 is medium-level risk, 0.5-0.7 is high-level risk, and above 0.7 is critical-level risk. By analyzing the historical disposal records, it was found that once fraud risks occur, they cause greater harm and have a long disposal cycle, requiring earlier intervention, while the disposal of operational risks is relatively simple and has a higher tolerance. Based on this, a hierarchical early warning threshold matrix was established. For example, for fraud risks, the critical-level threshold was adjusted to 0.65 and the high-level threshold was adjusted to 0.45; for operational risks, the original threshold remained unchanged. During daily monitoring, the system detected that the risk quantification value of a corporate loan customer was 0.68, exceeding the critical-level threshold of 0.65 for fraud risks, and immediately generated an original early warning mark. Based on this mark, the system extracted data such as the customer's basic information, loan situation, and risk performance, and generated a risk description report, the content of which included the risk description of "the enterprise obtained a loan by providing false financial statements", the risk type of "credit fraud", the occurrence location of "Company Bank Department, X Branch", the impact scope of "involving 3 loans totaling 5 million yuan", the specific indicators of "inconsistent financial data with tax declarations and abnormal large amounts of funds transferred recently", and the trend analysis of "the risk shows a deteriorating trend". Matching this report with historical fraud cases, treatment suggestions such as "immediately freeze the account, organize on-site inspections, and initiate legal procedures" were extracted to form a complete early warning signal package. According to the risk level and business attribution, the early warning recipients were determined as the director of the Company Bank Department, the risk control manager of the branch, and the Risk Management Department of the head office. They were notified urgently by text message and the complete early warning signal package was sent by e-mail at the same time to ensure that the relevant personnel could quickly take actions to respond to the risks.

[0062] The above describes the method for financial compliance risk assessment based on big data in the embodiments of the present application. Next, the system for financial compliance risk assessment based on big data in the embodiments of the present application will be described. Please refer to Figure 2 , an embodiment of the system for financial compliance risk assessment based on big data in the embodiments of the present application includes: A collection module 201, configured to collect and standardize financial transaction data, customer information data, regulatory requirement data, and market data to obtain a structured data set; A construction module 202, configured to construct a multi-level financial compliance risk assessment index system based on the structured data set, and the index system is a three-level structure including a risk area, a risk subcategory, and a risk index; A training module 203, configured to train a cascaded fusion network including a decision tree, a neural network, and a support vector machine by using the structured data set and the multi-level financial compliance risk assessment index system to obtain a financial compliance risk assessment model; An analysis module 204, configured to input the structured data set into the financial compliance risk assessment model for multi-dimensional real-time analysis to obtain a compliance risk assessment report; A quantification module 205, configured to identify potential compliance risk points according to the compliance risk assessment report, and assign weights to the compliance risk points for quantitative assessment to obtain a quantitative assessment result of the risk points; A generation module 206, configured to set a hierarchical warning threshold based on the quantitative assessment result of the risk points. When the risk quantification index exceeds the preset threshold, generate a warning signal including a risk point description, a risk level, and a processing suggestion, and send the warning signal to the corresponding risk management department.

[0063] Through the collaborative cooperation of the above-mentioned various components, through the organic combination of technical features such as the collection and standardized processing of multi-source data, the construction of a multi-level risk assessment index system, the training and application of a cascaded fusion network, multi-dimensional real-time analysis, quantitative assessment of risk points, and hierarchical early warning, the comprehensive assessment, accurate identification, and timely early warning of financial compliance risks have been achieved, and remarkable technical effects have been obtained. First, through the comprehensive collection and standardized processing of financial transaction data, customer information data, regulatory requirement data, and market data, a unified structured data set has been established, solving the problems of single data source, inconsistent format, and uneven quality in traditional methods, and providing a comprehensive and accurate data basis for subsequent risk analysis; second, the construction of a multi-level financial compliance risk assessment index system has realized the hierarchical management of risk areas, risk subclasses, and risk indicators, making the risk assessment more systematic and structured, and effectively avoiding the defects of scattered indicators and weak correlation in traditional methods; third, using the structured data set and multi-level index system to train a cascaded fusion network including decision trees, neural networks, and support vector machines, giving full play to the advantages of different machine learning algorithms. The decision tree model is good at handling risk judgments with clear rules, the neural network model is good at capturing complex non-linear relationships, and the support vector machine model performs well in high-dimensional spaces. By integrating the advantages of each algorithm through the fusion architecture, the accuracy and stability of risk identification have been significantly improved; at the same time, the structured data set is input into the financial compliance risk assessment model for multi-dimensional real-time analysis, generating a comprehensive compliance risk assessment report, realizing the dynamic and continuous monitoring of risks, and solving the problem of lagging risk assessment in traditional methods; in addition, potential compliance risk points are identified according to the compliance risk assessment report and quantitatively evaluated. Through scientific weight assignment and calculation methods, qualitative risks are transformed into quantitative indicators, making the risk assessment more objective and accurate; finally, based on the quantitative assessment results of risk points, hierarchical early warning thresholds are set, and early warning signals including risk point descriptions, risk levels, and handling suggestions are generated, realizing the early warning and accurate push of risks, and providing a time window and action guide for risk disposal. The technical solution of the present invention applies artificial intelligence algorithms and models in the specific field of financial compliance risk management, fully considering the contributions of different algorithm features to the solution, integrating the advantages of multiple algorithms through a cascaded fusion network, realizing the intelligent, accurate, and automated compliance risk assessment, effectively solving the problems of strong manual dependence, large subjectivity, and poor real-time performance in traditional methods, and providing a comprehensive, efficient, and accurate compliance risk management solution for financial institutions.

[0064] Above Figure 2 The financial compliance risk assessment system based on big data in the embodiments of the present invention is described in detail from the perspective of modular functional entities. Next, the financial compliance risk assessment device based on big data in the embodiments of the present invention is described in detail from the perspective of hardware processing.

[0065] Figure 3 FIG.

[0065] is a schematic structural diagram of a financial compliance risk assessment device based on big data provided by an embodiment of the present invention. The financial compliance risk assessment device 300 based on big data may vary greatly due to configuration or performance, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 for storing application programs 333 or data 332 (for example, one or more mass storage device ends). Among them, the memory 320 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the financial compliance risk assessment device 300 based on big data. Further, the processor 310 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the financial compliance risk assessment device 300 to implement the steps of the above-mentioned financial compliance risk assessment method based on big data.

[0066] The financial compliance risk assessment device 300 based on big data may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 3 the shown structural diagram of the financial compliance risk assessment device based on big data does not limit the financial compliance risk assessment device based on big data provided by the present invention, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0067] The present invention also provides a computer-readable storage medium. The computer-readable storage medium may be a non-volatile computer-readable storage medium, or may also be a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the financial compliance risk assessment method based on big data.

[0068] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, systems, and units may refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.

[0069] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a financial compliance risk assessment device based on big data (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0070] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A financial compliance risk assessment method based on big data, characterized in that: The method comprises: Collect and standardize financial transaction data, customer information data, regulatory requirements data, and market data to obtain structured data sets; Constructing a multi-level financial compliance risk assessment indicator system based on the structured data set, wherein the indicator system is a three-level structure including risk areas, risk subcategories and risk indicators; Using the structured data set and the multi-level financial compliance risk assessment indicator system to train a stacked fusion network including a decision tree, a neural network and a support vector machine to obtain a financial compliance risk assessment model; Inputting the structured data set into the financial compliance risk assessment model for multi-dimensional real-time analysis to obtain a compliance risk assessment report; Identify potential compliance risk points based on the compliance risk assessment report, and assign weights to the compliance risk points for quantitative assessment to obtain quantitative assessment results of the risk points; A graded warning threshold is set based on the quantitative assessment result of the risk point. When the risk quantification index exceeds the preset threshold, a warning signal including a risk point description, risk level, and handling suggestions is generated, and the warning signal is sent to the corresponding risk management department.

2. The financial compliance risk assessment method based on big data according to claim 1 is characterized in that: The financial transaction data, customer information data, regulatory requirements data and market data are collected and standardized to obtain a structured data set, including: Extract financial transaction data such as transaction records, fund flows, and account information from the internal systems of financial institutions, and obtain customer information data such as basic customer information, credit records, and transaction preferences from the customer management system to obtain the original data set; Performing data cleaning on the original data set to obtain a cleaned data set, and performing format conversion on the cleaned data set to obtain a data set with a unified format; Normalizing the numerical data in the unified format data set, and performing encoding conversion on the categorical data in the unified format data set to obtain a standardized data set; Annotating the standardized data set, adding risk-related labels, and obtaining a labeled data set; The labeled data set is integrated and stored in a data warehouse, and a data index and access control mechanism are established to form a structured data set.

3. The financial compliance risk assessment method based on big data according to claim 1 is characterized in that: The multi-level financial compliance risk assessment indicator system is constructed based on the structured data set. The indicator system is a three-level structure including risk areas, risk subcategories and risk indicators, including: Combined with the data features in the structured data set, the financial compliance risks are divided into anti-money laundering compliance risks, customer identity identification risks, internal control risks, market manipulation risks, conflict of interest risks, tax compliance risks, and cross-border business compliance risks, and a risk area classification framework is obtained; Subdividing each risk area in the risk area classification framework, establishing risk subcategories under each risk area, constructing a correlation between risk areas and risk subcategories, and obtaining a secondary risk classification structure; Based on the secondary risk classification structure, numerical indicators of transaction frequency, transaction amount, customer credit score, and regulatory compliance are designed for each risk sub-category to form a three-level indicator system; Assign weights to the indicators at each level in the three-level indicator system, determine the relative importance of each indicator in the overall evaluation through the analytic hierarchy process, and obtain a weighted indicator system; According to the weighted indicator system and historical data analysis, a risk threshold interval is formulated for each risk indicator, and a normal value range, a concern value range and a warning value range are divided to obtain an indicator system with thresholds; Establish a dynamic adjustment mechanism for the indicator system with thresholds, update indicator definitions and threshold settings according to changes in regulatory policies and new risk events, and form the multi-level financial compliance risk assessment indicator system.

4. The financial compliance risk assessment method based on big data according to claim 1 is characterized in that: The method uses the structured data set and the multi-level financial compliance risk assessment indicator system to train a stacked fusion network including a decision tree, a neural network and a support vector machine to obtain a financial compliance risk assessment model, including: Extracting training features from the structured data set according to the correspondence between risk areas, risk subcategories, and risk indicators in the multi-level financial compliance risk assessment indicator system, and dividing the training set, validation set, and test set into a ratio of 7:2:1 to obtain a model training data set; Using the risk indicator values ​​in the model training data set as input features and the risk levels corresponding to the risk indicators as target labels, a decision tree model, a five-layer feedforward neural network model, and a radial basis kernel support vector machine model are trained respectively to obtain three basic models; Applying the three basic models to the validation set, calculating the F1 score of each model for each risk area in the multi-level financial compliance risk assessment indicator system, and forming a performance matrix containing a mapping relationship between risk areas and model performance; A stacked fusion architecture is constructed based on the performance matrix. For each risk area, the model with the best performance is selected as the main model, and the remaining models are used as auxiliary models. The output results of the three basic models are combined by weighted voting to obtain an initial fusion network. Performing grid search parameter optimization on the initial fusion network on the validation set, adjusting the weight coefficient and decision threshold of each basic model, and obtaining a fusion network with optimized parameters; The parameter-optimized fusion network is applied to the test set, and the average precision and recall of each risk area are calculated to form the financial compliance risk assessment model, which is used to receive the structured data set and output a risk score and risk level.

5. The financial compliance risk assessment method based on big data according to claim 4 is characterized in that: The structured data set is input into the financial compliance risk assessment model for multi-dimensional real-time analysis to obtain a compliance risk assessment report, including: The structured data set is grouped according to the customer dimension, transaction dimension, business dimension and institution dimension, and according to the indicator definition of the multi-level financial compliance risk assessment indicator system, features corresponding to the risk indicators are extracted to obtain a multi-dimensional feature matrix; Performing data standardization and missing value processing on the multi-dimensional feature matrix to make it meet the input requirements of the financial compliance risk assessment model, and obtaining pre-processed feature data; The preprocessed feature data are respectively input into the decision tree model, the five-layer feedforward neural network model and the radial basis kernel support vector machine model in the financial compliance risk assessment model to obtain risk assessment values ​​of the three groups of models; According to the model performance mapping relationship of each risk field in the performance matrix, the risk assessment values ​​of the three groups of models are weighted fused and calculated, and the risk score of each risk indicator is generated by using the parameter-optimized fusion network to obtain a comprehensive risk assessment result; Based on the comprehensive risk assessment results, and in accordance with the hierarchical relationship of the multi-level financial compliance risk assessment indicator system, the scores of the risk indicator level are weighted and aggregated to obtain the risk scores of the risk sub-category level and the risk field level in turn, forming a hierarchical risk score table; The hierarchical risk scoring table is compared with the risk level classification standard to generate a risk level distribution chart and a risk trend analysis chart, which are combined into the compliance risk assessment report.

6. The financial compliance risk assessment method based on big data according to claim 1 is characterized in that: The identifying of potential compliance risk points according to the compliance risk assessment report and the quantitative assessment of the weights assigned to the compliance risk points to obtain the quantitative assessment results of the risk points include: Parsing the compliance risk assessment report, extracting risk indicators whose risk scores exceed a preset threshold, including risk indicator name, risk score, risk subcategory and risk field, to obtain a set of high-risk indicators; Analyze risk feature patterns based on the high-risk indicator set, identify abnormal transaction behaviors and compliance risks in business processes, and obtain a preliminary risk point list; Conduct data tracing for each risk point in the preliminary risk point list, extract associated original transaction records, customer information and related business data, form a risk point evidence chain, and obtain a risk point file; According to the three core factors of risk severity, impact scope and occurrence frequency in the risk point file, a weight coefficient is assigned to each risk point to obtain a risk point weight table; Using the risk point weight table and the risk indicator scores in the compliance risk assessment report, a composite calculation is performed to obtain a quantitative assessment value for each risk point, and direct risks and associated risks are comprehensively considered to obtain a quantitative value for the risk point; According to the quantitative value of the risk point, the risk point is divided into four risk levels: critical, high, medium and low, and the risk points are sorted according to the risk level to form the quantitative assessment result of the risk point.

7. The financial compliance risk assessment method based on big data according to claim 1 is characterized in that: The step of setting a graded warning threshold based on the quantitative assessment result of the risk point, generating a warning signal including a risk point description, risk level, and treatment suggestions when the risk quantitative index exceeds the preset threshold, and sending the warning signal to the corresponding risk management department includes: Performing statistical analysis on the quantitative assessment results of the risk points, determining the boundaries of the quantitative score intervals of different risk levels as the benchmark thresholds for risk warning, and obtaining a benchmark threshold table; According to the benchmark threshold table and historical risk disposal records, differentiated warning thresholds are set for different risk types and risk levels, a hierarchical warning threshold matrix is ​​established, and a warning threshold system is obtained; The risk point quantitative assessment results are monitored using the warning threshold system. When the risk quantitative index exceeds the corresponding warning threshold, the warning signal generation process is triggered to obtain the original warning trigger mark; Based on the original warning trigger mark and the risk point data in the risk point quantitative assessment result, a warning content including risk point description, risk type, occurrence location, impact scope, specific numerical indicators and development trend is generated to obtain a risk description report; Match and analyze the risk description report with the disposal cases in the historical risk database, generate disposal suggestions for different risk types and risk levels, and integrate them into an early warning signal package; According to the risk level and business field information in the early warning signal package, the early warning receiving department and personnel are determined, and the early warning signal package is sent to the corresponding risk management department via email, text message, or system notification to complete the early warning distribution.

8. A financial compliance risk assessment system based on big data, characterized in that: Used to implement the financial compliance risk assessment method based on big data as described in any one of claims 1 to 7, the financial compliance risk assessment system based on big data includes: The collection module is used to collect and standardize financial transaction data, customer information data, regulatory requirements data, and market data to obtain a structured data set; A construction module, used to construct a multi-level financial compliance risk assessment indicator system based on the structured data set, wherein the indicator system is a three-level structure including risk areas, risk subcategories and risk indicators; A training module, used to train a stacked fusion network including a decision tree, a neural network and a support vector machine using the structured data set and the multi-level financial compliance risk assessment indicator system to obtain a financial compliance risk assessment model; An analysis module, used to input the structured data set into the financial compliance risk assessment model for multi-dimensional real-time analysis to obtain a compliance risk assessment report; A quantification module is used to identify potential compliance risk points according to the compliance risk assessment report, and to perform quantitative assessment on the compliance risk points by assigning weights to obtain quantitative assessment results of the risk points; A generation module is used to set a graded warning threshold based on the quantitative assessment results of the risk points. When the risk quantification index exceeds the preset threshold, a warning signal including a risk point description, risk level, and processing suggestions is generated, and the warning signal is sent to the corresponding risk management department.

9. A financial compliance risk assessment device based on big data, characterized in that: It includes a memory and a processor, the memory stores a computer program that can be run on the processor, and the processor implements the big data-based financial compliance risk assessment method described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor is caused to execute the financial compliance risk assessment method based on big data as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Electrical operation risk assessment and early warning system based on big data analysis

    CN119539496A

Cited By

  • Online contract generation method and system

    CN120542405A

  • An online contract generation method and system

    CN120542405B

  • Risk data processing method and system based on multi-rule engine

    CN120654099A

  • Enterprise multi-dimensional data feature fusion method combined with AI technology

    CN120724400A

  • Purchase bid opening and evaluation risk labeling method based on supply chain cooperative processing

    CN120851032A