A financial pre-account opening risk assessment system based on intelligent identification
By employing multi-dimensional contextual feature acquisition, dynamic nearest neighbor search, weighted ensemble learning, knowledge graph updates, and crowdsourced knowledge collection, the problems of insufficient contextual awareness and lagging knowledge updates in traditional financial pre-account opening risk assessment have been solved, achieving efficient and accurate risk identification and continuous optimization.
Patent Information
- Application Number
- CN202511079742.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-08-04
AI Technical Summary
Traditional financial pre-account risk assessment methods are ill-equipped to deal with diverse and sophisticated fraud tactics, lack multi-dimensional contextual awareness, are outdated in terms of risk knowledge updates, and lack knowledge sharing among financial institutions, resulting in insufficient accuracy in risk identification and slow response speed.
It employs a multi-dimensional context feature acquisition module, a context analysis and processing module, a risk assessment module, a knowledge graph update module, and a crowdsourced knowledge collection module to achieve dynamic perception of multi-dimensional context features and rapid updating and sharing of risk knowledge. Through dynamic nearest neighbor search, weighted ensemble learning, knowledge graph construction, and crowdsourcing mechanisms, it improves the accuracy of risk identification and continuously optimizes the knowledge base.
It has improved the accuracy and speed of risk identification, reduced false positive and false negative rates, enabled the structured expression and rapid updating of risk knowledge, formed a continuously optimized risk assessment ecosystem, provided a clear path for risk interpretation, and enhanced the transparency and credibility of risk control decisions.
Smart Images

Figure CN120580070B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of financial risk control, more particularly, it relates to a financial pre-account opening risk assessment system based on intelligent identification. BACKGROUND
[0002] With the deepening of the digital transformation of financial services, online pre-account opening has become an important channel for financial institutions to acquire customers. However, with the popularization of online services, financial fraud behaviors have shown characteristics of diversification, intelligence and organization. Traditional risk assessment methods mainly rely on static rules and fixed models, which are difficult to cope with rapidly evolving fraud methods, especially the differentiated risk characteristics exhibited in different situations.
[0003] Currently, the financial pre-account opening risk assessment technology mainly adopts a combination of rule-based judgment systems and machine learning models. These methods usually analyze risks from the aspects of identity information, device information and behavior characteristics submitted by the applicant, but lack comprehensive consideration of multi-dimensional situational factors, resulting in insufficient risk identification accuracy in special time periods, specific regions or special network environments. At the same time, the existing technology has obvious lag in risk knowledge updating, making it difficult to quickly respond to the emergence of new fraud methods.
[0004] In addition, the risk knowledge sharing mechanism between financial institutions is not perfect, and the risk identification experience of each institution often forms an information island, which cannot fully utilize the risk identification experience scattered in different individuals. This knowledge fragmentation makes the entire financial system respond slowly to new fraud methods, increasing the systemic risk.
[0005] Therefore, a financial pre-account opening risk assessment method that can realize multi-dimensional situational awareness, rapid risk knowledge updating and crowdsourcing knowledge collection is needed. SUMMARY
[0006] The present application provides a financial pre-account opening risk assessment system based on intelligent identification, which solves the technical problems of insufficient situational awareness, lagging risk knowledge updating, risk knowledge island and insufficient risk assessment interpretability in related technologies.
[0007] The present application provides a financial pre-account opening risk assessment system based on intelligent identification, which includes:
[0008] A multi-dimensional situational feature acquisition module for acquiring multi-dimensional situational features in the pre-account opening process, including time dimension features, space dimension features, environment dimension features and event dimension features;
[0009] The context analysis processing module is configured to convert the multi-dimensional context features into a context feature vector, search for similar context cases in historical data by using a dynamic nearest neighbor search algorithm, and extract hierarchical risk features from the identified similar context cases;
[0010] The risk assessment module is configured to generate a risk score and a risk classification result by using a weighted ensemble learning method and integrating the hierarchical risk features;
[0011] The knowledge graph updating module is configured to extract structured risk knowledge from multi-source information based on the risk assessment result and related cases, integrate the extracted risk knowledge with an existing risk knowledge graph, perform quality assessment on the newly extracted risk knowledge, and dynamically update the risk knowledge graph based on the assessment result;
[0012] The crowdsourcing knowledge collection module is configured to integrate risk identification experience of front-line staff and experts through a crowdsourcing mechanism, apply the collected knowledge to the risk assessment process, and form a continuous optimization closed loop for risk assessment.
[0013] In a preferred embodiment, the time dimension features collected by the multi-dimensional context feature collection module include an application time point, a time period distribution, whether it is a holiday, a pre-post application time interval, and a seasonal activity cycle;
[0014] The spatial dimension features include a geographic location, an IP address, device positioning information, a historical activity area, and a cross-area activity frequency;
[0015] The environmental dimension features include a network state, device features, a system environment, a network access method, and a security software configuration;
[0016] The event dimension features include market fluctuations, regulatory policy changes, marketing activities, contemporaneous high-risk events, and internal process changes of financial institutions.
[0017] In a preferred embodiment, the process of converting the multi-dimensional context features into a context feature vector by the context analysis processing module specifically includes:
[0018] Adaptive normalization processing is performed on continuous features, and Z-score or Min-Max normalization methods are selected according to feature distribution characteristics;
[0019] Improved one-hot encoding is used to convert categorical features into binary vectors, and target encoding technology is used for high-base number features;
[0020] A pre-trained BERT model in the financial field is used to convert text features into context-aware dense vectors;
[0021] A feature fusion method based on an attention mechanism is used to dynamically adjust the weight coefficients of the features in each dimension according to the current context;
[0022] Feature dimension reduction is performed by the autoencoder to retain key information while reducing computational complexity.
[0023] In a preferred embodiment, the process of the context analysis processing module searching for similar context cases in historical data using a dynamic nearest neighbor search algorithm specifically includes:
[0024] A multi-level index structure is established for historical context feature vectors using a hierarchical local sensitive hashing algorithm;
[0025] The Euclidean distance, cosine similarity, Mahalanobis distance, and Jaccard similarity coefficient between the current context feature vector and the historical context feature vector are calculated;
[0026] An adaptive weighted fusion algorithm is used to generate a comprehensive similarity score, and the weight coefficients are dynamically adjusted according to the historical retrieval effect. A hierarchical screening strategy is adopted, which first quickly retrieves the candidate set, then accurately calculates the similarity, and selects the most similar historical context case;
[0027] The local sensitive hashing parameters, including the number of hash tables, the number of hash functions, and the bucket size, are dynamically adjusted based on an online learning mechanism to adapt to changes in data distribution.
[0028] In a preferred embodiment, during the process of the context analysis processing module extracting hierarchical risk features:
[0029] Account layer features include identity information consistency, account attribute features, historical credit records, biometric verification results, and certificate authenticity scores;
[0030] Transaction layer features include operation behavior features, input features, timing features, operation habit deviations, and abnormal interruption patterns;
[0031] Relationship layer features include associated account risks, social network features, group behavior patterns, device association graphs, and cross-institution risk correlation degrees.
[0032] In a preferred embodiment, the process of the knowledge graph updating module extracting structured risk knowledge from multiple sources of information specifically includes:
[0033] Risk-related information is obtained from internal business feedback data, risk event data, regulatory compliance information, external intelligence data, and industry shared risk libraries;
[0034] Text preprocessing is performed using domain-adaptive natural language processing techniques, including financial professional term identification, cleaning, segmentation, and stop word removal;
[0035] Named entity recognition is performed using a bidirectional long short-term memory network combined with a conditional random field model to identify key entities;
[0036] The relationship extraction is performed by using a contrast learning-based distant supervision multi-instance learning method to identify semantic relationships and causal relationships between entities.
[0037] The event extraction is performed by using a multi-head hierarchical attention mechanism to extract trigger words, participants, and spatio-temporal background information of risk events.
[0038] A multi-dimensional attribute system of risk knowledge is constructed, including risk level, occurrence probability, impact range, timeliness, and applicable customer groups.
[0039] In a preferred embodiment, the process of quality evaluation of newly extracted risk knowledge by the knowledge graph updating module specifically includes:
[0040] Knowledge consistency test is performed to check whether the new knowledge conflicts with the existing knowledge system logically by using ontology reasoning technology.
[0041] Knowledge importance evaluation is performed to evaluate the impact of new knowledge on risk assessment based on historical case backtesting and Monte Carlo simulation.
[0042] Knowledge timeliness evaluation is performed to establish a knowledge decay model to evaluate the effective period and time sensitivity of knowledge.
[0043] Knowledge source reliability evaluation is performed to construct a multi-dimensional trust scoring system to evaluate knowledge quality based on the historical accuracy, professional field authority, and information update frequency of the knowledge source.
[0044] Analytic Hierarchy Process is used to calculate the weight of each evaluation dimension to generate a comprehensive quality score, and a dynamic threshold is set to determine the knowledge adoption strategy.
[0045] In a preferred embodiment, the process of integrating frontline personnel and expert risk identification experience through crowdsourcing mechanism by the crowdsourcing knowledge collection module specifically includes:
[0046] Multi-modal knowledge contribution channels are provided, including structured forms, natural language interfaces, case labeling tools, rule editors, and visual knowledge construction tools.
[0047] A hierarchical incentive compatibility mechanism is constructed to give differentiated point rewards and professional certifications according to the quality, quantity, application effect, and innovation of the contributed knowledge.
[0048] A multi-dimensional evaluation framework of crowdsourcing knowledge quality is established, including knowledge accuracy, completeness, novelty, applicability, and complementarity with existing knowledge.
[0049] An evaluation method combining expert review and collective wisdom is adopted to form the quality evaluation results through Delphi method and weighted voting mechanism.
[0050] A knowledge integration algorithm is designed to semantically align and structurally fuse high-quality crowd-sourced knowledge with the existing risk knowledge graph and apply it to the risk assessment process.
[0051] In a preferred embodiment, the intelligent recognition-based financial pre-account opening risk assessment system further comprises a knowledge dissemination feedback closed-loop mechanism:
[0052] A knowledge application effect tracking system is established to monitor the performance of crowd-sourced knowledge in different business scenarios and customer groups in real time;
[0053] An A / B testing method is used to evaluate the impact of crowd-sourced knowledge on risk assessment accuracy, false positive rate, and false negative rate;
[0054] Based on application effect feedback, the knowledge structure, correlation strength, and application weight are optimized through reinforcement learning algorithm;
[0055] A multi-dimensional contributor reputation evaluation system is constructed, including historical contribution quality, knowledge verification rate, professional field coverage, innovation ability, and continuous contribution degree;
[0056] A knowledge co-creation community is established to promote cross-departmental and cross-professional risk knowledge exchange and collaborative innovation.
[0057] In a preferred embodiment, a computer-readable storage medium is used to store computer-readable instructions that can run an intelligent recognition-based financial pre-account opening risk assessment system when read by a computer.
[0058] The beneficial effects of the present application are:
[0059] The present application realizes precise identification and continuous optimization of financial pre-account opening risks through the synergistic effect of multi-dimensional context perception and crowd-sourced risk knowledge collection mechanism, and has the following specific technical effects:
[0060] The present application improves the accuracy of risk identification by collecting and analyzing multi-dimensional context features. The system can dynamically adjust the risk assessment strategy according to different time, place, environment, and other context conditions, effectively identifying abnormal behavior patterns in specific contexts. Especially in special situations such as holidays, special regions, or network activity peaks, the system shows better risk identification ability than traditional static models, significantly reducing false positive rate and false negative rate.
[0061] The risk knowledge graph constructed by the present application realizes the structured expression and rapid update of risk knowledge. By automatically extracting risk knowledge from multiple sources and performing quality evaluation, the system can respond to new financial risks in a timely manner, shorten the risk knowledge update cycle, and keep the risk assessment model always capable of identifying the latest fraud methods.
[0062] The crowdsourcing knowledge collection mechanism of the present application effectively solves the problem of risk knowledge island. By integrating the risk identification experience of front-line personnel and experts, the system establishes a continuously self-improving risk knowledge base. This mechanism not only makes full use of valuable experience scattered in different individuals, but also continuously optimizes risk identification rules through knowledge dissemination feedback loop, forming a continuous optimization ecology of risk assessment.
[0063] The evaluation method based on the risk knowledge graph of the present application provides a clear risk explanation path. The system can trace the decision basis of the risk assessment result, generate an interpretable risk report, effectively support risk control decision and regulatory review, reduce the difficulty of dispute handling, and improve the transparency and credibility of risk management. BRIEF DESCRIPTION OF DRAWINGS
[0064] Figure 1 is a module diagram of a financial pre-account opening risk assessment system based on intelligent identification of the present application;
[0065] Figure 2 is a bar chart of risk assessment accuracy rate comparison of the present application;
[0066] Figure 3 is a line chart of risk knowledge update timeliness of the present application;
[0067] Figure 4 is a risk assessment model performance radar chart of the present application;
[0068] Figure 5 is a scatter plot of context similarity and risk correlation of the present application;
[0069] Figure 6 is a pie chart of crowdsourcing knowledge contribution quality distribution of the present application. DETAILED DESCRIPTION
[0070] The subject matter described herein will now be discussed with reference to example implementations. It should be understood that discussions of these implementations are merely provided to give a more fuller understanding of the subject matter described herein and are not intended to limit the scope of the protection of the present specification. The functions and arrangements of the elements discussed can be changed, omitted, substituted, or added to, as needed, without departing from the scope of the protection of the present specification. In addition, features described in some examples can be combined in other examples.
[0071] In at least one embodiment of the present application, a financial pre-account opening risk assessment system based on intelligent identification is disclosed, as shown in Figure 1 The steps include:
[0072] A multi-dimensional context feature collection module is configured to collect multi-dimensional context features in the pre-account opening process, including time dimension features, space dimension features, environment dimension features, and event dimension features.
[0073] The method specifically comprises the following steps:
[0074] Step 1.1, context feature collection and preprocessing;
[0075] In the pre-account opening process, the system collects multi-dimensional context features, including:
[0076] Time dimension features: including application time point, time period distribution, whether it is a holiday, and previous and subsequent application time interval, etc. time-related features, to capture the regularity and abnormality of the application time.
[0077] Space dimension features: including geographical location, IP address, device positioning information, historical activity area, etc. space-related features, to identify the risk signals of geographical location abnormality or conflict.
[0078] Environment dimension features: including network status (such as Wi-Fi, mobile network), device features (such as terminal type, device fingerprint), system environment (such as operating system, browser), etc. features, to identify the risk of device or environment abnormality.
[0079] Event dimension features: including system internal market fluctuations, regulatory policy changes, marketing activities, and external context features such as contemporaneous high-risk events, to correlate the influence of external events on risk.
[0080] For the collected raw context data, the system performs outlier detection, missing value processing, and feature standardization to ensure data quality and consistency.
[0081] In some embodiments, the system adopts a distributed context feature collection architecture to simultaneously collect features at different levels:
[0082] Front-end collection: through SDK or embedded scripts, device features, behavior features, etc. are collected directly on the user device side, and preliminary encryption is performed.
[0083] Gateway collection: IP features, network features, request patterns, etc. are collected at the gateway layer where the network request passes through.
[0084] Back-end collection: the server side aggregates channel features, and combines historical databases and external data sources to supplement associated features.
[0085] For feature preprocessing, the following strategies can be flexibly selected according to actual scene requirements:
[0086] Outlier treatment: Optional outlier treatment methods include truncation (limiting values beyond a threshold to the threshold range), log transformation (reducing the impact of extreme values), or using robust statistics (such as median instead of mean).
[0087] Missing value treatment: In addition to the conventional mean / median filling, forward / backward filling, a collaborative filling method based on similar users can also be used, or a special missing value prediction model can be trained.
[0088] Feature encoding: For categorical features, in addition to one-hot encoding, target encoding, entity embedding, or hash encoding can also be used to balance expression ability and dimension explosion problem.
[0089] Step 1.2, context feature vector construction;
[0090] According to the embodiments of the present application, the system converts the preprocessed multi-dimensional context features into a computable numerical representation to construct a context feature vector. The specific process is as follows:
[0091] Continuous features (such as time intervals, geographic coordinates, etc.) are converted to numerical values in the range of [-1, 1] through normalization processing.
[0092] Discrete features (such as device type, network type, etc.) are converted to binary vectors through one-hot encoding.
[0093] Text features (such as address description, event description, etc.) are converted to dense vectors through text embedding technology.
[0094] After processing each dimension feature, they are combined to form a unified context feature vector, which includes time dimension, space dimension, environment dimension and event dimension feature sub-vectors.
[0095] In the financial pre-account opening scene, the implementation method of context feature vector construction is as follows:
[0096] Time dimension feature processing: Convert the time information of pre-account opening application into multiple time features
[0097] Time point feature: Convert the application time to hour system (0-23) and normalize it to a value in the range of [-1, 1].
[0098] Periodic feature: Use sine and cosine functions to process periodic time features (such as hours, days of the week, months, etc.) to avoid time wraparound problems.
[0099] Holiday feature: Encode holiday information into a binary feature (0 for non-holiday, 1 for holiday), and further subdivide it into different types of holidays.
[0100] Time interval features: Calculate the time interval from historical operations, such as the time interval from the last login, the time interval from the last transaction, etc., and normalize them through logarithmic transformation.
[0101] Spatial dimension feature processing: Process geographic and network location information
[0102] IP address features: Convert IP addresses into geographic location information (such as country, city), and further convert them into one-hot encoding vectors.
[0103] Geographic coordinate features: Standardize latitude and longitude coordinates to values within the range [-1, 1].
[0104] Geographic area features: Generate area type features (such as commercial area, residential area, etc.) based on geographic location and convert them into one-hot encoding vectors.
[0105] Movement trajectory features: Calculate features such as position change speed, direction, etc., to identify abnormal position jumps.
[0106] Environment dimension feature processing: Process device and network environment information.
[0107] Device features: Convert device type, operating system, browser type, etc. into one-hot encoding vectors.
[0108] Device fingerprint: Use a hash function to convert device fingerprint information into a fixed-length feature vector.
[0109] Network characteristics: Convert network type, connection quality, etc. into numerical or one-hot encoding vectors.
[0110] Security features: Convert security software status, VPN usage, etc. into binary feature vectors.
[0111] Event dimension feature processing: Process external events and context information.
[0112] Market events: Standardize market volatility indicators to values within the range [-1, 1].
[0113] Regulatory events: Convert regulatory policy changes into one-hot encoding vectors to represent different types of policy changes.
[0114] Marketing activities: Encode marketing activity information into binary features (0 represents no activity, 1 represents activity), and further subdivide them into different types of activities.
[0115] Risk events: Convert recent risk event information into feature vectors, including event type, severity, etc.
[0116] In the feature fusion stage, the dimension feature vectors are combined by feature importance weighting, and the weight coefficients are obtained by feature importance evaluation on the validation set.
[0117] Step 1.3, dynamic neighbor search and similarity calculation;
[0118] According to the embodiments of the present application, the system uses a dynamic neighbor search algorithm to find similar context cases in historical data based on the constructed context feature vector. The specific implementation is as follows:
[0119] The locality-sensitive hashing (LSH) algorithm is used to index the historical context feature vectors, improving the retrieval efficiency of large-scale data.
[0120] The current context feature vector and the historical context feature vector are calculated for multiple distance measures, such as Euclidean distance, cosine similarity, and Mahalanobis distance.
[0121] Combining the results of multiple distance measures, a comprehensive similarity score is generated by weighted fusion.
[0122] It should be noted that the formula for calculating the comprehensive similarity score is:
[0123] ;
[0124] Among them, represents the comprehensive similarity score between the current context feature vector and the historical context feature vector; represents the context feature vector of the current pre-opening application; represents the context feature vector in the historical data; represents the normalized Euclidean distance, which is used to measure the straight-line distance between two vectors in geometric space; represents the cosine similarity, which is used to measure the similarity of the direction of two vectors; represents the normalized Mahalanobis distance, which considers the correlation between features and data distribution; 、 、 are the weight coefficients of Euclidean distance, cosine similarity, and Mahalanobis distance, respectively.
[0125] According to the comprehensive similarity score, the top-K most similar historical context cases are selected as the basis for subsequent risk feature extraction and evaluation.
[0126] In the financial pre-opening risk assessment scenario, the implementation and application of the dynamic neighbor search algorithm are as follows:
[0127] Local sensitive hashing index construction:
[0128] Multiple hash functions are constructed for LSH using the random projection method, and high-dimensional context feature vectors are projected into a low-dimensional space.
[0129] For each hash function, a projection vector and a threshold are randomly generated, and the hash value is calculated.
[0130] Combining multiple hash functions forms a hash signature, mapping similar context feature vectors into the same bucket.
[0131] Multiple hash tables are constructed to improve recall rate, and each hash table uses a different combination of hash functions.
[0132] For high-dimensional sparse features in the financial scenario, MinHash technology is used to process the similarity calculation of classification features.
[0133] Dynamic parameter adjustment:
[0134] According to the query load and system resources, dynamically adjust the LSH parameters, including the number of hash tables, the number of hash functions per signature, etc.
[0135] In high-risk periods (such as before and after holidays) or when abnormal access patterns occur, automatically improve search accuracy and increase the number of hash tables.
[0136] Set adaptive threshold, dynamically adjust similarity threshold according to historical data distribution and current load.
[0137] Similarity calculation optimization:
[0138] Feature weighting: dynamically adjust feature weights according to the importance of each feature in different contexts. For example, during the holiday period, the weight of the time feature may increase.
[0139] Multi-granularity matching: simultaneously calculate similarity on different feature subsets, such as only based on spatial features, only based on device features, etc., and then integrate the similarity results of each subset.
[0140] Context-aware distance: adjust distance measurement according to the speed and direction of context change, giving higher attention to rapidly changing contexts.
[0141] Similar case screening and application:
[0142] Risk stratification screening: first filter out historical cases with similar risk levels, then further search for cases with similar contexts.
[0143] Time decay: apply a time decay factor to historical cases, giving higher weights to more recent similar cases.
[0144] Abnormal situation handling: When no enough similar historical cases can be found, use clustering algorithm to identify the differences between current situation and known situations, and generate abnormal situation alerts.
[0145] For new pre-account opening applications, the system extracts their multi-dimensional situation features, constructs a situation feature vector, quickly finds candidate similar cases through LSH, then calculates the exact similarity score, identifies the most similar historical cases and their risk assessment results, and provides a reference for current risk assessment.
[0146] The situation analysis processing module is used to convert multi-dimensional situation features into a situation feature vector, use a dynamic nearest neighbor search algorithm to find similar situation cases in historical data, and extract hierarchical risk features from the identified similar situation cases.
[0147] Specifically, the following steps are included:
[0148] Step 2.1, hierarchical risk feature extraction;
[0149] For similar situation cases identified from step 1, use hierarchical feature extraction method to extract risk features from three levels:
[0150] Account level features: including identity information consistency of pre-account opening applicants, account attribute features, historical credit records, and other features related to account subjects.
[0151] Identity information consistency: generate consistency score by comparing identity information in different dimensions (such as consistency of ID photo and real-time face, consistency of filled information and certificate information, etc.).
[0152] Account attribute features: including account type, expected transaction size, application channel, and other basic account characteristics.
[0153] Historical credit records: if there are associated historical credit data, extract relevant credit features.
[0154] Transaction level features: including operation behavior features, input features, and time sequence patterns in the pre-account opening process.
[0155] Operation behavior features: including mouse movement trajectory, click pattern, dwell time, operation sequence, and other behavior features.
[0156] Input features: including input rate, modification frequency, input error pattern, and other features.
[0157] Time sequence features: extract time sequence features of behavior by analyzing time sequence patterns of operation.
[0158] Relationship level features: including association with other accounts, social network features, and group behavior patterns.
[0159] Associated account risk: Identify associated accounts through device fingerprints, IP addresses, contacts, etc. Extract the risk status and risk propagation pattern of the associated account.
[0160] Social network features: If there is associated social network data, extract relevant social network features.
[0161] Group behavior pattern: Compare the current application behavior with the behavior pattern of the same group to identify abnormal deviations.
[0162] The mathematical representation of hierarchical risk feature extraction is:
[0163] ;
[0164] ;
[0165] ;
[0166] where, represents the account layer feature vector, which contains a set of various feature data related to the account subject; represents the transaction layer feature vector, which contains a set of various feature data related to the operation behavior; represents the relationship layer feature vector, which contains a set of various feature data related to the associated relationship; represents the account layer feature extraction function; represents the transaction layer feature extraction function; represents the relationship layer feature extraction function, which is used to extract valuable features from relationship information; represents the account information of the current pre-account opening application; represents the transaction information of the current pre-account opening application; represents the relationship information of the current pre-account opening application; represents the account information of similar historical cases; represents the transaction information of similar historical cases; represents the relationship information of similar historical cases, which is used for comparative analysis with the current case.
[0167] In the financial pre-account opening scenario, the specific implementation and application of the hierarchical feature extraction method are as follows:
[0168] Account layer feature extraction implementation:
[0169] Identity information consistency calculation: Use multi-modal matching algorithm to compare identity information from different sources;
[0170] Photo and real-time face comparison: Use a deep face recognition model to extract face features and calculate the cosine similarity;
[0171] Text Information Consistency: Using a combination of character-level edit distance and semantic similarity to compare the filled information with the certificate information.
[0172] Multi-source Information Cross-Verification: Construct a verification graph to check the consistency relationship between different information sources and calculate the information conflict degree.
[0173] Historical Record Correlation Calculation: Use entity linking technology to associate the applicant's historical records in different systems;
[0174] Fuzzy Matching: Use fuzzy matching technology on key fields such as name and ID number to handle possible spelling errors or format differences;
[0175] Multi-dimensional Matching: Combine multiple dimensions of information such as contact information and address to improve accuracy;
[0176] Association Strength Scoring: Calculate the credibility score of the association based on the number and degree of matching fields.
[0177] Transaction Layer Feature Extraction Implementation:
[0178] Operation Behavior Serialization: Convert operation behaviors into standardized sequence representation;
[0179] Behavior Coding: Encode basic operations such as clicks, inputs, and stays into fixed-length vectors;
[0180] Sequence Segmentation: Use sliding window and change point detection algorithms to segment operation sequences into meaningful paragraphs;
[0181] Sequence Feature Extraction: Apply recurrent neural networks or time convolution networks to extract sequence features.
[0182] Abnormal Pattern Detection: Identify abnormal operation patterns in the pre-opening process:
[0183] Rhythm Analysis: Calculate the distribution of time intervals between operations to detect abnormal rhythm changes;
[0184] Hesitation Indicator: Analyze the frequency of input, delete, and modify operations to identify hesitant or uncertain behavior patterns;
[0185] Proficiency Evaluation: Evaluate the applicant's operation proficiency based on indicators such as operation fluency and error rate.
[0186] Relationship Layer Feature Extraction Implementation:
[0187] Associated Account Network Construction: Construct an account relationship network based on multiple association signals;
[0188] Direct Association: Construct direct connections based on explicit associations such as shared devices, IP addresses, and contacts.
[0189] Indirect Association: Discover multi-hop connections through graph algorithms to identify hidden association paths.
[0190] Association Strength Quantification: Calculate association strength scores based on the number, timeliness, and uniqueness of association signals.
[0191] Risk Propagation Analysis: Evaluate the propagation characteristics of risks in the association network.
[0192] Risk Diffusion Model: Use graph diffusion models to simulate the propagation process of risks in the network.
[0193] Key Node Identification: Calculate node centrality indicators to identify key nodes in risk propagation.
[0194] Risk Aggregation Detection: Use community detection algorithms to identify risk account aggregation phenomena.
[0195] Feature Interaction and Combination:
[0196] Cross-layer Feature Interaction: Capture the interaction between features at different levels.
[0197] Feature Cross: Generate cross combinations of features at different levels, such as combinations of account attributes and operation behaviors.
[0198] Conditional Correlation: Analyze differences in behavior patterns under specific account conditions.
[0199] Hierarchical Attention Mechanism: Use attention mechanisms to dynamically adjust the importance weights of features at different levels.
[0200] Integration of Financial Domain Knowledge: Combine financial risk control domain knowledge to guide feature extraction.
[0201] Risk Label Features: Based on historical risk cases, extract features related to known risk types.
[0202] Industry Rule Conversion: Convert financial industry risk control rules into constraint conditions for feature extraction.
[0203] Regulatory Focus Mapping: According to regulatory priorities, strengthen the feature extraction strength of related dimensions.
[0204] Step 2.2, Multi-layer Feature Fusion and Risk Assessment
[0205] Use weighted ensemble learning methods to integrate multi-level risk features to generate the final risk assessment results. The specific implementation is as follows:
[0206] Feature Importance Weight Learning: Learn the importance weights of each feature to risk assessment through algorithms such as gradient boosting decision trees, ensuring that the model focuses on the most discriminative features.
[0207] The feature importance weight calculation formula is:
[0208] ;
[0209] Where, represents the importance weight of feature ; represents the importance score of feature in the decision tree ; represents the total number of decision trees in ensemble learning; represents the total number of features used in the model; is the index of the decision tree, from 1 to ; is the index of the feature, from 1 to ; The numerator part calculates the total importance of feature in all decision trees, and the denominator part calculates the total importance of all features in all decision trees.
[0210] Multi-model weighted ensemble: adopt multiple machine learning models for integration, including:
[0211] Logistic regression: captures linear relationships between features.
[0212] Random forest: handles non-linear relationships and feature interactions.
[0213] Gradient boosting decision tree: step-by-step optimization of model performance.
[0214] Neural network: learns complex feature patterns.
[0215] Each model outputs a risk score, and then a weighted average is generated to generate the final risk score:
[0216] ;
[0217] Where, represents the final risk score; represents the risk score of the th model; represents the weight coefficient of the th model; represents the total number of models participating in the integration; the sum of all model weight coefficients satisfies , ensuring that the final score remains within a reasonable proportion range.
[0218] Risk level division: based on the risk score, set threshold to divide the risk into different levels (such as low risk, medium risk, high risk and extremely high risk).
[0219] The risk level division formula is:
[0220] ;
[0221] wherein, represents the risk level, which is the final risk assessment result, divided into four levels of low risk, medium risk, high risk and extremely high risk; represents the risk score, which is a comprehensive risk score calculated by multi-model weighted integration; is the demarcation point between low risk and medium risk; is the demarcation point between medium risk and high risk; is the demarcation point between high risk and extremely high risk.
[0222] Risk factor explanation: the SHAP (SHapley Additive exPlanations) value is used to calculate the contribution of each feature to the risk score, providing an interpretable risk factor for the risk assessment result.
[0223] The calculation formula of SHAP value is:
[0224] ;
[0225] wherein, represents the SHAP value of feature , i.e. the contribution of feature to the model prediction result; is the set of all features; is the subset of features excluding feature ; represents the number of features in the subset; represents the total number of features; and are the factorial of the subset size and the factorial of the number of remaining features (excluding feature ), respectively; is the combination weight, used to balance the influence of different size subsets; represents the model prediction value containing feature and subset ; represents the model prediction value using only the feature subset ; represents the marginal contribution of the model prediction value after adding feature .
[0226] Risk assessment module, for adopting weighted integrated learning method, fusing hierarchical risk features, generating risk score and risk classification result;
[0227] Specifically comprising the following steps:
[0228] Step 3.1, Multi-source Risk Information Collection;
[0229] Risk-related information is obtained from the following multiple data sources:
[0230] Internal business feedback data:
[0231] Fraud case reports: Contains details of confirmed fraud cases, fraud methods and identification methods.
[0232] Customer complaint records: Contains customer feedback and dispute information on risk assessment.
[0233] Manual review records: Contains the amendments and supplementary opinions of risk control personnel on system evaluation results.
[0234] Risk event data:
[0235] Identified fraud patterns: Contains the characteristics of fraud patterns identified by the system automatically or manually.
[0236] Security event records: Contains information about account security-related events such as account theft, information leakage, etc.
[0237] Risk monitoring reports: Contains risk trends and abnormal phenomenon reports generated regularly.
[0238] Regulatory compliance information:
[0239] Risk control requirement updates: Contains the latest requirements of regulatory agencies on risk control.
[0240] Compliance check reports: Contains risk points and improvement suggestions found during compliance checks.
[0241] External intelligence data:
[0242] Industry risk alerts: Contains risk warning information published by industry organizations.
[0243] Third-party risk data: Contains risk intelligence obtained from third-party risk data providers.
[0244] Public security information: Contains information about security vulnerabilities, attack methods, etc. from public channels.
[0245] Step 3.2, Structured Extraction of Risk Knowledge;
[0246] According to the embodiments of the present application, the system uses natural language processing and knowledge extraction technology to extract structured risk knowledge from unstructured data. The specific implementation is as follows:
[0247] Text preprocessing: Perform preprocessing operations such as cleaning, word segmentation, and removing stop words on original text data to improve the quality of subsequent processing.
[0248] Named Entity Recognition: Identify key entities in text such as risk types, risk objects, risk indicators, etc. Implemented using a Bidirectional Long Short-Term Memory network with Conditional Random Fields (BiLSTM-CRF) model.
[0249] In the financial pre-account opening risk assessment scenario, the specific structure and implementation of the BiLSTM-CRF model are as follows:
[0250] Input Layer: Receives vector representations of each word in the text sequence;
[0251] Word Embedding Representation: Initialized using pre-trained word vectors (such as Word2Vec or GloVe) and fine-tuned through the training process;
[0252] Character-level representation: Extract character-level features of each word through a character-level convolutional neural network or LSTM;
[0253] Domain feature enhancement: Merge financial risk control domain dictionary features to mark common risk terms and entities.
[0254] BiLSTM layer: Capture sequence context information;
[0255] Forward LSTM: Process the sequence from left to right, capturing right context information;
[0256] Reverse LSTM: Process the sequence from right to left, capturing left context information;
[0257] Output concatenation: Concatenate the outputs of the forward LSTM and reverse LSTM in the feature dimension to obtain a representation containing bidirectional context,
[0258] CRF layer: Perform sequence labeling considering label dependencies;
[0259] Transition matrix: Learn the transition probabilities between labels (such as B-risk type followed by I-risk type);
[0260] Viterbi decoding: Use dynamic programming algorithm to efficiently find the optimal label sequence;
[0261] To adapt to the special needs of financial pre-account opening risk assessment, the standard BiLSTM-CRF model is optimized as follows in this application:
[0262] Entity type expansion: Customize the entity type system for the financial risk field, including risk levels, risk indicators, risk triggers, etc.
[0263] Domain dictionary integration: Integrate financial risk control professional dictionaries into the feature extraction process to improve the accuracy of professional term recognition;
[0264] Data augmentation strategies: Expand training data through synonym replacement, template generation, etc., especially for rare risk type samples.
[0265] Hierarchical annotation: Support nested entity recognition, handle complex expressions such as "application with {fake ID} risk".
[0266] Relationship extraction: Identify semantic relationships between entities, such as "cause", "contain", "associate", etc., and build a relationship network of risk knowledge. Relationship extraction mainly uses remote supervision multi-instance learning method.
[0267] Event extraction: Extract risk event triggers and participants from text and build event knowledge. Event extraction uses hierarchical attention mechanism to identify event triggers and then event elements.
[0268] Attribute extraction: Extract attribute information of risk knowledge, such as risk level, occurrence probability, impact range, etc., to enrich knowledge representation.
[0269] In the financial pre-account opening risk assessment scenario, the specific implementation and application of knowledge extraction technology are as follows:
[0270] Financial domain named entity recognition (NER):
[0271] Financial risk-specific entity type definition: Extend the standard NER system and define financial risk-specific entity types such as fraud types, risk factors, risk indicators, and regulatory requirements.
[0272] Domain dictionary enhancement: Build a financial risk-specific dictionary containing common fraud methods, risk terms, regulatory requirements, etc., and enhance the recognition ability of the BiLSTM-CRF model through dictionary features.
[0273] Annotated data construction: Use semi-supervised methods to expand annotated data, use pattern matching for preliminary annotation, manually verify high-confidence samples, and form an iterative annotation process.
[0274] Application case: Extract "credit card cashing" and "fake application" fraud type entities from risk event reports, and extract "customer identity recognition" regulatory requirement entities from regulatory notices.
[0275] Financial risk relationship extraction:
[0276] Relationship type system: Build a relationship type system for the financial risk domain, including causal relationships (cause, trigger), containment relationships (belongs to, includes), and association relationships (related, co-occurrence).
[0277] Remote supervision data construction: Utilize existing risk knowledge base as a source of remote supervision, automatically generate relationship annotation data.
[0278] Noise processing technology: Adopt sentence-level selective attention mechanism to reduce the impact of noise introduced by remote supervision.
[0279] Application case: Extract the causal relationship between "false identity information" and "refusal to open an account" from risk reports, and the correlation between "device fingerprint anomaly" and "gang fraud" from fraud cases.
[0280] Financial risk event extraction:
[0281] Event type definition: Define event types related to financial pre-account opening risks, such as fraud events, regulatory events, and market events.
[0282] Event element extraction: For each type of event, define specific element roles (such as fraud subject, fraud object, fraud means, time and place, etc.).
[0283] Cross-document event coreference resolution: Handle different descriptions of the same risk event in different documents to build complete event knowledge.
[0284] Application case: Extract the "batch of false pre-account opening applications in a certain region" event from news reports, including event elements such as time, place, fraud means, and involved institutions.
[0285] Financial risk attribute extraction:
[0286] Attribute type definition: Define key attribute types for financial risk entities and events, such as risk level, occurrence probability, impact range, and timeliness.
[0287] Numerical standardization: Standardize the risk level expressed in different texts (such as "extremely high risk" and "relatively high risk") into a unified numerical representation.
[0288] Uncertainty handling: Identify and retain uncertainty information (such as "may" and "expected") expressed in the text as a measure of the credibility of risk knowledge.
[0289] Application case: Extract the "credit card cashing risk" level attribute as "high risk", the impact range attribute as "nationwide", and the timeliness attribute as "continuous growth" from risk assessment reports.
[0290] Financial risk knowledge fusion:
[0291] Multi-source knowledge consistency verification: Verify the consistency of risk knowledge extracted from different sources, identify and resolve conflicting information.
[0292] Knowledge completion: Utilize knowledge reasoning techniques to complete the missing key information in the extraction process.
[0293] Knowledge quality evaluation: Evaluate the quality of extracted knowledge based on information source reliability, extraction confidence, and knowledge completeness.
[0294] Application case: Integrate the knowledge about "fake identity opening accounts" extracted from internal fraud case reports, external industry notifications, and regulatory notices to form a complete risk knowledge representation.
[0295] Therefore, in actual financial pre-account opening risk assessment systems, these knowledge extraction techniques automatically extract risk knowledge from a large amount of unstructured text, enrich the risk knowledge graph, and provide knowledge support for risk assessment. For example, when the system extracts a new fraud pattern about "cross-border IP + virtual mobile number + batch application" from multiple channels, it can quickly incorporate it into the risk assessment model, improving the ability to identify such fraudulent behavior.
[0296] Step 3.3, risk knowledge graph construction and integration;
[0297] Integrate the extracted structured risk knowledge with the existing risk knowledge graph to build a unified risk knowledge system. The specific implementation is as follows:
[0298] Knowledge representation: Represent risk knowledge in the form of entity-relation-entity triples, formally represented as where , represent two entities, and represent the relationship between entities.
[0299] Knowledge fusion: Handle the heterogeneity and redundancy of knowledge from different sources, mainly including:
[0300] Entity alignment: Identify and merge different entities representing the same concept, using entity representation learning and similarity calculation methods. The entity similarity calculation formula is:
[0301] ;
[0302] where represents the similarity between entities and ; represents the cosine similarity function, and represent the vector representation of entities and , represents the dot product (inner product) of two vectors, and represent the vector and Euclidean norm (vector length) of and, the entire formula calculates the cosine similarity between two entity vector representations, which measures the semantic similarity of two entities.
[0303] Relationship mapping: Establish mapping of relationship types in different knowledge sources to ensure semantic consistency.
[0304] Conflict resolution: When there is a knowledge conflict, decide which version to keep based on factors such as knowledge source reliability, timeliness, etc.
[0305] Knowledge graph expansion: Add new knowledge after fusion to the existing knowledge graph to expand and update the risk knowledge base.
[0306] Knowledge reasoning: Based on existing knowledge, infer and discover implicit risk knowledge. The following reasoning methods are mainly used:
[0307] Path reasoning: Reasoning based on path patterns in the knowledge graph.
[0308] Rule-based reasoning: Knowledge derivation based on predefined reasoning rules.
[0309] Embedded reasoning: Reasoning based on knowledge graph embedding model to predict the potential relationship between entities. The scoring function of the embedding model is:
[0310] ;
[0311] where, represents the scoring function of the embedding model; represents the vector representation of the head entity, i.e. the embedding vector of the starting entity of the relationship in the knowledge graph; represents the vector representation of the relationship, i.e. the embedding vector of the connection relationship between entities; represents the vector representation of the tail entity, i.e. the embedding vector of the target entity of the relationship in the knowledge graph; represents the distance norm between the head entity vector plus the relationship vector and the tail entity vector, the smaller the value, the more likely the triple relationship is established.
[0312] Step 4.1, risk knowledge quality assessment;
[0313] Specifically, the following steps are included:
[0314] Step 4.1, risk knowledge quality assessment;
[0315] Multi-dimensional quality assessment of newly extracted risk knowledge to ensure its reliability and effectiveness. Assessment dimensions include:
[0316] Knowledge consistency check: Check if the new knowledge conflicts with the existing knowledge system. Main methods include:
[0317] Logical consistency check: Check if the new knowledge is logically inconsistent with existing knowledge, such as whether conclusions derived through reasoning conflict with each other.
[0318] Numerical consistency check: Check if the numerical properties in the new knowledge (such as risk probability, impact degree, etc.) are significantly different from the existing knowledge.
[0319] Semantic consistency check: Check if the semantic expression of the new knowledge is consistent with the existing knowledge.
[0320] In some embodiments, more complex consistency checking methods can be used:
[0321] Hierarchical consistency check: Verify knowledge consistency at different levels of abstraction, from specific facts to abstract rules.
[0322] Contextual consistency check: Dynamically assess the consistency of knowledge according to different context conditions, allowing exceptions in specific situations.
[0323] Probability consistency check: For uncertain knowledge, use a probability framework to assess the compatibility of new and old knowledge.
[0324] Consistency score calculation formula:
[0325] ;
[0326] Where, represents the consistency score, ranging from 0 to 1, the higher the value, the better the consistency; represents the new knowledge and the existing knowledge conflict, that is, the number of knowledge items that are inconsistent or inconsistent between new and old knowledge; represents the total number of new knowledge, that is, the total number of new knowledge items; score represents the conflict ratio, 1 minus the ratio to get the final consistency score. When there is no conflict, the consistency score is 1; when all new knowledge conflicts with existing knowledge, the consistency score is 0.
[0327] Knowledge importance assessment: Assess the importance of new knowledge to risk assessment. Assessment indicators include:
[0328] Risk correlation: The degree of association between knowledge and known high-risk patterns.
[0329] Impact Scope: The extent of business areas or applicable scenarios that the knowledge influences.
[0330] Knowledge Scarcity: The scarcity of knowledge in the existing knowledge base, the more scarce the more important.
[0331] Importance Score Calculation Formula:
[0332] ;
[0333] Where, represents the total score of knowledge importance; represents the association degree score of knowledge with known high-risk patterns; represents the impact scope score of knowledge , reflecting the extent of business scenarios that the knowledge applies to; represents the scarcity score of knowledge in the existing knowledge base, the more scarce the knowledge, the higher the score; , , respectively are the weight coefficients of the association degree score, impact scope score and scarcity score.
[0334] Knowledge Time Sensitivity Assessment: Assess the validity period and time sensitivity of knowledge. The assessment dimensions include:
[0335] Knowledge Generation Time: The time of knowledge extraction or creation.
[0336] Knowledge Update Frequency: The historical update frequency of similar knowledge.
[0337] Knowledge Life Cycle: Estimate the validity period of knowledge according to the characteristics of the field.
[0338] For example, for knowledge related to specific activities or seasonal risks, a corresponding time sensitivity decay model can be set to dynamically adjust the weight of knowledge over time.
[0339] Time Sensitivity Score Calculation Formula:
[0340] ;
[0341] Where, represents the time sensitivity score; represents the current time, i.e. the time point of evaluation; represents the knowledge generation time; represents the expected life cycle of knowledge; is the time decay coefficient; represents the proportion of the time that the knowledge has existed to its expected life cycle; denotes an exponential decay function, the timeliness score of knowledge gradually decreases from 1 over time, but always remains greater than 0.
[0342] Knowledge source reliability assessment: assess knowledge quality based on the reliability of knowledge sources. Evaluation indicators include:
[0343] Source authority: the professional authority degree of the knowledge source.
[0344] Historical accuracy: the accuracy of historical knowledge provided by the source.
[0345] Source consistency: the consistency of different sources on the same knowledge.
[0346] In some embodiments, a more comprehensive source evaluation system can be constructed:
[0347] Multi-layer source evaluation: distinguish between direct and indirect sources, and evaluate the reliability of different levels of sources respectively.
[0348] Source relevance assessment: assess the professional relevance of knowledge sources to the field of knowledge content.
[0349] Source trust network: build a trust relationship network between sources, and calculate the comprehensive reliability through network propagation.
[0350] Source reliability calculation formula:
[0351] ;
[0352] where, denotes the reliability score of the knowledge source; denotes the number of knowledge provided by the source historically; denotes the knowledge verified accuracy; denotes the sum of the accuracy of all historical knowledge provided by the source; denotes the average of the accuracy of all historical knowledge provided by the source.
[0353] Comprehensive quality score calculation formula:
[0354] ;
[0355] where, denotes the comprehensive quality score of knowledge; denotes the knowledge consistency score, reflecting the logical consistency of new knowledge with the existing knowledge system; denotes the knowledge importance score, measuring the impact of knowledge on risk assessment; denotes the knowledge timeliness score, assessing the effective period and time sensitivity of knowledge; The reliability score of the knowledge source, based on the historical accuracy rate and professional authority of the knowledge provider; 、 、 、 are the weight coefficients of the knowledge consistency score, knowledge importance score, knowledge timeliness score, and reliability score of the knowledge source, respectively, and satisfy , ensuring weight normalization.
[0356] Step 4.2, dynamic updating of the risk knowledge graph;
[0357] Based on the quality evaluation results, dynamically update the risk knowledge graph to ensure the timeliness and accuracy of the risk knowledge. The specific implementation is as follows:
[0358] Update strategy formulation: determine the update strategy according to the knowledge quality score, mainly including:
[0359] Direct addition: for new knowledge with high quality and no conflict with existing knowledge, directly add it to the knowledge graph.
[0360] Replacement update: for new knowledge with higher quality than existing knowledge and conflict, replace the existing low-quality or outdated knowledge.
[0361] Conditional integration: for new knowledge that partially conflicts with existing knowledge but has its own value, perform conditional integration to clarify the applicable conditions.
[0362] Temporary storage verification: for new knowledge with low quality or major conflicts, temporarily store it for further verification.
[0363] Knowledge graph update operation: perform specific update operations on the knowledge graph, including:
[0364] Add operation: add new entities, relationships, or attributes to the knowledge graph.
[0365] Delete operation: delete outdated, incorrect, or redundant knowledge from the knowledge graph.
[0366] Modification operation: update the properties or relationships of existing knowledge in the knowledge graph.
[0367] Reorganization operation: adjust the structure of the knowledge graph to optimize knowledge organization.
[0368] Version management and rollback mechanism: through knowledge graph version control, support tracking and rollback of knowledge updates when necessary, including:
[0369] Incremental snapshot: record the content and time of each update.
[0370] Change log: detailed record of knowledge change history and reasons.
[0371] Rollback operation: Support rolling back the knowledge graph to a previous state when issues are discovered.
[0372] Knowledge evolution tracking: Analyze the evolution trends and patterns of risk knowledge, including:
[0373] Trend analysis: Identify long-term trends in knowledge changes, such as growth or decline trends of certain risks.
[0374] Periodic analysis: Discover periodic patterns in knowledge changes, such as seasonal risk variations.
[0375] Mutation detection: Identify sudden changes in knowledge, warning potential risk outbreaks.
[0376] The knowledge evolution model adopts a time series decomposition method to decompose the knowledge change sequence into trend items , seasonal items and residual items :
[0377] ;
[0378] By analyzing the change characteristics of each item, identify the laws and abnormalities of knowledge evolution.
[0379] Crowdsourcing knowledge collection module, for integrating the risk identification experience of front-line staff and experts through crowdsourcing mechanism, applying the collected knowledge to risk assessment process, forming a continuous optimization closed loop of risk assessment;
[0380] Specifically, the following steps are included:
[0381] Step 5.1, crowdsourcing knowledge collection mechanism construction;
[0382] Establish a systematic crowdsourcing knowledge collection channel and process, supporting efficient collection of risk knowledge from multiple sources. The specific implementation is as follows:
[0383] Knowledge contribution channel construction: Provide multiple channels to support front-line personnel and experts to contribute risk knowledge, including:
[0384] Structured form: Collect standardized risk knowledge through pre-set fields.
[0385] Natural language interface: Support describing risk knowledge in natural language and converting it to structured knowledge through natural language processing technology.
[0386] Case labeling tool: Support labeling risk features and judgment basis on actual cases.
[0387] Rule editor: Provide a visual rule editing interface to support experts directly editing risk judgment rules.
[0388] In some implementations, the following advanced contribution channels may also be provided:
[0389] Knowledge Graph Editor: Directly edit entities and relationships in the risk knowledge graph through a visual interface.
[0390] Scenario simulation tools: By simulating specific risk scenarios, experts can mark risk points in the simulated environment.
[0391] Intuitive API Interface: Provides a standardized knowledge contribution interface for third-party risk data service providers.
[0392] Incentive compatibility mechanism: Design incentive mechanisms to encourage high-quality knowledge contributions, including:
[0393] Knowledge Contribution Points System: Points are awarded based on the quality and quantity of knowledge contributed.
[0394] Professional Reputation System: Establish a professional reputation evaluation system to enhance the professional status of high-quality contributors.
[0395] Performance-based incentives: Provide feedback and rewards based on the actual application results of the contributed knowledge.
[0396] Alternatively, depending on organizational needs and culture, the following additional incentive methods may be adopted:
[0397] Collaborative innovation: Enabling contributors to participate in the improvement and application of their contributed knowledge, thereby enhancing their sense of accomplishment and participation.
[0398] Visualizing the impact: Visualizing the application effects and scope of the contributed knowledge enhances the contributors' sense of accomplishment.
[0399] Hierarchical permissions: As the quality and quantity of contributions increase, contributors are granted higher-level system permissions and participation permissions.
[0400] The incentive utility function is designed as follows:
[0401] ;
[0402] in, Indicates contributor The overall utility value; This indicates the number of points a contributor has earned. This indicates the magnitude of the impact of contributed knowledge in practical applications; This represents the time and effort invested by the contributors; It is the weighting coefficient for points rewards; It is the weighting coefficient of the actual impact; is the weight coefficient of cost. These three weight coefficients can be adjusted according to the incentive strategies of different institutions to balance short-term incentives and long-term value.
[0403] Knowledge collection strategy: adopt an active strategy to guide the collection of high-value knowledge, including:
[0404] Targeted collection: actively initiate knowledge collection targeting specific risk areas or issues.
[0405] Difference analysis: identify weak links in the knowledge base and supplement knowledge accordingly.
[0406] Hot spot tracking: track industry risk hotspots and collect relevant knowledge in a timely manner.
[0407] For example, when the system detects changes in risk patterns in a certain region or customer group, it can actively initiate targeted knowledge collection to frontline personnel in the relevant region to quickly obtain the latest risk knowledge.
[0408] Step 5.2, crowd-sourced knowledge quality assessment;
[0409] Quality assessment of collected crowd-sourced knowledge to filter high-quality risk knowledge. The specific implementation is as follows:
[0410] Multi-dimensional quality assessment: assess the quality of crowd-sourced knowledge from multiple dimensions, including:
[0411] Knowledge accuracy: assess the accuracy of knowledge content.
[0412] Knowledge completeness: assess the completeness of knowledge description.
[0413] Knowledge novelty: assess the innovation or uniqueness of knowledge.
[0414] Knowledge applicability: assess the applicability of knowledge in actual scenarios.
[0415] Optionally, the following assessment dimensions can also be included:
[0416] Knowledge timeliness: assess the relevance and timeliness of knowledge to the current risk situation.
[0417] Knowledge universality: assess the applicable scope of knowledge in different scenarios.
[0418] Knowledge operability: assess whether the knowledge can be converted into specific risk assessment rules or operations.
[0419] Quality assessment adopts multi-dimensional weighted scoring:
[0420] ;
[0421] wherein, Representing knowledge The quality score, Representing knowledge In the Scores in each dimension Representing dimensions The weight, This represents the total number of dimensions assessed (i.e., the number of quality assessment factors considered). This formula uses a weighted summation method to comprehensively consider the performance of knowledge across multiple dimensions, including accuracy, completeness, novelty, and applicability, to arrive at the final quality score. Weights It can be dynamically adjusted according to different business scenarios and risk types to highlight the importance of key dimensions.
[0422] Expert review and consensus mechanism: Combining expert review and group consensus to form a more reliable quality assessment, including:
[0423] Expert review process: Designated experts in the field review important knowledge.
[0424] Group voting mechanism: assessing the quality of knowledge through multiple votes.
[0425] Hierarchical review: Different review levels are set according to the importance and complexity of the knowledge.
[0426] In some implementations, a more efficient review mechanism can be employed:
[0427] Crowdsourcing verification: Different experts verify the same knowledge separately, and the reliability of the evaluation is enhanced by comparing opinions from multiple parties.
[0428] Case verification: Test the effectiveness of knowledge through real-world cases to verify the quality of knowledge in an empirical manner.
[0429] Incremental review: First, apply the technology to low-risk scenarios and conduct small-scale tests, then gradually expand the application scope based on performance.
[0430] Consensus score calculation formula:
[0431] ;
[0432] in, Representing knowledge The consensus score is the average degree to which the knowledge item is recognized by the evaluators. Indicates the first Each evaluator's assessment of knowledge The rating (ranging from 0 to 1, where 0 indicates complete disapproval and 1 indicates complete approval). This indicates the total number of evaluators who participated in assessing this knowledge item; This represents the arithmetic mean of all evaluators' scores. This formula quantifies the degree of collective acceptance of a particular knowledge item by calculating the average of all evaluators' scores.
[0433] Contributor Reputation Evaluation System: Establish a reputation evaluation system based on the quality and accuracy of contributions, including:
[0434] Historical contribution quality: Evaluates the average quality of a contributor's historical contributions.
[0435] Knowledge validation rate: The percentage of contributors whose knowledge is validated as correct.
[0436] Expertise Coverage: The scope of an contributor's expertise.
[0437] Reputation score calculation formula:
[0438] ;
[0439] in, Indicates contributor Reputation score Indicates contributor The entire collection of knowledge contributed Indicates contributor The total amount of knowledge contributed. Expressing gratitude to contributors Summing all the knowledge, Representing knowledge The quality score is calculated by averaging the quality scores of all the contributor's knowledge to assess the contributor's overall reputation level.
[0440] Step 5.3, Crowdsourced Knowledge Application and Optimization;
[0441] The quality-assessed crowdsourced knowledge is integrated into a risk knowledge graph and applied to the real-time risk assessment process. The specific implementation is as follows:
[0442] Crowdsourced knowledge integration: Integrating high-quality crowdsourced knowledge into the risk knowledge graph, including:
[0443] Knowledge structuring transformation: Transforming crowdsourced knowledge into a structured representation compatible with knowledge graphs.
[0444] Knowledge association establishment: Establishing connections between crowdsourced knowledge and existing knowledge.
[0445] Knowledge conflict resolution: Addressing potential conflicts between crowdsourced knowledge and existing knowledge.
[0446] Real-time risk assessment applications: Utilizing crowdsourced knowledge in the risk assessment process, including:
[0447] Rule embedding: transforming crowdsourced knowledge into risk assessment rules and embedding them into the assessment process.
[0448] Feature extraction guidance: using crowdsourced knowledge to guide the extraction and selection of risk features.
[0449] Decision basis support: providing risk decision-making basis based on crowdsourced knowledge.
[0450] Knowledge dissemination feedback loop: building a feedback mechanism for knowledge application effectiveness, forming a continuous optimization closed-loop system, including:
[0451] Application effect tracking: tracking the effectiveness of crowdsourced knowledge in actual application.
[0452] Knowledge impact assessment: assessing the impact of crowdsourced knowledge on risk assessment accuracy.
[0453] Feedback and optimization: based on application effect feedback, optimize knowledge structure and content.
[0454] Knowledge impact assessment formula:
[0455] ;
[0456] Where, represents the impact of knowledge , that is, the degree of improvement of the knowledge on risk assessment accuracy; represents the risk assessment accuracy after applying knowledge , that is, the ratio of correct identification of risks in the risk assessment system after adding the knowledge; represents the risk assessment accuracy without applying knowledge , that is, the ratio of correct identification of risks in the risk assessment system without the knowledge. This formula calculates the difference in accuracy before and after adding a specific knowledge, quantifies the actual contribution of the knowledge to the risk assessment effect, and provides an objective basis for knowledge value assessment.
[0457] Application examples of the present embodiment:
[0458] Application scenario introduction:
[0459] According to the embodiments of the present application, the following is an application example of the financial pre-account opening risk assessment method based on intelligent identification in the retail financial business of a large commercial bank. The bank launched an online pre-account opening service on the Internet bank platform, allowing customers to submit pre-account opening applications through mobile banking APP or online banking, complete preliminary identity verification and information collection, and then directly go to the offline branch to complete the simplified account opening confirmation process.
[0460] With the growth of online pre-account opening business, the bank faces the risk of daily 50,000 pre-account opening application review pressure, and the traditional manual review and fixed rule risk control system has been difficult to cope with. Especially during the marketing campaign, the pre-account opening volume has increased dramatically, and the fraud methods have been constantly updated, leading to an increase in risk missed detection and misjudgment rate, causing economic losses and affecting the service experience of normal customers.
[0461] In this context, the bank implemented the financial pre-account opening risk assessment method based on intelligent identification proposed in this application, and the following is the specific implementation process and effect.
[0462] Multi-dimensional context perception implementation example:
[0463] In actual application, the system collects and processes the following multi-dimensional context features:
[0464] Time dimension context perception: the system records the time information of the pre-account opening application and automatically converts it into multiple time features. For example, during a certain promotion campaign, the system found that the pre-account opening application volume increased abnormally in the 2-4 am period, which increased by 15 times compared with the regular period.
[0465] Through context perception, the system automatically increases the risk checking strength of this period, and identifies that these applications have a highly similar behavior pattern.
[0466] Application case: the system detects 50 pre-account opening applications submitted at 3 am, although each individual indicator is within the normal range, but through the matching of time context features with historical similar cases, it is found that it is highly similar to a gang fraud case confirmed a month ago. The system automatically marks this batch of applications as high risk and transfers them to the manual review queue, and finally confirms that it is a new round of attack attempt of the same fraud gang.
[0467] Spatial dimension context perception: the system integrates IP address, GPS positioning, communication base station information and other spatial information to construct the spatial behavior features of the applicant. The system pays special attention to the consistency and rationality of spatial information, such as the matching degree of IP address and GPS location, the rationality of location change, etc.
[0468] Application case: the system detects that the IP address of a certain applicant shows that it is located in a city in China, but the GPS positioning and base station information point to another city thousands of kilometers away, and completes login attempts in 3 different cities within 10 minutes. Through spatial context feature analysis, the system captures this spatial anomaly and automatically marks the application as high risk. Subsequent verification confirms that it is a fraud attempt using IP proxy and GPS simulator.
[0469] Device Environment Context Awareness: The system collects environment information such as device model, operating system, browser features, screen resolution, and for mobile devices, additional sensor data, accelerometer data, and other features to build a fine-grained device fingerprint.
[0470] Application Case: During a marketing campaign, the system detected 40 pre-account opening applications from different usernames and IP addresses, but their device fingerprint features were highly similar, especially in sensor data and touchscreen operation characteristics. Through environmental context feature analysis, the system identified that these applications may come from the same batch of simulation devices and triggered a risk alert. Subsequent investigation confirmed that it was a fraud using virtual devices for batch account opening.
[0471] Risk Knowledge Graph Construction Example:
[0472] In application, the system continuously extracts risk knowledge from multiple sources and constructs a risk knowledge graph:
[0473] Risk Expert Knowledge Collection: The bank organized 30 risk control experts, through the knowledge contribution interface provided by the system, initialized the risk knowledge graph. Experts defined 12 main risk types, 56 risk subtypes, and more than 200 risk indicators and their associated relationships. These knowledge is structured as entities and relationships in the knowledge graph.
[0474] Operational Data Knowledge Extraction: The system uses natural language processing technology to extract more than 2,800 risk knowledge from the past two years of risk control reports, fraud case analysis and audit records, including high-risk patterns in specific time periods, regional fraud features, and device environment anomaly indicators. These knowledge is automatically converted into structured risk rules and integrated into the risk knowledge graph.
[0475] Crowdsourcing Risk Knowledge Collection: The system deploys a crowdsourcing risk knowledge collection mechanism to allow frontline staff to easily contribute risk findings. In the first three months after going online, the system received 823 risk knowledge contributions, of which 618 were accepted and integrated into the risk knowledge graph after quality evaluation and expert review.
[0476] Application Case: A branch manager noticed a new type of fraud: fraudsters use specific types of mobile virtual positioning software to fake location information. The manager submitted this discovery through the crowdsourcing channel, including the characteristic identification of the suspicious positioning software. The system evaluated and quickly incorporated this knowledge into the risk graph and deployed corresponding detection rules across the bank. The rule successfully identified 42 similar fraud attempts in the first week after going online, avoiding potential losses.
[0477] Knowledge updating and evolution tracking: The system continuously tracks the effectiveness of risk knowledge and dynamically adjusts the importance of knowledge based on actual application results. For example, for the risk indicator "device fingerprint anomaly", the system finds that the effectiveness of this indicator varies significantly in different time periods and different customer groups by real-time statistics of its trigger rate and accuracy. Accordingly, the system automatically adjusts the risk weight of this indicator in different situations, improving the accuracy of overall risk assessment.
[0478] Real-time risk assessment case:
[0479] The following is an example of the risk assessment process of the system in actual business:
[0480] Multi-dimensional context feature extraction: When the user submits a pre-account opening application, the system collects multi-dimensional context features in real time, including time features (submitted at 10:25 on Wednesday morning), spatial features (IP address located in a business district in a city, GPS location matches IP location), device features (iOS15.4 device, Safari browser, device has been used for 6 months), etc.
[0481] Similar context case matching: The system uses a dynamic nearest neighbor search algorithm to retrieve the 50 most similar cases from 12 million historical cases. Analysis shows that among the historical cases in similar contexts, 92% are normal applications and 8% are fraud attempts, and the common feature of these fraud cases is the use of fake identity documents.
[0482] Hierarchical risk feature extraction: The system extracts identity information consistency features from the account level and finds that the similarity score between the applicant's ID photo and real-time face is 0.83, which is lower than the normal value range (usually >0.95); from the transaction level, it extracts operation behavior features and finds that the user has an abnormal modification pattern and pause when filling out personal information; from the relationship level, it detects that the device has indirect contact with a risk device marked 3 months ago.
[0483] Risk assessment and decision-making: Based on the above features, the system calculates the risk score of the application as 78 points (high risk threshold is 70 points), and the main risk factors include "low identity information consistency", "abnormal operation behavior" and "associated device risk". The system automatically transfers the application to the manual review queue and marks the specific risk points for reference by the auditors.
[0484] Risk knowledge updating: Manual review confirms that the application is indeed a fraud attempt using a PS modified ID photo. The system automatically extracts this newly discovered fraud pattern (specific image tampering features) as new risk knowledge, adds it to the risk knowledge graph after quality evaluation, and applies it in subsequent risk assessment.
[0485] Technical effect verification:
[0486] The financial pre-account opening risk assessment method based on intelligent identification proposed in the present application has achieved remarkable technical effects in the actual application of the bank:
[0487] Risk identification accuracy rate is improved: after implementing the method, the pre-account opening fraud detection rate is improved from 76.3% to 94.8%, an increase of 18.5 percentage points;
[0488] At the same time, the misjudgment rate is reduced from 7.2% to 2.8%, a decrease of 4.4 percentage points. Especially during the peak of marketing activities, when the daily application volume exceeds 100,000, the system still maintains a fraud detection rate of more than 90%, while the traditional system's detection rate decreases to about 60% under similar load.
[0489] Risk knowledge update efficiency is improved: in the traditional risk control system, it takes an average of 18.5 days from discovering a new fraud pattern to deploying corresponding prevention and control measures in the entire bank;
[0490] After implementing the method, this period is shortened to 1.2 days, an increase of 93.5%. Especially for high-risk new fraud patterns, the system can complete the update and deployment in the entire bank within 4 hours.
[0491] Operation cost is reduced: after implementing the method, the proportion of pre-account opening applications that need to be manually reviewed is reduced from 25% to 8%, reducing the human review workload by 68%.
[0492] At the same time, the system's automatic risk assessment shortens the average processing time of each pre-account opening application from 15 minutes to 45 seconds, improving processing efficiency and improving customer experience.
[0493] Risk explanation is enhanced: the system provides clear basis and explanation for each risk judgment, including the triggered risk rules, matched historical cases, and associated risk knowledge. This enables risk control personnel to better understand the risk judgment logic and provides sufficient basis for customer complaint handling, reducing the average dispute handling time by 65%.
[0494] Adaptability is improved: within 6 months of system deployment, the financial market has experienced multiple regulatory policy changes and marketing activity fluctuations. The system successfully adapts to these changes through situational awareness and knowledge graph dynamic updates, maintaining stable risk control effects. Especially during a nationwide marketing activity, the system successfully coped with the pressure peak of daily application volume increasing to 5 times the regular volume, maintaining normal risk assessment performance.
[0495] As shown in Figures 2 to 6 , respectively, are risk assessment accuracy rate comparison; risk knowledge update timeliness; risk assessment model performance; situational similarity and risk relevance; and crowd-sourced knowledge contribution quality distribution.
[0496] The above describes the embodiments of the present application, but the embodiments are not limited to the specific implementation described above, which is only illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the inspiration of the embodiments, which are all within the protection scope of the embodiments.
Claims
1. An intelligent recognition-based financial pre-account opening risk assessment system, characterized in that, The method comprises the following steps: A multi-dimensional context feature acquisition module is used to acquire multi-dimensional context features in the pre-opening process, including time dimension features, space dimension features, environment dimension features, and event dimension features; A context analysis processing module is used to convert the multi-dimensional context features into a context feature vector, and a dynamic nearest neighbor search algorithm is used to search for similar context cases in historical data, and hierarchical risk feature extraction is performed on the identified similar context cases, specifically including: Using a hierarchical local sensitive hashing algorithm to establish a multi-level index structure for historical context feature vectors; Calculating the Euclidean distance, cosine similarity, Mahalanobis distance, and Jaccard similarity coefficient between the current context feature vector and the historical context feature vector; Generating a comprehensive similarity score through an adaptive weighted fusion algorithm, and dynamically adjusting the weight coefficient according to the historical retrieval effect; using a hierarchical screening strategy, first quickly retrieving a candidate set, then accurately calculating the similarity, and selecting the most similar historical context case; Based on an online learning mechanism, dynamically adjusting the local sensitive hashing parameters, including the number of hash tables, the number of hash functions, and the bucket size, to adapt to changes in data distribution; A risk assessment module is used to use a weighted ensemble learning method to integrate hierarchical risk features to generate risk scores and risk classification results; A knowledge graph updating module extracts structured risk knowledge from multiple sources based on risk assessment results and related cases and integrates it with existing risk knowledge graphs, performs quality assessment on newly extracted risk knowledge, dynamically updates the risk knowledge graph based on the assessment results, specifically including: Performing knowledge consistency testing, using ontology reasoning technology to test whether the new knowledge conflicts with the existing knowledge system; Performing knowledge importance evaluation, based on historical case backtesting and Monte Carlo simulation to evaluate the impact of new knowledge on risk assessment; Performing knowledge timeliness evaluation, establishing a knowledge decay model to evaluate the effective period and time sensitivity of knowledge; Performing knowledge source reliability evaluation, constructing a multi-dimensional trust scoring system, and evaluating knowledge quality based on the historical accuracy, professional field authority, and information update frequency of the knowledge source; Using the analytic hierarchy process to calculate the weight of each evaluation dimension to generate a comprehensive quality score, and setting a dynamic threshold to determine the knowledge adoption strategy; A crowdsourcing knowledge collection module is used to integrate the risk identification experience of front-line staff and experts through a crowdsourcing mechanism, and apply the collected knowledge to the risk assessment process to form a continuous optimization closed loop for risk assessment, specifically including: Providing multiple modal knowledge contribution channels, including structured forms, natural language interfaces, case labeling tools, rule editors, and visual knowledge construction tools; Building a hierarchical incentive compatible mechanism to give differentiated point rewards and professional certifications based on the quality, quantity, application effect, and innovation of the contributed knowledge; Establishing a multi-dimensional evaluation framework for the quality of crowdsourced knowledge, including accuracy, completeness, novelty, applicability, and complementarity to existing knowledge; Using an evaluation method that combines expert review and collective wisdom, forming a quality evaluation result through the Delphi method and weighted voting mechanism; Design an integration algorithm to align and fuse high-quality crowd-sourced knowledge with the existing risk knowledge graph semantically, and apply it to the risk assessment process.
2. The intelligent recognition-based financial pre-account opening risk assessment system according to claim 1, characterized in that, The time dimension features collected by the multi-dimensional context feature collection module include the application time point, time period distribution, whether it is a holiday, the interval between the previous and subsequent application times, and seasonal activity cycles. The spatial dimension features include geographic location, IP address, device positioning information, historical activity area, and cross-region activity frequency. The environmental dimension features include network status, device characteristics, system environment, network access method, and security software configuration. The event dimension features include market fluctuations, regulatory policy changes, marketing activities, contemporaneous high-risk events, and internal process changes of financial institutions.
3. The financial pre-account opening risk assessment system based on intelligent identification according to claim 1, characterized in that, The context analysis processing module converts multi-dimensional context features into context feature vectors, which includes: Adaptive normalization processing for continuous features, selecting Z-score or Min-Max normalization method according to feature distribution characteristics; For categorical features, convert them into binary vectors through improved one-hot encoding, and use target encoding technology for high-base features; For text features, convert them into context-aware dense vectors through pre-trained financial domain BERT model; Use feature fusion method based on attention mechanism to dynamically adjust the weight coefficients of each dimension feature according to the current context; Reduce the dimensionality of features through self-encoder to retain key information while reducing computational complexity.
4. The financial pre-account opening risk assessment system based on intelligent identification according to claim 1, characterized in that, In the process of hierarchical risk feature extraction by the context analysis processing module: Account-level features include identity information consistency, account attribute features, historical credit records, biometric verification results, and certificate authenticity scores; Transaction-level features include operation behavior features, input features, time series features, operation habit deviations, and abnormal interruption patterns; Relationship-level features include associated account risks, social network features, group behavior patterns, device association graphs, and cross-institution risk correlation.
5. The financial pre-account opening risk assessment system based on intelligent identification according to claim 1, characterized in that, The knowledge graph updating module extracts structured risk knowledge from multiple sources, including: Obtain risk-related information from internal business feedback data, risk event data, regulatory compliance information, external intelligence data, and industry shared risk library; Use domain-adaptive natural language processing techniques for text preprocessing, including financial professional term identification, cleaning, segmentation, and stop word removal; Use bidirectional long short-term memory network and conditional random field model to identify key entities through named entity recognition; Use a remote supervision multi-instance learning method based on contrastive learning for relation extraction to identify semantic relationships and causal associations between entities; Use multi-head hierarchical attention mechanism for event extraction to extract trigger words, participants, and temporal and spatial background information of risk events; Construct a multi-dimensional attribute system for risk knowledge, including risk level, occurrence probability, impact range, timeliness, and applicable customer groups.
6. The intelligent recognition-based financial pre-account opening risk assessment system according to claim 1, characterized in that, It also includes a knowledge dissemination feedback closed-loop mechanism: Establish a knowledge application effect tracking system to monitor the performance of crowd-sourced knowledge in different business scenarios and customer groups in real time; Use A / B testing method to evaluate the impact of crowd-sourced knowledge on risk assessment accuracy, false positive rate, and false negative rate; Based on the application effect feedback, the knowledge structure, the correlation strength and the application weight are optimized through the reinforcement learning algorithm; A multi-dimensional contributor reputation evaluation system is constructed, including historical contribution quality, knowledge verification rate, professional field coverage, innovation ability and continuous contribution degree. A knowledge co-creation community is established to promote cross-department and cross-professional field risk knowledge exchange and collaborative innovation.
7. A computer readable storage medium characterized in that, The computer readable instructions are stored in the computer readable storage medium, and when the computer readable instructions are read by the computer, the computer readable instructions can run the intelligent identification-based financial pre-account opening risk assessment system according to any one of claims 1-6.
Citation Information
Patent Citations
Risk management method and system for cross-border e-commerce transaction behavior
CN118469715A
Security risk dynamic assessment system and method based on multi-source heterogeneous data analysis
CN118898397A