Commercial bank blacklist management method based on big data platform and Elasticsearch

By integrating commercial bank blacklist data through the big data platform and Elasticsearch, building a real-time retrieval engine and associated risk map, we solved the data silos and real-time processing bottlenecks in commercial bank blacklist management, achieved bank-wide data sharing and millisecond-level risk interception, and improved risk control and security.

CN120634712APending Publication Date: 2025-09-12BANK OF GUIYANG CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510955919.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Commercial banks' blacklist management faces data silos, technical architecture performance bottlenecks, single risk management dimensions, inefficient dynamic updates, and security and privacy protection challenges, resulting in low data integration efficiency, poor real-time processing performance, low risk identification accuracy, and insufficient system scalability and compliance security.

Method used

A big data platform is used to integrate multi-source heterogeneous data, Elasticsearch is used to establish a distributed real-time retrieval engine, and an associated risk map model is constructed to achieve dynamic risk assessment and real-time updates. Combined with the stream processing platform and blockchain technology, data sharing and real-time interception are achieved.

Benefits of technology

It has achieved unified management of bank-wide blacklist data, improved the linkage interception effect between systems, supported millisecond-level risk interception, enhanced the defense capabilities against fraudulent transactions and credit risks, and improved data security and real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634712A_ABST
    Figure CN120634712A_ABST
Patent Text Reader

Abstract

The invention discloses a commercial bank blacklist management method based on a big data platform and Elasticsearch, and the method comprises the steps: building a blacklist data center based on the big data platform, integrating multi-source heterogeneous data, and carrying out the real-time data cleaning and standardization processing; establishing a distributed real-time retrieval engine by utilizing Elasticsearch, and constructing a reverse index for the processed data; establishing an associated risk map model, dynamically analyzing the association degree between the client and the blacklist main body through a map calculation engine, generating a risk conduction coefficient, and storing the risk conduction coefficient; during business handling, a multi-condition combination query request is initiated to Elasticsearch through an ESB (Enterprise Service Bus) real-time calling interface; a dynamic updating mechanism is adopted, and the blacklist state is automatically updated when a risk threshold value is triggered. According to the method, by integrating multi-source heterogeneous data, millisecond risk interception is realized, a dynamic association graph is constructed, the active defense capability of a bank on risks such as fraudulent transactions, credit default and money laundering behaviors is improved, and the method is suitable for risk management and control of core business scenes such as pre-loan auditing, transaction monitoring and anti-money laundering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of financial technology and bank risk management technology, and more specifically to a commercial bank blacklist management method based on a big data platform and Elasticsearch. Background Art

[0002] With the rapid development of internet information technology, financial institutions should enhance their risk control capabilities across all their businesses. Currently, many commercial banks nationwide are experiencing situations where customers and their associated customers who are blacklisted in certain bank systems can still apply for other banking services. While commercial bank blacklist management is a core component of risk prevention and control, existing technologies have significant flaws, primarily in the following five areas:

[0003] 1. Data silos and lack of sharing mechanisms:

[0004] Commercial banks operate a wide range of businesses, involving numerous systems. Each business system has its own blacklist management module, but these systems lack effective data sharing, creating information silos. Traditionally, blacklist data is dispersed across various internal bank systems, such as credit card, loan, and anti-fraud departments, as well as external third-party institutions like credit reporting platforms and judicial databases, lacking a unified integration mechanism. Independent maintenance of lists by each department leads to data fragmentation and weak cross-departmental risk prevention capabilities. For example, a customer may be blacklisted by the loan department, but the transaction system may be unable to intercept their fraudulent transactions due to a lack of real-time access to this information.

[0005] The cost of purchasing external data is high, and centralized sharing models (such as traditional databases) have the risk of single point failure, data is easily tampered with, and there is a lack of mutual trust mechanism among financial institutions.

[0006] 2. Technical architecture performance bottleneck

[0007] When faced with massive amounts of data (tens to hundreds of millions), the query response delay of the list management system based on relational databases can be as high as several seconds (for example, the average response time of a certain bank system is 6.5 seconds), which cannot meet the real-time transaction risk control requirements (millisecond-level interception is required).

[0008] Insufficient complex association analysis capabilities: Traditional technologies are difficult to support multi-dimensional association queries, such as uncovering fraud gangs through device IDs, mobile phone numbers, and social relationship chains. They can only rely on simple rule matching, such as precise ID card matching, and the missed detection rate remains high.

[0009] 3. Single risk management dimension

[0010] Existing systems often rely on static list matching and lack dynamic risk assessment capabilities. For example, they are unable to identify "high-risk contagion targets" with strong ties to blacklisted customers, such as companies with equity control, collateral chains, or close financial transactions. Nor do they incorporate behavioral data, such as unusual transaction frequency or geographic location changes, to adjust risk levels in real time. Rule engines rely on manual experience and are unable to adapt to emerging fraud patterns, such as the "79% return rate" circumvention threshold detection.

[0011] 4. Inefficient dynamic updates and maintenance

[0012] Blacklist updates rely on batch end-of-day operations, which are subject to severe time lags. For example, newly added risky customers take effect 24 hours later. Data cleaning, deduplication, and conflict resolution are highly dependent on manual operations, resulting in high error rates and huge operation and maintenance costs.

[0013] 5. Security and Privacy Protection Challenges

[0014] Centralized storage is vulnerable to attacks, leading to data leaks, and sensitive information (ID number, transaction records) is not effectively desensitized.

[0015] Therefore, the current commercial bank blacklist management has systemic bottlenecks in data integration efficiency, real-time processing performance, risk identification accuracy, system scalability and compliance security. It is urgent to reconstruct the risk control architecture through technologies such as big data distributed computing, real-time search engines, and intelligent graph analysis to achieve the transformation from "passive interception" to "intelligent defense." Summary of the Invention

[0016] In view of this, the present invention provides a commercial bank blacklist management method based on a big data platform and Elasticsearch, which at least partially solves the above technical problems.

[0017] In order to achieve the above object, the present invention adopts the following technical solutions:

[0018] The present invention provides a commercial bank blacklist management method based on a big data platform and Elasticsearch, comprising the following steps:

[0019] S1. Build a blacklist data center based on the big data platform, integrating multi-source heterogeneous data from the bank's internal credit system, external credit reporting agencies, judicial databases, and third-party risk intelligence platforms, and performing real-time data cleansing and standardization using distributed ETL tools;

[0020] S2. Use Elasticsearch to establish a distributed real-time search engine and construct an inverted index for the data processed in step S1. The index fields include at least the customer's ID number, mobile phone number, enterprise unified social credit code, IP address, device fingerprint, and associated person information;

[0021] S3. Build a correlation risk graph model. Use the graph computing engine to dynamically analyze the client's equity relationship, guarantee chain, capital transactions, and social network connections with blacklisted entities. Generate a risk transmission coefficient and store it in the _source metadata field in Elasticsearch.

[0022] S4. Real-time interception during business processing. After the front-end system receives the customer identification information, it initiates a multi-condition combination query request to Elasticsearch through the ESB real-time call interface.

[0023] S5. Adopt a dynamic update mechanism to monitor customer behavior data in real time based on the stream processing platform. When the risk threshold is triggered, the blacklist status is automatically updated and the index is refreshed in real time through the Update By Query API of Elasticsearch.

[0024] Furthermore, the multi-source heterogeneous data in step S1 includes:

[0025] Internal bank data: credit overdue records exceeding the preset period; identification of substandard, suspicious, and loss-making customers;

[0026] External credit reporting agencies: tax anomaly lists, fraudulent accounts flagged by third-party payment institutions;

[0027] Judicial database: records of judicial execution;

[0028] Third-party risk intelligence platform: Provides real-time behavioral data including: abnormal changes in the geographic location of customer login IP addresses and high-frequency exploratory transactions.

[0029] Furthermore, distributed ETL tools are used to perform real-time data cleaning and standardization, including:

[0030] Real-time data extraction steps: Use distributed connectors to simultaneously pull data streams from the bank's internal system, external credit API, and judicial blockchain nodes, and only capture incremental change data;

[0031] Distributed real-time cleaning steps: Utilize streaming statistical models to identify anomalies in real time, dynamically determine thresholds for fields related to transaction amounts and geographic location jumps, and mark or correct values ​​that exceed the range; and use context-aware filling strategies to fill missing values.

[0032] Real-time standardization step: Use the regular expression engine to convert the data format and map non-standard enumeration values ​​to unified codes;

[0033] Distributed processing steps: The data processed by the above steps is processed by the big data platform to generate three core tables, including: the list master table containing the list ID, customer name and ID number; the list details table containing the details field and delimiter storage structure; the association relationship table containing equity, guarantee, group, and natural person association attributes.

[0034] Furthermore, the real-time data extraction step also includes: allocating an independent thread pool to each data source and setting a back-pressure mechanism to automatically limit the flow when the data influx speed exceeds the processing capacity to prevent the cluster from crashing.

[0035] Furthermore, in the real-time data extraction step, only incremental change data is captured, including:

[0036] Call the stored procedure through the distributed scheduling engine, compare the temporary table with the current partition data, identify the newly added / exited / updated records, and record the timestamp.

[0037] Furthermore, the inverted index in step S2 adopts the following optimization strategy:

[0038] Enable doc_values ​​to accelerate aggregate queries for ID card number and mobile phone number fields;

[0039] Use the keyword type for the risk tag field and configure eager_global_ordinals to improve filtering efficiency;

[0040] Add a routing_path routing policy for high-risk customers to force their indexes to be assigned to independent physical shards.

[0041] Furthermore, the construction of the associated risk map model in step S3 includes:

[0042] Node type: natural person, corporate entity, guarantor, fund recipient; the natural person includes spouse and / or actual controller;

[0043] Definition of marginal relationship: equity ratio ≥ 5%, guarantee amount ≥ 30% of loan principal, capital transactions ≥ 10 times in the past 6 months;

[0044] Risk transmission rules: If the directly related node is on the blacklist, the target customer risk score will increase by 2n%; if it is indirectly related, it will increase by n%.

[0045] Furthermore, the multi-condition combination query in step S4 adopts a dynamic weight strategy:

[0046] Comprehensive matching degree = ID card matching × 1.0 + device fingerprint matching × 0.3 + IP risk matching × 0.5 + risk transmission coefficient × 0.7;

[0047] Among them, the basic weight is 1.0 for ID card number matching or 0.8 for mobile phone number matching; the enhanced weight is 0.3 for device fingerprint consistency and 0.5 for IP address and historical fraud behavior overlap; the associated risk weight is risk transmission coefficient × 0.7.

[0048] Furthermore, it also includes setting up a blacklist hierarchical interception mechanism:

[0049] Completely restricted: When criminal offenses or malicious debt evasion are involved, new credit, guarantees, and account openings are prohibited;

[0050] Restricted level: When guarantee compensation or risk contagion of affiliated companies is triggered, only mortgage / pledge business is allowed;

[0051] Prompt level: When there is a slight abnormality in the credit report and the person is on the grey list, manual review is triggered.

[0052] Furthermore, the dynamic update of the risk threshold in step S5 includes:

[0053] Short-term threshold: 10 or more failed transactions per day or 20 or more accounts bound to the same device;

[0054] Long-term threshold: repayment delay for three consecutive months and risk flags from external data sources ≥3 times;

[0055] Dynamic scores are calculated through Elasticsearch's ScriptedMetric Aggregation, and blacklisting is automatically triggered when the score exceeds the threshold of 85.

[0056] Furthermore, it also includes query optimization based on flow control:

[0057] Configure an independent thread pool for the blacklist query interface and limit the maximum number of concurrent connections to ≤1000 per node;

[0058] Enable heap memory back pressure control. When the node memory usage is ≥ 80%, non-real-time query requests are automatically rejected.

[0059] Furthermore, it also includes a blacklist exit mechanism:

[0060] Automatic exit: After the debt is settled and there are no risk events for 24 consecutive months, the system will automatically remove you from the blacklist;

[0061] Manual appeal: Customers submit judicial rulings and proof of repayment through the blockchain evidence storage platform, which triggers a status update after verification by Elasticsearch.

[0062] Furthermore, it also includes security protection mechanisms:

[0063] Deploy IP blacklists and whitelists in the Elasticsearch cluster to prevent untrusted IP addresses from accessing the blacklist index.

[0064] For sensitive fields in the query results, including ID card number and mobile phone number, enable FieldMasking desensitization.

[0065] It can be seen from the above technical solutions that, compared with the prior art, the present invention has the following technical effects:

[0066] The present invention first integrates the blacklist data from multiple upstream systems into the big data platform, and then obtains the bank-wide blacklist data through unified processing on the big data platform, and uniformly manages the bank's blacklist, breaking the information silos formed by each business system managing its own blacklist and failing to achieve effective data sharing.

[0067] By providing a highly available blacklist query service that supports complex queries through an Elasticsearch cluster, downstream systems can query bank-wide blacklist data through real-time ESB calls. This prevents customers blacklisted in a particular bank system and their associated customers from continuing to apply for other bank services, improving the effectiveness of inter-system linkage interception and the bank's risk control capabilities. By integrating multi-source heterogeneous data, achieving millisecond-level risk interception, and building a dynamic correlation graph, this method enhances the bank's proactive defense capabilities against risks such as fraudulent transactions, credit defaults, and money laundering. It is suitable for risk management in core business scenarios such as pre-loan review, transaction monitoring, and anti-money laundering. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0069] Figure 1 This is a flow chart of the commercial bank blacklist management method based on the big data platform and Elasticsearch provided by the present invention.

[0070] Figure 2 This is a diagram of the data processing principle of step S1 provided by the present invention.

[0071] Figure 3 This is a schematic diagram of the Elasticsearch cluster structure provided by the present invention.

[0072] Figure 4 This is a schematic diagram of the association map analysis provided by the present invention. DETAILED DESCRIPTION

[0073] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0074] The embodiment of the present invention discloses a commercial bank blacklist management method based on a big data platform and Elasticsearch, referring to Figure 1 As shown, it includes steps S1 to S5: Data silos are resolved by fusing multi-source heterogeneous data, and the Elasticsearch real-time search engine is used to break through the performance bottleneck of traditional databases, reducing query latency from seconds to milliseconds. In addition, dynamic analysis of associated risk maps increases the risk identification dimension and solves the problem of single static matching. Finally, dynamic updates driven by stream processing can eliminate the time lag caused by traditional batch updates.

[0075] S1. Build a blacklist data center based on the big data platform, integrating multi-source heterogeneous data from the bank's internal credit system, external credit reporting agencies, judicial databases, and third-party risk intelligence platforms, and performing real-time data cleansing and standardization using distributed ETL tools; Figure 2 This is the schematic diagram of the data processing principle of the entire step S1.

[0076] In this step, multi-source heterogeneous data includes:

[0077] Internal bank data: credit delinquency records exceeding 30 days, for example; customers identified as substandard, suspicious, or loss-making;

[0078] External credit reporting agencies: tax anomaly lists, fraudulent accounts flagged by third-party payment institutions;

[0079] Judicial database: records of judicial execution;

[0080] Third-party risk intelligence platform: Provides real-time behavioral data including: abnormal changes in the geographic location of the customer's login IP address (e.g., Δ≥500km / 1h) and high-frequency exploratory transactions (e.g., ≥5 times / minute).

[0081] Taking the data warehouse as an example, for example, the bank's blacklist data is collected by directly connecting to the upstream system database through a read-only user, including 30+ types of blacklist original data such as asset write-offs, credit overdue, telecommunications fraud, anti-money laundering, gambling-related, case-related, and dishonest debtors extracted from retail credit, corporate credit, credit pre-investigation, anti-money laundering, core, and asset preservation systems.

[0082] Among them, real-time data cleaning and standardization are performed through a distributed ETL tool, including:

[0083] 1) Real-time data extraction steps: Streams of data are pulled simultaneously from the bank's internal system, external credit investigation API, judicial blockchain nodes, etc. through distributed connectors (such as Kafka Connect, FlinkKafkaSource), supporting millisecond-level latency; The stored procedure is called through a distributed scheduling engine to compare the data in the temporary table with the current partition data, identify new / withdrawn / updated records, and record the timestamp, so as to achieve only capturing incremental change data and avoiding resource waste caused by full-scan; Then, an independent thread pool is allocated for each data source, and a backpressure mechanism is set up to automatically limit the flow when the data influx speed exceeds the processing capacity, preventing the cluster from crashing.

[0084] For example, the data is first loaded into the database temporary table; Subsequently, the stored procedure is called to compare the data in the temporary table with the data in the current latest partition, and the following three parts of data are obtained: (1) The blacklist newly added on the current day; (2) The blacklist withdrawn on the current day; (3) The blacklist with updated information on the current day; The three parts of data are processed respectively, recording the time when the blacklist enters / withdraws, and updating the corresponding records, etc.

[0085] 2) Distributed real-time cleaning steps: Abnormalities are identified in real-time using streaming statistical models (such as Z-Score or IQR), and dynamic threshold judgments are made on relevant fields such as transaction amount and geographical location jumps, and records are marked or corrected if they exceed the range; In addition, a context-aware filling strategy is adopted:

[0086] Numeric fields: Filled with the moving window mean (such as the average amount of the same type of transaction in the last 10 minutes);

[0087] Categorical fields: Inferred through associated fields (such as filling the missing address with the user's historical address associated with the bank card number).

[0088] Duplicate removal and primary key conflict resolution:

[0089] For data with the same primary key (such as "transaction ID + timestamp"), the latest version is recorded in the distributed state, and duplicate-arriving records are discarded. The context-aware filling strategy is used for missing value filling;

[0090] 3) Real-time standardization steps: The regular expression engine (such as Flink SQL REGEXP_EXTRACT) is used to forcibly convert the data format, and non-standard enumeration values (such as "Male" / "M" / "男") are mapped to a unified code ("M");

[0091] 4) Distributed processing steps: The big data platform processes the data processed by the above steps to generate three core tables, including: the list master table containing the list ID, customer name and ID number; the list details table containing the details field and delimiter storage structure; and the relationship table containing equity, guarantee, group, and natural person association attributes.

[0092] After the big data platform obtains the original data through SFTP, it processes the list master table, list details table and list association table according to the agreed format, distinguishes the specific list type through the list number field, and associates the main table to the list details data through the data ID. At the same time, it associates it to the corresponding association data of the list (including equity association, guarantee association, group association, and natural person association) through the association relationship number.

[0093] The specific processing logic is as follows:

[0094] (1) Processing the master table: All internal lists use the same master table, which has 17 fields, including list ID, list number, customer name, public / private logo, private ID number, organization, unified social credit code, etc.

[0095] (2) Processing details table. The details table has two fields: list ID and list details field. Since the details field of each blacklist is different, multiple values ​​in the details field are processed and stored by separating them with special delimiters. The application will parse the specific list details.

[0096] (3) Processing of association relationships: Currently, association relationships include equity association, group association, guarantee association, and natural person association. The association relationship table has 8 fields, mainly including association relationship ID, association relationship type code, subject name, certificate type, certificate number, relationship attribute, relationship name, etc. The first is equity association: find out the shareholder and investment customer information of all corporate customers of the bank from external industrial and commercial data, credit report, and credit system data, and process the data into the association relationship table; the second is group association: find out which companies are under a certain group from the corporate credit system data; the third is guarantee association: process the association relationship data of guarantors and guaranteed persons from the corporate credit and retail credit data; and the last is natural person association: process the direct association relationship of natural persons from the credit, corporate, and retail data.

[0097] S2. Use Elasticsearch to establish a distributed real-time search engine and construct an inverted index for the data processed in step S1. The index fields include at least the customer's ID number, mobile phone number, enterprise unified social credit code, IP address, device fingerprint, and associated person information; an index compression algorithm can be used for compressed storage.

[0098] Write the full blacklist data into the Elasticsearch cluster, such as Figure 3 As shown in the figure, the inverted index is used to support multi-field retrieval; by calling the metering engine microservice, the full blacklist data for the day can be written to the latest index of the Elasticsearch cluster through the API.

[0099] The inverted index in step S2 adopts the following optimization strategy:

[0100] (1) Enable doc_values ​​to accelerate aggregate queries for ID card number and mobile phone number fields;

[0101] (2) Use keyword type for risk tag fields (such as "gambling-related" and "fraud-related") and configure eager_global_ordinals to improve filtering efficiency;

[0102] (3) Add routing_path routing policy for high-risk customers to force their indexes to be assigned to independent physical shards.

[0103] Elasticsearch: A distributed search engine technology, also known as ES. For example, it can be deployed in a three-machine cluster to ensure high performance, high availability, and data loss prevention for query services. A batch program loads the latest and complete blacklist data into the ES cluster daily. Leveraging its inverted index, it can quickly process complex queries, supporting fuzzy searches, multi-field searches, and word segmentation searches, supporting complex and diverse downstream search requests. Finally, its ability to calculate the relevance score between the query string and the searched field can be leveraged to implement relevance calculations in screening rules.

[0104] S3. Build a correlation risk graph model. Use the graph computing engine to dynamically analyze the client's equity relationship, guarantee chain, capital transactions, and social network connections with blacklisted entities. Generate a risk transmission coefficient and store it in the _source metadata field in Elasticsearch.

[0105] Among them, the construction of the associated risk map model, such as Figure 4 Shown, including:

[0106] Node type: natural person, corporate legal person, guarantor, fund recipient; the natural person includes spouse and / or actual controller;

[0107] Definition of marginal relationship: equity ratio ≥ 5%, guarantee amount ≥ 30% of loan principal, capital transactions ≥ 10 times in the past 6 months;

[0108] Risk transmission rules: If the directly related node is on the blacklist, the target customer risk score will increase by 40%; if it is indirectly related, it will increase by 20%.

[0109] Calculate the risk transmission coefficient RiskScore through the GNN algorithm:

[0110] RiskScore=α·GNN(graph)+β·BehaviorIndex

[0111] Where α and β are adaptive weights, GNN(graph) is the correlation output by the graph neural network, and BehaviorIndex is a dynamic index based on behavioral data. This overcomes the limitations of static rules and enables adaptive risk prediction.

[0112] S4. Real-time interception during business processing: After receiving the customer identification information, the front-end system initiates a multi-condition combination query request to Elasticsearch through the ESB real-time call interface;

[0113] Multi-condition combination query uses dynamic weight strategy:

[0114] Comprehensive matching degree = ID card matching × 1.0 + device fingerprint matching × 0.3 + IP risk matching × 0.5 + risk transmission coefficient × 0.7;

[0115] Among them, the basic weight is 1.0 for ID card number matching or 0.8 for mobile phone number matching; enhanced weight: 0.3 for device fingerprint consistency and 0.5 for IP address and historical fraud behavior overlap; associated risk weight: risk transmission coefficient × 0.7.

[0116] In practice, the query interface supports flexible and differentiated list screening rule configuration, offering over 50 list categories to choose from. Downstream business personnel can independently determine the screening rules for their own systems based on their specific business circumstances. Screening rules primarily include: matching rules (ID number matching, name matching, and correlation requirements), screening categories (which types of lists need to be checked, and whether correlations need to be checked), and screening types (personal, public, and general). Additionally, the application offers features such as manual list inclusion, list querying, and whitelist management.

[0117] The blacklist query service is published through ESB for real-time call by downstream systems. The downstream calls the list query microservice tf-interface through the HTTP protocol. The microservice queries Elasticsearch through the API and determines whether the list is matched based on the configured screening rules and encapsulates the returned results.

[0118] In implementation, the application service uses VUE+Java to develop front-end and back-end components, and uses Spring Boot and Nacos for deployment. It provides functions such as user management, list query, whitelist management, and screening rule configuration. It supports flexible configuration of screening rules for downstream business systems (such as retail credit, corporate credit, and credit cards), primarily by configuring which lists and relationships can be queried by a particular system. Whitelists support time-limited configuration, allowing users to release blacklisted customers within a certain period of time under special circumstances.

[0119] The list query service is a pure Java backend component, deployed using Spring Boot and Nacos. It provides an ESB publishing interface and 24-hour list query service. Downstream systems call this service in real time through the ESB. The service receives and parses request messages, retrieves the screening rule details from Redis, organizes the query statement based on the screening rules and the request message, calls Elasticsearch through the API to query the list, and assembles the message to return the results.

[0120] In one embodiment, the method of the present invention further includes query optimization based on flow control:

[0121] Configure an independent thread pool for the blacklist query interface and limit the maximum number of concurrent connections to ≤1000 per node;

[0122] Enable heap memory back pressure control. When the node memory usage is ≥ 80%, non-real-time query requests are automatically rejected.

[0123] S5. Adopt a dynamic update mechanism to monitor customer behavior data in real time based on the stream processing platform. When the risk threshold is triggered, the blacklist status is automatically updated and the index is refreshed in real time through the Update By Query API of Elasticsearch.

[0124] The dynamic update of risk thresholds includes:

[0125] Short-term threshold: 10 or more failed transactions per day or 20 or more accounts bound to the same device;

[0126] Long-term threshold: repayment delay for three consecutive months and risk flags from external data sources ≥3 times;

[0127] Dynamic scores are calculated through Elasticsearch's ScriptedMetric Aggregation, and blacklisting is automatically triggered when the score exceeds the threshold of 85.

[0128] The risk threshold can be trained through the federated learning model to protect data privacy:

[0129] Threshold=FedAvg(local_models)+λ·AnomalyDetect

[0130] Among them, FedAvg is the federated averaging algorithm, λ is the adjustment factor, and AnomalyDetect is the real-time anomaly detection output.

[0131] In one embodiment, the method of the present invention further includes a blacklist exit mechanism:

[0132] Automatic exit: After the debt is settled and there are no risk events for 24 consecutive months, the system will automatically remove you from the blacklist;

[0133] Manual appeal: Customers submit judicial rulings and proof of repayment through the blockchain evidence storage platform, which triggers a status update after verification by Elasticsearch.

[0134] In one embodiment, the method of the present invention further includes setting a blacklist hierarchical interception mechanism:

[0135] Completely restricted level: When criminal offenses or malicious debt evasion are involved, new credit, guarantees, and account opening are prohibited; restricted level: When guarantee compensation is triggered or risks of affiliated companies spread, only mortgage / pledge business is allowed; prompt level: When there are slight credit abnormalities or the person is on the gray list, manual review is triggered.

[0136] In one embodiment, the method of the present invention further includes a safety protection mechanism:

[0137] Deploy IP blacklists and whitelists in the Elasticsearch cluster to prevent untrusted IP addresses from accessing the blacklist index.

[0138] For sensitive fields in the query results, including ID card number and mobile phone number, enable FieldMasking desensitization.

[0139] This invention integrates multi-source heterogeneous data of blacklists based on a big data platform, builds a real-time retrieval engine based on Elasticsearch, integrates graph computing to mine associated risks, and uses blockchain to ensure the credibility of shared data. It aims to solve the systemic bottlenecks in current commercial bank blacklist management in terms of data integration efficiency, real-time processing performance, risk identification accuracy, system scalability and compliance security. Through technologies such as big data distributed computing, real-time search engines, and intelligent graph analysis, the risk control architecture is reconstructed to achieve the transformation from "passive interception" to "intelligent defense."

[0140] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0141] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A commercial bank blacklist management method based on a big data platform and Elasticsearch, characterized in that: The following steps are involved: S1. Build a blacklist data center based on the big data platform, integrating multi-source heterogeneous data from the bank's internal credit system, external credit reporting agencies, judicial databases, and third-party risk intelligence platforms, and performing real-time data cleansing and standardization using distributed ETL tools; S2. Use Elasticsearch to establish a distributed real-time search engine and construct an inverted index for the data processed in step S1. The index fields include at least the customer's ID number, mobile phone number, enterprise unified social credit code, IP address, device fingerprint, and associated person information; S3. Build a correlation risk graph model. Use the graph computing engine to dynamically analyze the client's equity relationship, guarantee chain, capital transactions, and social network connections with blacklisted entities. Generate a risk transmission coefficient and store it in the _source metadata field in Elasticsearch. S4. Real-time interception during business processing: After receiving the customer identification information, the front-end system initiates a multi-condition combination query request to Elasticsearch through the ESB real-time call interface; S5. Adopt a dynamic update mechanism to monitor customer behavior data in real time based on the stream processing platform. When the risk threshold is triggered, the blacklist status is automatically updated and the index is refreshed in real time through the Update By Query API of Elasticsearch.

2. The method according to claim 1, characterized in that The multi-source heterogeneous data in step S1 includes: Internal bank data: credit overdue records exceeding the preset period; identification of substandard, suspicious, and loss-making customers; External credit reporting agencies: tax anomaly lists, fraudulent accounts flagged by third-party payment institutions; Judicial database: records of judicial execution; Third-party risk intelligence platform: Provides real-time behavioral data including: abnormal changes in the geographic location of customer login IP addresses and high-frequency exploratory transactions.

3. The method according to claim 2, characterized in that Real-time data cleaning and standardization through distributed ETL tools, including: Real-time data extraction steps: Use distributed connectors to simultaneously pull data streams from the bank's internal system, external credit API, and judicial blockchain nodes, and only capture incremental change data; Distributed real-time cleaning steps: Utilize streaming statistical models to identify anomalies in real time, dynamically determine thresholds for fields related to transaction amounts and geographic location jumps, and mark or correct values ​​that exceed the range; and use context-aware filling strategies to fill missing values. Real-time standardization step: Use the regular expression engine to convert the data format and map non-standard enumeration values ​​to unified codes; Distributed processing steps: The data processed by the above steps is processed by the big data platform to generate three core tables, including: the list master table containing the list ID, customer name and ID number; the list details table containing the details field and delimiter storage structure; the association relationship table containing equity, guarantee, group, and natural person association attributes.

4. The method according to claim 1, wherein The inverted index in step S2 adopts the following optimization strategy: Enable doc_values ​​to accelerate aggregate queries for ID card number and mobile phone number fields; Use the keyword type for the risk tag field and configure eager_global_ordinals to improve filtering efficiency; Add a routing_path routing policy for high-risk customers to force their indexes to be assigned to independent physical shards.

5. The method according to claim 1, wherein The construction of the associated risk map model in step S3 includes: Node type: natural person, corporate entity, guarantor, fund recipient; the natural person includes spouse and / or actual controller; Definition of marginal relationship: equity ratio ≥ 5%, guarantee amount ≥ 30% of loan principal, capital transactions ≥ 10 times in the past 6 months; Risk transmission rules: If the directly related node is on the blacklist, the target customer risk score will increase by 2n%; if it is indirectly related, it will increase by n%.

6. The method according to claim 1, characterized in that The multi-condition combination query in step S4 adopts a dynamic weight strategy: Comprehensive matching degree = ID card matching × 1.0 + device fingerprint matching × 0.3 + IP risk matching × 0.5 + risk transmission coefficient × 0.7; Among them, the basic weight is 1.0 for ID card number matching or 0.8 for mobile phone number matching; the enhanced weight is 0.3 for device fingerprint consistency and 0.5 for IP address and historical fraud behavior overlap; the associated risk weight is risk transmission coefficient × 0.

7.

7. The method according to claim 1, characterized in that It also includes setting up a blacklist hierarchical interception mechanism: Completely restricted: When criminal offenses or malicious debt evasion are involved, new credit, guarantees, and account openings are prohibited; Restricted level: When guarantee compensation or risk contagion of affiliated companies is triggered, only mortgage / pledge business is allowed; Prompt level: When there is a slight abnormality in the credit report and the person is on the grey list, manual review is triggered.

8. The method according to claim 1, characterized in that The dynamic update of the risk threshold in step S5 includes: Short-term threshold: 10 or more failed transactions per day or 20 or more accounts bound to the same device; Long-term threshold: repayment delay for three consecutive months and risk flags from external data sources ≥3 times; Dynamic scores are calculated through Elasticsearch's ScriptedMetric Aggregation, and blacklisting is automatically triggered when the score exceeds the threshold of 85.

9. The method according to claim 1, characterized in that Also includes query optimization based on flow control: Configure an independent thread pool for the blacklist query interface and limit the maximum number of concurrent connections to ≤1000 per node; Enable heap memory back pressure control. When the node memory usage is ≥ 80%, non-real-time query requests are automatically rejected.

10. The method according to claim 1, characterized in that Also includes a blacklist exit mechanism: Automatic exit: After the debt is settled and there are no risk events for 24 consecutive months, the system will automatically remove you from the blacklist; Manual appeal: Customers submit judicial rulings and proof of repayment through the blockchain evidence storage platform, which triggers a status update after verification by Elasticsearch.

Citation Information

Cited By

  • Cleaning processing method and system for medicine distribution data, electronic equipment and medium

    CN120872945A