Bank customer portrait knowledge graph construction method and system based on multi-source heterogeneous data fusion

By fusing multi-source heterogeneous data to construct a bank customer profile knowledge graph, and using deep learning and machine learning models to extract entities and relationships, the graph is dynamically updated. This solves the problems of data fusion difficulties and inaccurate entity recognition in bank customer profiling, and enables efficient business decision-making and operational support.

CN120996169APending Publication Date: 2025-11-21JIANGSU SUNING BANK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511055271.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing bank customer profiling technology solutions suffer from problems such as difficulty in data integration, insufficient real-time performance, coarse label granularity, and long label expiration cycles, resulting in low profile accuracy, insufficient coverage, slow real-time response, and high misjudgment rate, making it difficult to achieve precise marketing and risk control.

Method used

A knowledge graph construction method for bank customer profiles is adopted, which integrates multi-source heterogeneous data. By collecting structured and unstructured data, deep learning models and machine learning models are used to extract entities and relationships, and a dynamically updated knowledge graph is constructed. The graph database and hybrid inference engine are then used for evaluation and repair.

Benefits of technology

It enables efficient processing of entities and relationships, improves the accuracy and timeliness of customer profiling, optimizes business decision-making and operations, solves the problems of data fusion difficulties and inaccurate entity identification, and supports precision marketing and risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996169A_ABST
    Figure CN120996169A_ABST
Patent Text Reader

Abstract

The invention relates to the field of financial science and technology, and discloses a bank customer portrait knowledge graph construction method and system based on multi-source heterogeneous data fusion, and the key point of the technical scheme is that the method comprises the following steps: collecting bank customer data; obtaining a hierarchical classification rule of clients and products, constructing a knowledge framework of client portraits, and sorting the collected client data through the knowledge framework; based on the sorted customer data, through an entity extraction rule, extracting entities from the structured data; extracting entities from the unstructured data through a deep learning model; based on the extracted entity, extracting an entity relationship from the client data through an entity relationship extraction rule; meanwhile, multi-dimensional features related to entities are extracted and input into a machine learning model with a confidence coefficient threshold value, and entity relations higher than the confidence coefficient threshold value are output; and according to the entity, the entity relationship and the knowledge framework, constructing a customer portrait knowledge graph, and setting a dynamic updating mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of financial technology, and more specifically, to a method and system for constructing a bank customer profile knowledge graph based on the fusion of multi-source heterogeneous data. Background Technology

[0002] As financial markets develop and digitalization deepens, banking operations are becoming increasingly diversified and complex. Building accurate and comprehensive customer profiles is becoming increasingly crucial for banks to improve their business performance, supporting decisions in areas such as targeted marketing, personalized services, and risk control. However, banks currently face difficulties in data integration and technological bottlenecks when building customer profiles.

[0003] Current mainstream customer profiling technologies mainly include: static rule-based tagging systems, behavior scoring models, and periodic profiling based on CRM data warehouses. These methods suffer from drawbacks such as insufficient data real-time performance, coarse tag granularity, and long tag expiration periods. Quantitative comparisons show that traditional solutions have an accuracy rate below 60%, coverage below 80%, a real-time response time of 72 hours, and a tag richness of less than 200, with a misjudgment rate as high as 35% in retail credit scenarios. Existing technologies suffer from significant profiling distortion due to weak data fusion capabilities and insufficient model dynamism, necessitating the use of multi-dimensional data fusion technologies to improve the accuracy and timeliness of profiling. Summary of the Invention

[0004] The purpose of this invention is to provide a method and system for constructing a bank customer profile knowledge graph based on the fusion of multi-source heterogeneous data. This method and system can achieve more accurate profile construction, better entity and relationship processing, and the constructed knowledge graph can realize knowledge fusion expansion and high-quality maintenance.

[0005] The above-mentioned technical objective of this invention is achieved through the following technical solution: a method for constructing a bank customer profile knowledge graph based on multi-source heterogeneous data fusion, comprising the following steps: S1. Collect bank customer data, including structured and unstructured data. S2. Obtain hierarchical classification rules for customers and products, construct a knowledge framework for customer profiles, and organize the collected customer data through the knowledge framework; S3. Based on the sorted customer data, entities are extracted from structured data using entity extraction rules; entities are also extracted from unstructured data using a deep learning model. S4. Based on the extracted entities, entity relationships are extracted from customer data through entity relationship extraction rules; at the same time, multi-dimensional features related to the entities are extracted and input into a machine learning model with a set confidence threshold, and the output is entity relationships that are higher than the confidence threshold. S5. Based on entities, entity relationships, and knowledge framework, construct a customer profile knowledge graph and establish a dynamic update mechanism for the customer profile knowledge graph.

[0006] As a preferred technical solution of the present invention, in S1, after the customer data is collected, encryption and sensitive information masking are performed. Customer data can be collected in three ways: scheduled collection, real-time collection, and incremental collection.

[0007] As a preferred technical solution of the present invention, in S2, the hierarchical classification rules for customers and products include: multi-level customer classification and multi-level product classification; by connecting the categories of multi-level customer classification and multi-level product classification, a knowledge framework for customer profiles is constructed.

[0008] As a preferred embodiment of the present invention, in S3, the method for extracting entities from unstructured data is as follows: Word embedding technology is used to convert words in the text into low-dimensional vectors. The low-dimensional vector is then input into a deep learning model using a BiLSTM+CRF architecture. The entity category label assignment is determined by maximizing the model probability, thus completing entity recognition.

[0009] As a preferred technical solution of the present invention, in S3, after the entity is extracted, knowledge disambiguation is performed, and entity links are established with entities in the knowledge graph or entity library. For entities that fail knowledge disambiguation, they are pushed to the manual review module. For entity links that conflict, they are detected according to preset detection rules and statistical analysis methods, and conflict handling is performed based on the detection results.

[0010] As a preferred technical solution of the present invention, in S4, the multidimensional features related to entities include: word vector features, part-of-speech features, syntactic structure features, and entity distance features.

[0011] The machine learning model is also used to output entity relationships below the confidence threshold to the human review module, and to update the machine learning model with entity relationships that have been approved by the human review module.

[0012] As a preferred technical solution of the present invention, in S5, when constructing the knowledge graph, the entities are also optimized through ontology mapping technology and entity alignment technology.

[0013] As a preferred embodiment of the present invention, after S5, S6 is also executed, the content of which is: Based on a customer profile knowledge graph with a dynamic update mechanism, it is stored in a graph database; After storage, customer importance is assessed by analyzing the customer profile knowledge graph; customer business risks and recommendation services are assessed through a hybrid inference engine; and the customer profile knowledge graph is quality checked and repaired through contradictory inference.

[0014] As a preferred technical solution of the present invention, in the process of model application in S3 and S4, the model is evaluated by accuracy, recall and F1 score, and the model is optimized based on the evaluation results.

[0015] A knowledge graph construction system for bank customer profiling based on multi-source heterogeneous data fusion, comprising: The data acquisition module is used to collect bank customer data, including both structured and unstructured data. The knowledge framework building module is used to obtain hierarchical classification rules for customers and products, build a knowledge framework for customer profiles, and organize the collected customer data through the knowledge framework. The entity extraction module, based on the sorted customer data, extracts entities from structured data using entity extraction rules; and extracts entities from unstructured data using a deep learning model. The entity relationship extraction module extracts entity relationships from customer data based on the extracted entities and according to entity relationship extraction rules. At the same time, it extracts multi-dimensional features related to the entities and inputs them into a machine learning model with a set confidence threshold, outputting entity relationships that are higher than the confidence threshold. The knowledge graph building module is used to construct a customer profile knowledge graph based on entities, entity relationships, and knowledge frameworks, and to establish a dynamic update mechanism for the customer profile knowledge graph.

[0016] In summary, the present invention has the following beneficial effects: This invention enables more accurate and efficient entity and relationship processing, achieves knowledge fusion and dynamic graph optimization, and continuously improves profile quality to optimize business decisions, solving the problem of data fusion difficulties for banks when building customer profiles.

[0017] This invention features different data update modes, each with its own advantages, enabling timely updates to the map. Regarding version consistency, it demonstrates excellent entity alignment and map merging performance, supports version rollback, and facilitates subsequent system operation. It solves the problems of traditional methods for eliminating data security risks and multi-source data conflicts.

[0018] This invention improves entity recognition accuracy by utilizing deep learning models and related processing techniques to more accurately identify entities in unstructured text, effectively enriching and ensuring the accuracy of knowledge graph construction. Simultaneously, it enhances the efficiency of unstructured text data processing through a rational workflow. Overall, it makes the process of extracting entities from various types of data and integrating them into the knowledge graph smoother and more efficient, laying a solid data foundation for subsequent business applications such as knowledge graph-based analysis and decision-making. It solves the problems of inaccurate entity recognition, confusion due to homonyms, and low efficiency in traditional methods.

[0019] This invention utilizes the PageRank algorithm to accurately assess customer importance, enabling precise allocation of marketing resources and improved business conversion rates. Leveraging a hybrid inference engine, it enhances inference efficiency and quality, handling complex business reasoning more accurately and efficiently, providing reliable data for credit risk assessment and product recommendations. Simultaneously, a contradiction-correction mechanism ensures the consistency of the knowledge graph, preventing erroneous reasoning from impacting business applications and maintaining the overall quality of the graph, thereby comprehensively optimizing the bank's business decisions and operations. This invention solves the problems of traditional methods' inability to comprehensively and objectively measure customer importance and the insufficient accuracy of single inference methods. Attached Figure Description

[0020] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is the complete flowchart of the present invention. Detailed Implementation

[0021] The present invention will be further described in detail below with reference to the accompanying drawings.

[0022] This invention provides a method for constructing a bank customer profile knowledge graph based on multi-source heterogeneous data fusion, comprising the following steps: S1. Collect bank customer data, including structured and unstructured data. Specifically, customer data is collected from multiple channels, both internal and external to the bank. These channels include: internally, using ETL tools to connect to the bank's core business systems according to protocols such as JDBC / ODBC, and extracting basic customer information, credit information, and transaction details according to preset data collection rules; externally, using RESTful APIs to obtain relevant data from social media and e-commerce platforms, while ensuring data format verification and other safeguards. Customer data includes: basic customer information, credit records, and transaction details obtained from the bank's internal system, as well as customer-related data collected from external channels; basic customer information includes name, gender, date of birth, ID number, and contact information; credit records include loan amount, loan term, repayment status, and overdue records; and transaction details include transaction time, transaction amount, transaction type, and counterparty information.

[0023] After collecting customer data, the data undergoes encryption and sensitive information masking. In the sensitive information encryption and desensitization process, the AES-256 symmetric encryption algorithm is used to encrypt sensitive fields. Secure key management is used to store and distribute keys, and data is processed using CBC mode, among other methods. Field masking involves replacing key parts of different sensitive fields according to rules, with strict desensitization rule management and verification mechanisms.

[0024] The purpose of encryption and masking is to protect customer privacy data, prevent data leaks and unauthorized access, and reduce security risks during data storage, transmission, and processing. On one hand, it prevents unauthorized personnel from accessing sensitive information; even internal management personnel cannot view genuine customer privacy data without the appropriate permissions. On the other hand, it can be used to counter external attacks; for example, when hackers attack a database, encrypted and masked data is difficult to directly interpret and exploit. Furthermore, this complies with relevant laws and regulations regarding personal information protection, helping banks fulfill their data security protection obligations.

[0025] Data Collection Authorization and Legal Compliance: According to the "Personal Information Protection Law of the People's Republic of China," banks must adhere to the principles of legality, legitimacy, and necessity when collecting customer privacy data. Excessive collection is prohibited, and the data must be limited to the minimum necessary to achieve the processing objective. Therefore, in this invention, when a customer conducts business, the bank clearly informs the customer of the purpose, method, scope, and retention period of data collection through relevant agreements, authorization letters, etc., and obtains the customer's authorization and consent.

[0026] When collecting customer data from external channels, banks must ensure the source is legal and compliant. Contracts must be signed with third parties to clearly define the rights and obligations of the data provider, guaranteeing that they have obtained customer authorization or are collecting publicly available information only within the legally permissible scope. Furthermore, for sensitive personal information, such as ID card numbers and financial account information, banks must, in accordance with legal regulations, inform customers of the necessity of processing and its impact on their personal rights, and obtain their separate consent in accordance with the law.

[0027] Examples of data related to social media platforms include: user registration information, location, gender, age, etc. AI facial recognition technology can be used to analyze user age and gender, and also to obtain user registration information and location. This includes user-posted videos and images, as well as records of actions such as comments, likes, and shares on other content. Social media platforms such as Xiaohongshu, Douyin, Kuaishou, Bilibili, and Weibo can all collect this behavioral data, which this invention can use to understand customer interests and consumption tendencies. Data such as following lists, number of followers, and friend relationships can be collected to capture follower data of relevant industry bloggers and analyze customer social circles and influence. Keywords and hashtags in posted content can also be analyzed. By analyzing this data, we can understand the hot topics users follow, determine their areas of interest and consumption trends. For example, by analyzing hashtags on Weibo, we can understand customers' attention to and demand for financial products.

[0028] Customer data is collected through scheduled collection, real-time collection, and incremental collection, featuring multiple data update modes. Scheduled collection involves daily batch updates performed in the early morning using Spark SQL for full synchronization, with checkpointing for breakpoint resumption and data partitioning for parallel processing. Real-time collection utilizes a Kafka message queue to receive real-time data, and the Flink stream processing engine parses and triggers incremental graph updates in real time, with strict monitoring to ensure stability. Incremental collection relies on timestamp comparison combined with MVCC to ensure data consistency under concurrency, dynamically monitoring the update volume. S2. Obtain hierarchical classification rules for customers and products, construct a knowledge framework for customer profiles, and organize the collected customer data through the knowledge framework; that is, construct a hierarchical classification system for customers and products, clarify the semantic relationships between concepts to form a knowledge framework.

[0029] The hierarchical classification rules for customers and products include: multi-level customer classification and multi-level product classification; by connecting the categories of multi-level customer classification and multi-level product classification, a knowledge framework for customer profiles is constructed.

[0030] Specifically, the customer concept is subdivided into individual customers and corporate customers. Individual customers are further subdivided according to age, income level, and occupation, while corporate customers are subdivided according to company size and industry. The product concept is divided into deposit products, loan products, wealth management products, and credit card products according to business type, and each subcategory is further subdivided according to specific product characteristics.

[0031] In customer segmentation, for individual customers, after collecting relevant data from multiple channels, they are further segmented by age (divided according to set intervals), income level (distinguished by established ranges), and occupation (classified according to standards) to comprehensively depict their characteristics. For corporate customers, the size is determined based on the company's total assets and number of employees, and the industry is classified according to national industry classification standards. When segmenting product concepts, financial products are first divided into several major categories based on core business attributes: deposits, loans, wealth management, and credit cards. Then, they are further subdivided based on their respective characteristics. For example, deposit products are segmented by term, interest rate type, and withdrawal method; loan products by purpose, repayment method, and term; wealth management by investment targets, risk level, and return cycle; and credit cards by credit limit, annual fee, and preferential benefits.

[0032] In the process of clarifying semantic relationships and constructing a knowledge framework, the inherent logic and business connections of each sub-concept are analyzed, and relationships such as purchase and holding are sorted out. Knowledge graph technology (such as RDF triples) is used to integrate the relationship between customers and products and corresponding concepts, and a hierarchical and logically clear knowledge framework is constructed. This provides a structured knowledge foundation for subsequent work, making the originally complex and scattered information more organized, facilitating various business operations and analysis, helping to accurately build customer profiles and promote cross-departmental business collaboration within the bank.

[0033] S2 clarifies knowledge structures, organizing previously scattered and disorganized customer and product information into a clear and structured framework for easier viewing and automated processing. It also powerfully supports more accurate customer profiling within banks, comprehensively showcasing customer financial behavior characteristics through detailed classification dimensions, thereby improving the quality of business decisions. Furthermore, it optimizes internal business collaboration within banks, breaking down information barriers between departments and facilitating efficient cross-departmental collaboration across multiple business processes, thus improving overall operational efficiency. It resolves the past problems of coarse customer profiling, vague product understanding, and obstacles to business collaboration.

[0034] S3. After the customer data is sorted out, entities are extracted from structured data through preset entity extraction rules; entities are extracted from unstructured data through deep learning models; and after the entities are extracted, knowledge disambiguation is performed to establish entity links with entities in knowledge graphs or entity databases. Specifically, regarding structured data, entities are extracted based on carefully formulated business rules. For example, in various bank data tables, the "name" field is clearly defined as the "customer name" entity, the "account number" field as the "bank account number" entity, etc., to ensure that the extracted entities conform to the banking business logic and meet the needs of subsequent analysis.

[0035] Specifically, the entity extraction method in unstructured data is as follows: First, the text data is vectorized. Word embedding techniques (such as Word2Vec and GloVe) are used to transform the words in the text into low-dimensional vectors. This is achieved through training on a large-scale corpus, ensuring that semantically similar words have close vector distances, thus converting text words into low-dimensional vectors. The low-dimensional vector sequence is then input into a deep learning model employing a BiLSTM+CRF architecture. BiLSTM consists of forward and backward LSTMs. The forward pass captures positive semantic dependencies in the text order, while the backward pass captures negative associations, comprehensively understanding sentence semantics. Each time step outputs a sequence of hidden state vectors rich in semantic information. The CRF layer receives the hidden state vector sequence output by the BiLSTM. Based on Markov random field theory, it considers the global constraints of the entire label sequence and optimizes the local prediction results output by the BiLSTM. Specifically, it defines a transition probability matrix... and an emission probability matrix To model this. Transition probability matrix. It describes the transition probabilities between different entity category labels, such as the probability of transitioning from the "customer name" entity label to the "bank account number" entity label; emission probability matrix. This represents the probability of outputting a certain entity category label in a given hidden state, which is the probability that each hidden state vector corresponds to a specific entity category label; entity category label assignment is determined by maximizing the model probability, thus completing entity recognition.

[0036] The probability calculation formula for a deep learning model using the BiLSTM+CRF architecture is as follows: ,in For entity category label sequence, This is a sequence of vectorized text data. The transition probability matrix, Here is the emission probability matrix. The sequence length is given.

[0037] Regarding "determining entity category label assignment by maximizing model probabilities to complete entity recognition," it is necessary to clarify the following: The core logic of BiLSTM+CRF is to jointly learn "sequence semantic features" and "label transfer patterns": The BiLSTM part is responsible for extracting semantic features from the text sequence, transforming each word into a vector representation containing contextual information, and capturing the semantic relationship between the preceding and following text of an entity. For example, when "credit card" appears near "repayment", it is more likely to be a financial entity.

[0038] CRF section: Based on the BiLSTM output, label transition probabilities are introduced. For example, it is more reasonable to follow "address" with "telephone". The corresponding label transition matrix constrains the rationality of the label sequence and avoids label combinations that do not conform to business logic, such as "person's name" directly followed by "amount".

[0039] "Maximizing model probability" means that the model will traverse all possible entity category label sequences, such as possible label combinations for a text: [person name, organization name, empty, financial product name, etc.]. By calculating the joint probability of "semantic features + label transfer" through CRF, the model selects the label sequence with the highest probability to achieve a reasonable allocation of entity category labels.

[0040] The meaning of the model output probability: In the formula It is "the conditional probability of entity category label sequence I given a text data vector sequence o".

[0041] The model calculates "how likely this text is to conform to semantic logic and label transfer rules when labeled with the current label sequence".

[0042] Specifically, for a single entity, a word or phrase in the text: each position in the label sequence corresponds to an entity label, and the overall probability is the joint probability of all labels in the sequence and the transitions between labels (not the independent probability of a single entity label, but the probability contribution of a single label can be analyzed).

[0043] Logic for determining entity tags: One entity corresponds to one label: Entity recognition is essentially "sequence labeling". Each segment to be identified in the text belongs to an entity category. For example, "Zhang San" is a [person's name] and "XX Bank" is an [institution's name].

[0044] Select the label of the sequence with the highest probability: The model calculates the probability of all possible label sequences using CRF, selects the complete sequence with the highest probability, and the label at each position in the sequence is used as the category label of the corresponding entity.

[0045] During training, the model learns "what text features correspond to what labels and how to transition between labels more reasonably"; during prediction, it uses the learned patterns to calculate the probability of all possible sequences and selects the optimal solution.

[0046] After entity recognition, the similarity between entities is measured using methods such as cosine similarity. Knowledge disambiguation is then performed using knowledge graphs and domain knowledge bases to eliminate ambiguities such as entities with the same name but different identities. Next, the processed entities are linked to their corresponding entities in the knowledge graph or entity database to establish associations. If conflicts arise in entity links, they are detected using preset detection rules and statistical analysis methods, such as comparing key attributes. If disambiguation fails, the process reverts to manual review, and can be limited to a preset timeframe for processing.

[0047] Specifically, the preset detection rules are set around three core logic categories: attribute consistency, business logic, and data credibility. Strong constraints on key attributes: If unique identifiers such as ID card number and mobile phone number are inconsistent, a conflict is determined. Business logic validation: Conflicts are triggered when the loan amount exceeds the credit rating or when transaction behavior violates the laws of time and space. Source credibility weighting: Data from internal core systems takes precedence over third-party channels, and in case of conflict, the most credible source is adopted first.

[0048] The detection method achieves automated detection through attribute comparison, time series analysis, and correlation verification. Attribute comparison: Verify multi-source data field by field, such as name, gender, etc., and mark discrepancies as abnormal; Time series analysis: Tracking time-related data such as transactions and credit to identify abnormal fluctuations in frequency / amount; Association verification: Construct an entity relationship diagram to identify unreasonable associations, such as customers being strongly bound to unrelated products or conflicting kinship relationships.

[0049] Results processing follows the principles of hierarchical processing, human-machine collaboration, and closed-loop management. Tiered processing: Resolve conflicts of critical information first, and address secondary information later; Human-machine collaboration: Complex conflicts are pushed to the manual review module, where personnel combine business experience and multi-source evidence to make judgments and corrections; Closed-loop management: Record the causes of conflicts and the handling process, feed back into rule optimization and quality assessment, and continuously iterate the detection logic.

[0050] The entire S3 process improves entity recognition accuracy by utilizing deep learning models and related processing techniques to more accurately identify entities in unstructured text, effectively enriching and ensuring the accuracy of knowledge graph construction. Simultaneously, a streamlined processing workflow enhances the efficiency of unstructured text data processing, making the extraction of entities from various data types and their integration into the knowledge graph smoother and more efficient. This lays a solid data foundation for subsequent business applications such as knowledge graph-based analysis and decision-making. It solves the problems of inaccurate entity recognition, name confusion and ambiguity, and low efficiency inherent in traditional methods.

[0051] S4. Based on the extracted entities, entity relationships are extracted from customer data through entity relationship extraction rules; at the same time, multi-dimensional features related to the entities are extracted and input into a machine learning model with a set confidence threshold, and the output is entity relationships that are higher than the confidence threshold.

[0052] Specifically, entity relationships are extracted using rules: The entity relationship extraction rules are rules formulated based on banking business logic and domain knowledge. For example, entity relationships are determined from structured data such as transaction records and credit records by field association, such as determining the "guarantee" relationship between customers based on account association.

[0053] Specifically, the machine learning model is one of the following: Support Vector Machine (SVM), Convolutional Neural Network (CNN), or Bidirectional Long Short-Term Memory Network (BiLSTM).

[0054] The process of extracting entity relationships using machine learning models is as follows: Construct feature vectors that integrate word vector features, part-of-speech features, syntactic structure features, and entity distance features.

[0055] More specifically, word vector features originate from word embedding processing of unstructured text data containing entities; part-of-speech features are obtained through part-of-speech tagging analysis of the text; syntactic structure features are obtained by parsing the text using syntactic analysis methods; and entity distance features are quantified based on the distance between the positions of entities in the text. These features are combined to assist in extracting entity relationships.

[0056] After obtaining the feature vectors, one of the following models is selected for training: Support Vector Machine (SVM), Convolutional Neural Network (CNN), or Bidirectional Long Short-Term Memory (BiLSTM). SVM distinguishes relation categories by finding the optimal hyperplane; CNN extracts features and outputs predictions using convolutional and pooling layers; BiLSTM captures semantic judgments of relations in both directions.

[0057] When the model outputs, a confidence threshold is preset. Results above the threshold are reliable and can be directly adopted, while low-confidence results below the threshold are sent to a manual review module for verification by bank staff based on their business knowledge and experience. Afterward, the correct results confirmed by the manual review are fed back into the training process, initiating a semi-supervised learning mechanism to continuously optimize the model.

[0058] By leveraging S4, combining rules and machine learning to extract entity relationships, the accuracy of extraction is improved, the introduction of erroneous relationships is reduced, and the quality of entity relationships in the knowledge graph is enhanced. Setting confidence thresholds, along with manual review and semi-supervised learning mechanisms, optimizes the collaborative process between manual and automated processes, improving overall work efficiency. Simultaneously, it contributes to improving the quality of the knowledge graph, expanding its application in multiple banking business scenarios, and providing strong data support for accurate decision-making. This solves the problems of relying solely on a single method to extract entity relationships, the difficulty in guaranteeing machine learning model results, and the challenge of continuous model optimization.

[0059] S5. Based on entities, entity relationships, and knowledge framework, construct a customer profile knowledge graph and establish a dynamic update mechanism for the customer profile knowledge graph.

[0060] Entities: Entities derived from the S3 steps after disambiguation and non-conflict filtering, including various entity contents representing concepts such as customers and products extracted from structured and unstructured data.

[0061] Entity Relationships: The entity relationships extracted and filtered in step S4 clarify the interrelationships between entities in the business scenario, such as the specific business relationship between customers and products.

[0062] Knowledge Framework: The knowledge framework of the customer profile built in step S2 provides an overall conceptual hierarchy and semantic logic guide for building a knowledge graph, clarifying the business scope and structural patterns that entities and relationships should follow.

[0063] Building a knowledge graph: Using input entities as nodes and entity relationships as edges, and based on the logical and semantic rules defined by the knowledge framework, knowledge graph construction technology is used to integrate these elements to form a networked customer profile knowledge graph structure. This structure intuitively presents the complex relationships between various elements such as customers and products, making it a knowledge carrier that reflects the overall picture of the bank's customer business.

[0064] When constructing the knowledge graph, entity optimization is also performed using ontology mapping and entity alignment techniques.

[0065] Ontology mapping determines the mapping relationship between cross-domain concepts through terminology similarity and structural similarity methods, while entity alignment merges entity attributes by comparing key information such as mobile phone numbers and ID card numbers to achieve entity uniqueness.

[0066] Specifically, in terms of ontology mapping, cross-domain concept mapping relationships are first determined through term similarity calculation. Methods such as cosine similarity and edit distance are used to measure the degree of term similarity, such as analyzing the association between terms like "fixed-term deposit" and "fixed-term wealth management". At the same time, structural similarity analysis is performed to construct a tree structure to compare the differences in the knowledge structure of concepts in different domains. The mapping relationship is determined by combining the results of both methods, so as to break down knowledge barriers and integrate knowledge from multiple domains.

[0067] Entity alignment relies on comparing key information such as mobile phone numbers and ID card numbers. After precise matching extracted from various data sources, the two are determined to be the same entity if they are completely identical. Subsequently, their attributes are merged, such as integrating customer transaction and risk preference information from different systems, to eliminate entity redundancy and ensure entity uniqueness and data accuracy.

[0068] Dynamic Update Mechanism: By establishing a series of data monitoring mechanisms, such as monitoring data source update timestamps and tracking changes in key business data, combined with preset update trigger conditions and corresponding update rules, a mechanism is built that allows the knowledge graph to automatically update in real-time or periodically as business develops and data changes, ensuring it always stays aligned with the latest business situation. The dynamic update mechanism first establishes data change monitoring methods: for structured data, it monitors update timestamps and change logs; for unstructured data, it uses text mining techniques to determine information changes. Upon detecting changes, updates are triggered according to preset rules, adding or modifying entities and relationships in the knowledge graph. Adhering to the consistency principle, the updated graph is ensured to be reasonable and accurate, enabling it to reflect changes in multiple types of data in real time and enhancing the timeliness of the knowledge graph.

[0069] A customer profile knowledge graph with a dynamic update mechanism can accurately reflect current customer business relationships and has the ability to automatically adjust according to subsequent business changes, providing a basic data structure for subsequent application, evaluation and optimization operations.

[0070] Therefore, by leveraging S5's ontology mapping, cross-domain knowledge fusion and expansion are achieved, allowing the knowledge graph to encompass richer and more comprehensive information, thus improving the accuracy of analysis. Entity alignment ensures entity uniqueness, reducing data redundancy and inconsistency, and providing a high-quality data foundation for subsequent applications. The dynamic update mechanism of the knowledge graph enhances its timeliness, enabling it to adapt to changes in business data in real time and better serve various business decisions within the bank. This solves the problems of difficulty in cross-domain knowledge integration, entity redundancy and inconsistency, and lagging knowledge graph updates.

[0071] S6. Based on a customer profile knowledge graph with a dynamic update mechanism, it is stored in a graph database; after storage, the importance of customers is assessed by analyzing the customer profile knowledge graph; customer business risks and recommendation business are assessed through a hybrid inference engine; and the customer profile knowledge graph is quality checked and repaired through contradictory reasoning.

[0072] Specifically, for storage: choose a graph database (such as Neo4j) to store the knowledge graph. Nodes represent entities such as customers and financial products, edges represent the relationships between entities, and each node and edge also includes corresponding attribute data such as basic customer information and product attributes. This contains rich information about customer and product entities and their relationships, and can be dynamically updated to adapt to business changes. Using its specific query language, such as Cypher, complex entity relationships in the graph can be easily queried; for example, quickly retrieve all product information associated with a specific customer, providing an efficient data organization format for subsequent operations.

[0073] Assessing Importance: By analyzing factors such as the customer node's entry and exit from the chain and its correlation with other nodes, and combining these with parameters such as the set damping coefficient, the influence score of the customer in the entire financial business network is calculated iteratively. This measures the relative importance of the customer and provides a reference for business decisions.

[0074] Specifically, the PageRank algorithm is used to assess customer importance, which is applied to customer nodes in the graph. Its core principle is to measure influence based on node link relationships, using a damping coefficient, typically set to 0.85, to reflect the probability of a user randomly jumping to another node. The specific calculation follows an iterative formula. The process involves initializing the PageRank value of each node and then iteratively updating it based on factors such as the number of incoming and outgoing chains until convergence. In the graph, if a customer node is pointed to by many important customer nodes or closely connected to high-value product nodes, its PageRank value will be high according to the algorithm, indicating that the customer is highly important in the financial business network and deserves close attention.

[0075] Hybrid Reasoning: Utilizing a hybrid reasoning engine, based on the entity information such as customers and products stored in the knowledge graph and the relationships between them, and following pre-set business rules and reasoning logic learned through graph neural networks on the graph structure and node features, the engine assesses the customer's business risks, recommends suitable businesses to the customer, uncovers the customer's potential business needs and risk characteristics, and assists the bank in carrying out precise business operations.

[0076] On the one hand, inference rules are constructed and presented in the form of triplets of "condition set, relational operator, conclusion," such as "({credit score < 600, number of large transactions in the past 3 months > 5}, AND, credit risk level = high)." Based on factors such as customer credit, transaction behavior, and product preferences, these rules are used to determine credit risk, recommend products, and perform other business-related tasks. On the other hand, a hybrid inference engine combines rule-based inference with graph neural networks. Rule-based inference matches graph nodes and attributes according to set rules to draw conclusions and update graph information; graph neural networks utilize graph structure and node features to capture complex nonlinear relationships, helping to improve inference accuracy. For example, when inferring a customer's interest in a new product, the system integrates information from surrounding nodes to refine the result.

[0077] Conflict-based reasoning: Through a conflict-based reasoning mechanism, all nodes, edges, and related attribute information in the knowledge graph are traversed. Logical judgment, business rule verification, and data consistency checks are used to identify logical contradictions and data inconsistencies. Once a contradiction is found, source analysis is performed promptly to determine the cause, and corresponding remedial measures are taken. This ensures the logical rigor and data accuracy of the knowledge graph while optimizing its overall quality, enabling it to better serve subsequent business applications.

[0078] Specifically, a contradiction inference correction process is in place to address potential inconsistencies in hybrid inference. Once a contradiction is detected, such as differing assessments of customer risk levels, the source is immediately traced and analyzed, including rules, node attributes, and the intermediate inference process, to determine the cause of the contradiction. Subsequently, if rule weights are unreasonable, they are adjusted; if data issues exist, the relevant data is verified and corrected. For contradictory results related to critical business operations, the results are pushed to a manual review module, requiring correction within a preset timeframe. This ensures the consistency of the knowledge graph and the accuracy of the inference results, guaranteeing uninterrupted business applications.

[0079] After application optimization and its own logical optimization, the customer profile knowledge graph has been improved in terms of storage management, customer importance representation, business risk assessment and recommendation, and its own data quality. It can more effectively provide strong support for various business decisions and operations of the bank, and can continue to serve as the basic data structure for further optimization and application.

[0080] S6 leverages the PageRank algorithm to accurately assess customer importance, helping banks focus on key customers, achieve precise allocation of marketing resources, and improve business conversion rates. Its hybrid inference engine enhances inference efficiency and quality, enabling more accurate and efficient handling of complex business reasoning, providing reliable data for credit risk assessment and product recommendations. Simultaneously, a contradiction-correction mechanism ensures the consistency of the knowledge graph, preventing erroneous reasoning from impacting business applications and maintaining the overall quality of the graph, thereby comprehensively optimizing the bank's business decisions and operations. This addresses the limitations of traditional methods in comprehensively and objectively measuring customer importance and the insufficient accuracy of single inference methods.

[0081] As an optimization scheme of the present invention, in the process of model application in S3 and S4, the model is evaluated by accuracy, recall and F1 score, and the model is optimized according to the evaluation results.

[0082] Step S3 includes the process and results of entity recognition and extraction, such as the entity extraction rules used, the specific operations for entity extraction from structured and unstructured data, and the final output entity information after disambiguation and non-conflict filtering. These reflect the overall performance of the entity recognition process.

[0083] Step S4 includes the process and results of entity relation extraction and filtering, such as the preset entity relation extraction rules, the specific circumstances of inputting multidimensional features into the machine learning model, and the entity relation content output by the model that is higher than the confidence threshold, which reflects the actual effectiveness of the entity relation extraction stage.

[0084] The optimization process involves: conducting multi-dimensional performance evaluation of the designed model, including: comprehensively evaluating the entity recognition and extraction in S3 and the entity relationship extraction and screening in S4 using three key indicators: accuracy, recall, and F1 score.

[0085] Accuracy measures the proportion of correct results to predicted results in entity recognition and relation extraction. It reflects the model's ability to avoid misjudgments by calculating the ratio of correctly identified or extracted entities to incorrectly identified or extracted entities; the expression is: ,in To correctly identify or extract the quantity, The number of incorrectly identified or extracted items.

[0086] Recall measures the proportion of correct results to all results that should have been identified or extracted; it is calculated as the ratio of correctly identified or extracted results to the actual number of unidentified or extracted results, reflecting the degree to which the model covers true information. The expression is: ;in The actual number of unidentified or unsampled entities.

[0087] The F1 score is the harmonic mean of precision and recall, which comprehensively evaluates the overall performance of a model on these two tasks, avoiding the bias of a single metric. It is expressed as: , Based on the results of the above evaluation indicators, we will conduct an in-depth analysis of the problems existing in the model. If the accuracy is low, it may mean that the model structure is not reasonable enough, the training data contains noise, or the feature extraction is insufficient. To address these possible causes, we will take corresponding optimization measures, such as adjusting the model parameter settings, adding high-quality training data to expand the training samples, and improving feature engineering methods to enhance the effectiveness of features. We will also conduct targeted optimization of the relevant models used for entity recognition and relationship extraction to improve their performance in subsequent tasks, thereby improving the accuracy and reliability of the entire customer profile knowledge graph construction method.

[0088] After the above processing, we can obtain more optimized models for entity recognition and relation extraction. When these optimized models are applied again to steps S3 and S4 or similar entity recognition and relation extraction tasks, they can output more accurate and comprehensive results, providing strong model support for continuously improving the quality of customer profile knowledge graphs and the effectiveness of bank customer business-related applications.

[0089] In addition to optimizing S3 and S4, the resulting knowledge graph is also evaluated. The overall performance evaluation of the knowledge graph relies on coverage, completeness, and consistency metrics. Coverage is determined by the ratio of the number of entities to relations. The theoretically required number of entities and relations is compared with the actual number contained in the graph to determine the degree of coverage; a higher value indicates a more complete and comprehensive reflection of the business picture. Completeness is measured by the entity attribute filling ratio, which is the ratio of the number of entities with valid attributes to the total number of entities, thus judging the richness of entity information. Consistency focuses on checking for contradictions in the graph, such as conflicting attribute records for the same entity in different places. By traversing each element of the graph, inconsistencies are identified and marked according to logic and business rules to ensure data logical consistency.

[0090] This optimization scheme enables the improvement of knowledge graph quality and business support based on these evaluation results. If entity recognition accuracy is low, the analysis indicates problems with model parameters, training data, or feature extraction, leading to targeted model adjustments such as resetting hyperparameters or adding high-quality data. If knowledge graph coverage is insufficient, data collection channels are expanded to incorporate more entities and relationships. For incompleteness, the data entry and acquisition processes are checked, or more data sources are linked, and attribute information is improved using fill-in algorithms. For consistency issues, contradictory data is corrected according to preset rules and business logic.

[0091] The current solution utilizes multi-dimensional evaluation metrics to comprehensively and accurately measure the performance of entity recognition and relation extraction tasks, as well as the overall performance of the knowledge graph, providing a reliable basis for subsequent optimization. Targeted optimization measures based on the evaluation results can continuously improve the quality of entity recognition and relation extraction tasks, reducing misjudgments and omissions. Simultaneously, it continuously improves the coverage, completeness, and consistency of the knowledge graph, making it more comprehensive, information-rich, and consistent. This effectively supports various banking businesses, enhancing the scientific rigor and accuracy of risk assessment, customer marketing, and product recommendations, thereby increasing customer satisfaction. This solution addresses the problem of previous evaluation methods using only single and incomplete metrics for task and knowledge graph performance.

[0092] Corresponding to the above method, the present invention also provides a bank customer profile knowledge graph construction system based on multi-source heterogeneous data fusion, comprising: The data acquisition module is used to collect bank customer data, including both structured and unstructured data; it accesses multi-source data through ETL tools and multiple protocols. The knowledge framework building module is used to obtain hierarchical classification rules for customers and products, build a knowledge framework for customer profiles, and organize the collected customer data through the knowledge framework. The entity extraction module, based on the sorted customer data, extracts entities from structured data using entity extraction rules; and extracts entities from unstructured data using a deep learning model. The entity relationship extraction module extracts entity relationships from customer data based on the extracted entities and according to entity relationship extraction rules. At the same time, it extracts multi-dimensional features related to the entities and inputs them into a machine learning model with a set confidence threshold, outputting entity relationships that are higher than the confidence threshold. The knowledge graph building module is used to construct a customer profile knowledge graph based on entities, entity relationships, and knowledge frameworks, and to establish a dynamic update mechanism for the customer profile knowledge graph.

[0093] The manual review module is used to connect with technical personnel and their equipment, publish review tasks, and obtain manual processing results.

[0094] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for constructing a bank customer profile knowledge graph based on multi-source heterogeneous data fusion, characterized by: Includes the following steps: S1. Collect bank customer data, including structured and unstructured data. S2. Obtain hierarchical classification rules for customers and products, construct a knowledge framework for customer profiles, and organize the collected customer data through the knowledge framework; S3. Based on the sorted customer data, entities are extracted from structured data using entity extraction rules; entities are also extracted from unstructured data using a deep learning model. S4. Based on the extracted entities, entity relationships are extracted from customer data through entity relationship extraction rules; at the same time, multi-dimensional features related to the entities are extracted and input into a machine learning model with a set confidence threshold, and the output is entity relationships that are higher than the confidence threshold. S5. Based on entities, entity relationships, and knowledge framework, construct a customer profile knowledge graph and establish a dynamic update mechanism for the customer profile knowledge graph.

2. The method for constructing a bank customer profile knowledge graph based on multi-source heterogeneous data fusion as described in claim 1, characterized in that: In S1, after collecting customer data, encryption and sensitive information masking are performed. Customer data can be collected in three ways: scheduled collection, real-time collection, and incremental collection.

3. The method for constructing a bank customer profile knowledge graph based on multi-source heterogeneous data fusion as described in claim 1, characterized in that: In S2, the hierarchical classification rules for customers and products include: multi-level customer classification and multi-level product classification; by connecting the categories of multi-level customer classification and multi-level product classification, a knowledge framework for customer profiles is constructed.

4. The method for constructing a bank customer profile knowledge graph based on multi-source heterogeneous data fusion according to claim 1, characterized in that: in S3, the entity extraction method in unstructured data is: Word embedding technology is used to convert words in the text into low-dimensional vectors. The low-dimensional vector is then input into a deep learning model using a BiLSTM+CRF architecture. The entity category label assignment is determined by maximizing the model probability, thus completing entity recognition.

5. The method for constructing a bank customer profile knowledge graph based on multi-source heterogeneous data fusion according to claim 1, characterized in that: S3 In the process, after the entity is extracted, knowledge disambiguation is performed, and entity links are established with entities in the knowledge graph or entity library. For entities that fail knowledge disambiguation, they are pushed to the manual review module. For conflicting entity links, detection is performed according to preset detection rules and statistical analysis methods, and conflict handling is carried out based on the detection results.

6. The method for constructing a bank customer profile knowledge graph based on multi-source heterogeneous data fusion according to claim 1, characterized in that: In S4, the multidimensional features related to entities include: word vector features, part-of-speech features, syntactic structure features, and entity distance features; The machine learning model is also used to output entity relationships below the confidence threshold to the human review module, and to update the machine learning model with entity relationships that have been approved by the human review module.

7. The method for constructing a bank customer profile knowledge graph based on multi-source heterogeneous data fusion as described in claim 1, characterized in that: In S5, when constructing the knowledge graph, entity optimization is performed using ontology mapping and entity alignment techniques.

8. The method for constructing a bank customer profile knowledge graph based on multi-source heterogeneous data fusion as described in claim 1, characterized in that: After S5, S6 is executed, with the following content: Based on a customer profile knowledge graph with a dynamic update mechanism, it is stored in a graph database; After storage, customer importance is assessed by analyzing the customer profile knowledge graph; customer business risks and recommendation services are assessed through a hybrid inference engine; and the customer profile knowledge graph is quality checked and repaired through contradictory inference.

9. The method for constructing a bank customer profile knowledge graph based on multi-source heterogeneous data fusion according to claim 1, characterized in that: During the application of models S3 and S4, the models are evaluated using accuracy, recall, and F1 score, and then optimized based on the evaluation results.

10. A knowledge graph construction system for bank customer profiles based on multi-source heterogeneous data fusion, characterized by: include: The data acquisition module is used to collect bank customer data, including both structured and unstructured data. The knowledge framework building module is used to obtain hierarchical classification rules for customers and products, build a knowledge framework for customer profiles, and organize the collected customer data through the knowledge framework. The entity extraction module, based on the sorted customer data, extracts entities from structured data using entity extraction rules; and extracts entities from unstructured data using a deep learning model. The entity relationship extraction module extracts entity relationships from customer data based on the extracted entities and according to entity relationship extraction rules. At the same time, it extracts multi-dimensional features related to the entities and inputs them into a machine learning model with a set confidence threshold, outputting entity relationships that are higher than the confidence threshold. The knowledge graph building module is used to construct a customer profile knowledge graph based on entities, entity relationships, and knowledge frameworks, and to establish a dynamic update mechanism for the customer profile knowledge graph.