Multi-source data fusion-based credit business risk control decision method, apparatus and device, and medium
By fusing multi-source data to construct a knowledge graph, dynamically adjusting data weights and resolving conflicts, a comprehensive credit score is generated. This solves the problems of insufficient data coverage and semantic conflicts in traditional credit scoring models, enabling efficient and accurate risk control decisions.
Patent Information
- Application Number
- CN202510818406.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-06-18
Smart Images

Figure CN120833210A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of financial technology and artificial intelligence, and in particular to a credit business risk control decision method and device based on multi-source data fusion, equipment and medium. BACKGROUND
[0002] In the field of financial credit, accurately assessing the credit risk of credit subjects is the core link of risk management for financial institutions. Traditional credit scoring models mainly rely on static indicators of a single data source (such as credit reporting, bank flow), which has the problems of limited data coverage and insufficient real-time performance, and is difficult to fully reflect the dynamic credit status of credit subjects. With the rapid development of financial technology, multi-source heterogeneous data has gradually become an important information source to supplement traditional credit data.
[0003] However, the time-consuming proportion of cross-system data alignment is more than 60%, and there are semantic conflicts (such as large differences in the definition of "income stability" by different systems), which leads to high rejection rate of the rule engine-based method, and cannot identify cross-device associated risks, resulting in distorted risk assessment. In addition, there are significant differences in data update frequency and reliability between different data sources, for example, Internet of Things device data may be updated at a second level but there is sensor noise, and credit data has a long update cycle but high authority, thereby reducing the reliability of risk assessment. SUMMARY
[0004] The present application provides a credit business risk control decision method and device based on multi-source data fusion, which solves the problem of poor risk assessment reliability caused by multi-source heterogeneous data in the prior art.
[0005] The present application provides a credit business risk control decision method based on multi-source data fusion, comprising: Obtaining multi-source data related to the subject identifier of the credit subject, the multi-source data including credit data, Internet of Things device data and social network data; Performing entity analysis and alignment on the multi-source data based on the semantic features of each entity in the multi-source data to construct an initial knowledge graph for the credit subject; Based on the confidence scores of each data source in the initial knowledge graph and the corresponding time decay weight, the multi-source data is fused and conflict is resolved through a spatio-temporal decay fusion model, and based on the fusion result, entity filtering is performed on the initial knowledge graph to obtain an optimized knowledge graph; For each entity node associated with a heterogeneous data source in the optimized knowledge graph, a scoring model adapted to its data characteristics is called respectively, and a comprehensive credit score is generated through a model integration mechanism; Based on the comprehensive credit score, a risk control decision for the credit subject is determined.
[0006] The credit business risk control decision-making method based on multi-source data fusion provided by the application comprises the following steps: Extracting semantic features of each entity in the multi-source data; Taking each entity as a node, defining the association edge between nodes according to the business rules, and calculating the topological distance weight of the edge: Based on the semantic features of each node and the topological distance weight of the edge, the similarity between each two nodes is calculated; Based on the similarity between each two nodes and the preset similarity threshold, each node is merged.
[0007] The credit business risk control decision-making method based on multi-source data fusion provided by the application comprises the following steps: The similarity between each two nodes is determined based on the following formula:
[0008] Wherein, is the similarity between nodes and is the semantic feature of node is the semantic feature of node is the similarity calculation function, is the topological distance weight of the edge between node and node .
[0009] The credit business risk control decision-making method based on multi-source data fusion provided by the application comprises the following steps: Based on the confidence score and the corresponding time decay weight of each data source in the initial knowledge graph, the multi-source data is fused and conflict is resolved through a space-time decay fusion model, comprising: Based on the difference between the time decay weights of each data source and the time fusion confidence, the conflict resolution confidence is calculated; The time fusion confidence and the conflict resolution confidence are fused to obtain the final fusion confidence.
[0010] According to the credit risk control decision method based on multi-source data fusion provided by the application, before the time-weighted fusion of the multi-source data based on the confidence score of each data source and the corresponding time decay weight, the method further comprises the following steps: extracting multi-source data from an initial knowledge graph, each data source comprising a confidence score and a data timestamp; based on the data timestamp, calculating a time decay weight for each data source based on a time decay function.
[0011] According to the credit risk control decision method based on multi-source data fusion provided by the application, for each heterogeneous data source associated with an entity node in the optimized knowledge graph, a scoring model adapted to the data characteristics thereof is respectively called to generate a comprehensive credit score through a model integration mechanism, comprising the following steps: based on the data type of each data source, calling a scoring model adapted to the data type to obtain an individual credit score of each scoring model; adjusting the dynamic weight of each scoring model; based on the adjusted dynamic weight, weighting and fusing the individual credit scores to obtain the comprehensive credit score.
[0012] According to the credit risk control decision method based on multi-source data fusion provided by the application, the method of adjusting the dynamic weight of each scoring model comprises the following steps: adjusting the dynamic weight of each scoring model based on the following formula:
[0013] In the formula, is the dynamic weight of the scoring model , is the value of the scoring model , is the sum of the values of all scoring models, is the sum of the values of all scoring models, is a feature drift index.
[0014] The application further provides a credit risk control decision device based on multi-source data fusion, comprising: a data acquisition unit configured to acquire multi-source data related to a subject identifier of a credit subject, the multi-source data comprising credit investigation data, Internet of Things device data and social network data; a graph construction unit configured to perform entity analysis and alignment on the multi-source data based on the semantic characteristics of each entity in the multi-source data to construct an initial knowledge graph for the credit subject; a data fusion unit configured to fuse and resolve conflicts of the multi-source data based on confidence scores of each data source in the initial knowledge graph and corresponding time decay weights through a spatio-temporal decay fusion model, perform entity screening on the initial knowledge graph based on a fusion result, and obtain an optimized knowledge graph; a credit scoring unit configured to call a scoring model adapted to data characteristics of each entity node associated with a heterogeneous data source in the optimized knowledge graph, respectively, and generate a comprehensive credit score through a model integration mechanism; a decision determination unit configured to determine a risk control decision for the credit subject based on the comprehensive credit score.
[0015] The application also provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the credit business risk control decision method based on multi-source data fusion according to any one of the above.
[0016] The application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and the computer program is executable by a processor to implement the credit business risk control decision method based on multi-source data fusion according to any one of the above.
[0017] The application also provides a computer program product including a computer program, and the computer program is executable by a processor to implement the credit business risk control decision method based on multi-source data fusion according to any one of the above.
[0018] The credit business risk control decision method, device, equipment and medium based on multi-source data fusion provided by the application integrate multi-source heterogeneous data such as credit investigation, Internet of Things devices and social networks, implement high-precision entity analysis and alignment to construct an initial knowledge graph based on a BERT-Graph model, dynamically adjust data weights and resolve conflicts using a spatio-temporal decay fusion model, call an adaptive scoring model integration to generate a comprehensive credit score for heterogeneous data sources after screening and optimizing the knowledge graph, and finally implement accurate and efficient risk control decisions, significantly improve risk identification accuracy, real-time performance and model robustness, reduce false rejection rate and fraud loss, and enhance financial risk control efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0020] Figure 1is one of the process schematic diagrams of the credit business risk control decision method based on multi-source data fusion provided by the application.
[0021] Figure 2 is a structural schematic diagram of the credit business risk control decision device based on multi-source data fusion provided by the application.
[0022] Figure 3 is a structural schematic diagram of the electronic device provided by the application. DETAILED DESCRIPTION
[0023] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described below in connection with the drawings in the present application, which will make the technical solutions in the present application clearer, complete and more comprehensible. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0024] In view of the problem of poor risk assessment reliability caused by multi-source heterogeneous data, the embodiment of the present application provides a knowledge-driven credit business risk control decision method fusing multi-source heterogeneous data and supporting dynamic time decay and conflict resolution. In the method, first, multi-source data related to the subject identifier of the credit subject is acquired, the multi-source data including credit investigation data, Internet of Things device data and social network data; the multi-source data is parsed and aligned based on the semantic features of each entity in the multi-source data, so as to construct an initial knowledge graph for the credit subject, convert the multi-source heterogeneous data into a standardized entity-relation graph, unify the field expressions of different data sources, eliminate data islands and provide a structured basis for subsequent fusion calculation; based on the confidence scores and the corresponding time decay weights of each data source in the initial knowledge graph, the multi-source data is fused and conflict is resolved through a time-space decay fusion model, the sensitivity of the model to dynamic risks is improved, entity filtering is performed on the initial knowledge graph based on the fusion result, and an optimized knowledge graph is obtained; for each entity node associated with a heterogeneous data source in the optimized knowledge graph, a scoring model adapted to the data features thereof is called respectively, and a comprehensive credit score is generated through a model integration mechanism; based on the comprehensive credit score, a risk control decision for the credit subject is determined.
[0025] The embodiment of the present application integrates multi-source heterogeneous data such as credit investigation, Internet of Things device and social network, realizes high-precision entity parsing and alignment based on a BERT-Graph model to construct an initial knowledge graph, dynamically adjusts data weights and resolves conflicts through a time-space decay fusion model, generates a comprehensive credit score through an adaptive scoring model integration after filtering and optimizing the knowledge graph, finally realizes accurate and efficient risk control decision, significantly improves risk identification accuracy, real-time performance and model robustness, reduces false rejection rate and fraud loss, and enhances financial risk control efficiency.
[0026] The embodiments of the present application can be applied to the credit risk control decision-making scenarios, such as bank, supply chain finance, cross-border payment and the like. The execution subject of the method can be an electronic device such as terminal device, computer, server, server cluster or specially designed risk control decision-making device, or a risk control decision-making device arranged in the electronic device, which can be realized by software, hardware or combination of both.
[0027] In the description of the embodiments of the present application, the meaning of "a plurality of" is two or more than two, unless otherwise explicitly specified.
[0028] Figure 1 is one of the flowcharts of the credit risk control decision-making method based on multi-source data fusion provided by the present application, as Figure 1 shown, the method comprises the following steps: Step 110, acquiring multi-source data related to the subject identification of the credit subject, the multi-source data including credit investigation data, Internet of Things device data and social network data.
[0029] Specifically, the credit subject refers to an individual or enterprise applying for credit service to a financial institution. Its core features include subject identification such as ID number, unified social credit code, credit record and financial status, which is the object of credit evaluation.
[0030] Multi-source data is a collection of data from different systems or platforms, with heterogeneous (structured, unstructured mixed) and multi-modal features (text, numerical, time series data). Three types are involved in this step, credit investigation data, Internet of Things device data and social network data.
[0031] Among them, the credit investigation data refers to the authoritative credit record provided by the credit investigation agency, such as loan history, repayment status, debt situation. The data source can include credit investigation center, third-party credit investigation agency, etc. The credit report of the credit subject is synchronized regularly through API interface or data file (such as XML, JSON). The credit investigation data includes: basic information such as name, ID number, occupation; credit record such as loan amount, repayment status, overdue times; public records such as tax arrears, judicial litigation, etc.
[0032] The Internet of Things device data refers to the real-time behavior data collected by sensors and intelligent terminals, such as GPS location, device usage time, temperature and humidity. The data source is the sensor deployed on the device associated with the credit subject, such as smart phone GPS, vehicle OBD, smart home, etc. The device data is uploaded to the Internet of Things platform in real time through MQTT, HTTP protocol. The Internet of Things device data specifically includes: device ID, geographic location (latitude and longitude), usage time, abnormal events such as device fault code, etc.
[0033] Social network data refers to the behavior and relationship data of users on social platforms, such as friend networks, interaction frequency, and consumption preferences. Friend relationship chains, interaction behaviors, and consumption records can be obtained through user-authorized APIs, or public social graph data such as fan numbers and community affiliations can be crawled through web crawlers under compliance.
[0034] Preferably, the multi-source data can also be pre-processed and standardized, such as uniform field format, default value filling, outlier removal, and noise reduction.
[0035] Step 120, based on the semantic features of each entity in the multi-source data, the multi-source data is parsed and aligned to construct an initial knowledge graph for the credit subject.
[0036] Specifically, each entity in the multi-source data refers to an object with a clear business meaning, such as a customer entity: a credit applicant, which can be an individual or an enterprise; a device entity: an Internet of Things device associated with the customer, such as a smartphone or a vehicle terminal; a location entity: a geographic location involved in the customer's activities, such as a home address or a work unit.
[0037] The semantic features of each entity refer to the semantic information contained in the natural language or data fields. Semantic features can be achieved through text embedding based on BERT model or feature encoding.
[0038] Entity resolution refers to identifying entities in multi-source data that point to the same real-world object, such as associating "Zhang San" in different systems as the same customer. Entity alignment refers to mapping the parsed entities to standardized nodes in the knowledge graph, eliminating the problem of homonymy.
[0039] Here, the semantic mapping rule library can be used in conjunction with entity alignment. The semantic mapping rule library contains 2000+ cross-system term alignment logic, such as defining "monthly income" = salary stream + e-commerce consumption + tax data.
[0040] The initial knowledge graph obtained in this way is a structured network centered on the credit subject, constructed through entity relationships such as "customer-device binding" and "device-location association", which is used to support subsequent credit evaluation and risk analysis.
[0041] In some possible implementations, step 120 specifically includes: Step 121, extracting the semantic features of each entity in the multi-source data; Step 122, taking each entity as a node, defining the association edges between nodes according to business rules, and calculating the topological distance weight of the edges: Step 123, based on the semantic features of each node and the topological distance weight of the edges, calculating the similarity between each two nodes; Step 124, based on the similarity between each two nodes and the preset similarity threshold, merging each node.
[0042] Specifically, for text type entities, the customer name, device ID description, address text, etc. are converted into high-dimensional semantic vectors using a pre-trained BERT model to obtain semantic features.
[0043] For numerical feature encoding, after standardization, it can be mapped to a low-dimensional vector through a fully connected layer.
[0044] For category type feature encoding, it is converted into a dense vector through an Embedding layer.
[0045] An entity relationship graph is constructed through a Graph neural network. Each entity, such as a customer, a device, and a location, is taken as a node in the graph, and the node features are semantic feature vectors (BERT vectors + numerical / category embedding vectors). According to business rules, the association edges between nodes are defined, such as a “customer-device” binding relationship and a “device-location” usage record, and the topological distance weight of the edge is calculated.
[0046] Directly associated relationships, such as customer A explicitly binding device X, can set the topological distance weight as w = 1.0; indirect associated relationships, such as customer A→device X→location Y, can be attenuated according to the path length, such as setting w = 0.6.
[0047] Considering that the constructed entity relationship graph may contain similar entities, in order to improve the cross-system entity matching accuracy and eliminate the problems of homonymy and synonymy, a unified credit subject portrait is constructed, and the similarity between each two nodes can be calculated.
[0048] Here, the similarity between the node pair can be calculated by combining the semantic similarity and the topological distance weight. The semantic similarity can be obtained by calculating the cosine similarity of the BERT vector. The higher the cosine similarity, the closer the semantics of the two nodes.
[0049] In some embodiments, the similarity between each two nodes is determined based on the following formula:
[0050] wherein, is the similarity between nodes and , is the semantic feature of node , is the semantic feature of node , is a similarity calculation function, is a node with nodes topological distance weight of edges between nodes.
[0051] After obtaining the similarity between each two nodes, the similar nodes can be merged. The similarity threshold can be set according to business requirements, combined with statistical distribution or expert experience. If the similarity is greater than or equal to the similarity threshold, it is determined that the two nodes are the same entity, and merging is performed. At the same time of merging, the high confidence attribute is inherited, such as updating the customer address with the latest data. All associated edges of the merged entity are retained, and the weight of the directly associated edge is improved. If the similarity is less than the similarity threshold, it is determined that the two nodes are different entities, and are independently retained. In addition, for the conflict field, a rule base or voting mechanism is used to solve. Thus, the initial knowledge graph representing the credit subject portrait is constructed.
[0052] In step 130, based on the confidence scores of each data source in the initial knowledge graph and the corresponding time decay weight, the multi-source data is fused and conflict is resolved through a space-time decay fusion model, and based on the fusion result, entity screening is performed on the initial knowledge graph to obtain an optimized knowledge graph.
[0053] Specifically, the node attributes in the initial knowledge graph have been preliminarily standardized through semantic parsing and alignment. Before step 130 is performed, it includes: Extracting multi-source data from the initial knowledge graph, each data source including a confidence score and a data timestamp; based on the data timestamp, calculating the time decay weight of each data source based on a time decay function.
[0054] In this step, the initial knowledge graph nodes are traversed, and the multi-source data associated with each entity and its source are recorded, such as customer A associated with credit data and device X associated with Internet of Things data. Extract the entity-data source mapping in the initial knowledge graph, each data source including a confidence score and a data timestamp. The confidence score is a predicted confidence indicator output by each data source, with a value range of [0, 1], reflecting data reliability. For example, the credit score confidence of the credit data output by the credit model is high, such as 0.9; the Internet of Things data is scored by sensor data quality; the confidence score of the social network data is obtained by scoring the authenticity of user behavior data.
[0055] The time decay weight measures the dynamic weight of data freshness, which is exponentially decayed with the age of data. The calculation formula of the time decay weight can be expressed as:
[0056] In the formula, represents the time decay weight of the i-th data source when the data age is t i, = 0.85.
[0057] After obtaining the confidence scores of each data source and its corresponding time decay weight, the multi-source data is fused and conflict resolved using a spatiotemporal decay fusion model. The spatiotemporal decay fusion model dynamically adjusts the contribution of multi-source data using time decay weights and confidence scores, and resolves conflicting data using a weighted fusion model, achieving the effect of recent high-confidence data dominating and the impact of conflicting data being dynamically suppressed. In some embodiments, step 130 specifically includes: Step 131: Based on the confidence scores of each data source and its corresponding time decay weight, perform time-weighted fusion on the multi-source data to obtain a time fusion confidence score; Step 132: Calculate the conflict resolution confidence based on the difference between the time decay weights of the data sources and the time fusion confidence. Step 133: Fusing the temporal fusion confidence with the conflict resolution confidence to obtain a final fusion confidence.
[0058] For each entity, the weighted average confidence of its multi-source data is calculated, that is, the time fusion confidence. The formula is: Time fusion confidence = ,in For the The confidence score of each data source, The total number of data sources.
[0059] Through the product term Quantify the degree of data conflict. If the time decay weights of each data source are close and the difference is small, then If the time decay weights of the data sources are different, then Increase to suppress low-weight conflict data.
[0060] The confidence level of conflict resolution is expressed as: Conflict resolution confidence =
[0061] The final fusion confidence is the sum of the temporal fusion confidence and the conflict resolution confidence. Based on this fusion result, the initial knowledge graph is screened for entities to obtain an optimized knowledge graph. First, a dynamic threshold is set. Entity nodes and their associated edges whose final fusion confidence is less than the set dynamic threshold are deleted. The knowledge graph structure is then updated to obtain the optimized knowledge graph.
[0062] This embodiment uses time-decayed weighting to assign greater decision-making power to recent data, adapting to dynamic environments, such as sudden changes in a customer's credit status. Dynamically suppressing interference from highly conflicting data through product terms reduces scoring errors. Confidence filtering and redundancy elimination improve the accuracy and consistency of graph structure and attributes, supporting more precise risk management decisions.
[0063] Step 140, for each entity node associated with the heterogeneous data source in the optimized knowledge graph, call the scoring model adapted to its data characteristics respectively, and generate a comprehensive credit score through model integration mechanism.
[0064] Specifically, the optimized knowledge graph processed by the spatio-temporal decay fusion model contains high-confidence entity nodes and their associated relationships, and the data quality and consistency have been significantly improved, which can be used as a reliable basis for credit assessment.
[0065] Each entity node associated with the heterogeneous data source has data type heterogeneity, semantic heterogeneity and temporal dynamics. Data types include structured data, unstructured data and semi-structured data. Therefore, the scoring model is designed for specific data source characteristics to assess credit risk. For example: credit scoring model based on logistic regression or XGBoost, using historical loan records, repayment behavior, etc. to calculate credit score. Internet of Things data model uses time series convolution network (TCN) or LSTM to mine device usage behavior patterns such as device active period and abnormal event frequency. Social network data model based on graph neural network (GNN) to mine social credit risk through friend relationships and interaction behavior.
[0066] The output results of multiple scoring models are fused into a comprehensive credit score through weighting or other strategies to improve the comprehensiveness and robustness of the evaluation.
[0067] In some embodiments, step 140 specifically includes: Step 141, based on the data type of each data source, call the scoring model adapted to the data type to obtain individual credit scores for each scoring model; Step 142, adjust the dynamic weight of each scoring model; Step 143, based on the adjusted dynamic weight, weight and fuse the individual credit scores to obtain a comprehensive credit score.
[0068] Specifically, the data type of each data source specifically includes: Structured data: tabular form, fields are clear, such as "loan amount" and "repayment status" in credit data.
[0069] Unstructured data: no fixed format, such as text (social comments), images (device photos).
[0070] Semi-structured data: between the two, such as JSON format Internet of Things logs.
[0071] Time series data: data arranged in chronological order, such as temperature values recorded by device sensors every minute.
[0072] Graph data: a network of relationships represented as nodes and edges, such as friend connections in a social network.
[0073] Score models adapted to data types refer to predictive models designed for specific data types to assess credit risk of credit subjects, for example: Logistic regression: a linear risk prediction suitable for structured data.
[0074] XGBoost / LightGBM: an efficient tree model for structured data.
[0075] Temporal convolutional network (TCN): modeling local patterns in time series data such as device sensor logs.
[0076] Long short-term memory network (LSTM): capturing long-term dependencies in time series data.
[0077] Graph neural network (GNN): analyzing node relationship risks in graph data.
[0078] Adjust the weight coefficient according to the recent performance of the model, such as AUC, KS value and data quality, to control the contribution of each model in the comprehensive credit score, that is, adjust the dynamic weight of each scoring model.
[0079] In some embodiments, first calculate the performance indicators of each model on recent data, including AUC value: measures the ability of the model to distinguish between positive and negative samples, the higher the value the better the performance; KS value: assesses the discrimination of credit score for good and bad samples; KS drift indicator: detects the impact of feature distribution changes on model performance. The dynamic weight of each scoring model considers both model performance and data stability.
[0080] The dynamic weight of each scoring model can be adjusted based on the following formula:
[0081] In the formula, is the dynamic weight of the scoring model , is the value of the scoring model , is the sum of all values of the scoring model , and is the feature drift indicator.
[0082] The individual credit scores of each model are weighted and summed according to the dynamic weight to obtain the comprehensive credit score.
[0083] In some embodiments, the scoring model can be iteratively optimized by incremental learning and quantum optimization technology. The feature selection task is converted into a quadratic unconstrained binary optimization (QUBO) problem, and a quantum processor (such as a quantum annealing machine) is used to efficiently solve the QUBO problem to obtain the optimal feature subset, thereby significantly improving the feature selection speed and model iteration efficiency.
[0084] In step 150, a risk control decision for the credit subject is determined based on the comprehensive credit score.
[0085] After obtaining the comprehensive credit score, a risk control decision for the credit subject is determined according to the comprehensive credit score. For example, if the comprehensive credit score indicates low risk, the risk control decision can be automatic lending and dynamically increasing the loan amount; if the comprehensive credit score indicates medium risk, the risk control decision can be biometric authentication plus manual review; and if the comprehensive credit score indicates high risk, the risk control decision can be blockchain blacklist marking plus automatic transaction link break.
[0086] The method provided by the embodiment of the application integrates multi-source heterogeneous data such as credit investigation, Internet of Things equipment and social networks, realizes high-precision entity analysis and alignment to construct an initial knowledge graph based on a BERT-Graph model, dynamically adjusts data weights and eliminates conflicts by using a time-space attenuation fusion model, filters and optimizes the knowledge graph, calls an adaptive scoring model set for heterogeneous data sources to generate a comprehensive credit score, and finally realizes accurate and efficient risk control decisions, significantly improves risk identification accuracy, real-time performance and model robustness, reduces false rejection rate and fraud loss, and enhances financial risk control efficiency.
[0087] It should be noted that the edge computing platform can separate the internal network (production data) and the external network (Internet) through a secure isolation barrier to block external attacks and improve data security.
[0088] The credit business risk control decision device based on multi-source data fusion provided by the application will be described below. The credit business risk control decision device based on multi-source data fusion described below can be correspondingly referred to the credit business risk control decision method based on multi-source data fusion described above.
[0089] Based on any of the above embodiments, Figure 2 is a structural schematic diagram of the credit business risk control decision device based on multi-source data fusion provided by the application, as Figure 2 shown, the device comprises: A data acquisition unit 210 is configured to acquire multi-source data related to the subject identifier of the credit subject, wherein the multi-source data comprises credit investigation data, Internet of Things equipment data and social network data. The graph construction unit 220 is configured to perform entity resolution and alignment on the multi-source data based on semantic features of each entity in the multi-source data, so as to construct an initial knowledge graph for the credit subject. The data fusion unit 230 is configured to perform fusion and conflict resolution on the multi-source data based on confidence scores of each data source in the initial knowledge graph and corresponding time decay weights, and perform entity screening on the initial knowledge graph based on a fusion result, so as to obtain an optimized knowledge graph. The credit scoring unit 240 is configured to call a scoring model adapted to data features of each entity node in the optimized knowledge graph for each heterogeneous data source associated with the entity node, and generate a comprehensive credit score through model integration. The decision determination unit 250 is configured to determine a risk control decision for the credit subject based on the comprehensive credit score.
[0090] Based on the above embodiment, the graph construction unit is specifically configured to: extract semantic features of each entity in the multi-source data; define an associated edge between each entity as a node according to a business rule, and calculate a topological distance weight of the edge: calculate a similarity between each two nodes based on the semantic features of each node and the topological distance weight of the edge; merge each node based on the similarity between each two nodes and a preset similarity threshold.
[0091] Based on the above embodiment, the graph construction unit is specifically configured to: determine the similarity between each two nodes based on the following formula:
[0092] wherein, is a similarity between the node and the node , is a semantic feature of the node , is a semantic feature of the node , is a similarity calculation function, is a topological distance weight of an edge between the node and the node .
[0093] Based on the above embodiment, the data fusion unit is specifically configured to: perform time-weighted fusion on the multi-source data based on the confidence scores of each data source and corresponding time decay weights, to obtain a time fusion confidence; Calculating conflict resolution confidence based on the difference between the time decay weights of the data sources and the time fusion confidence; The temporal fusion confidence is fused with the conflict resolution confidence to obtain a final fusion confidence.
[0094] Based on the above embodiment, the data fusion unit is specifically used for: Extract multi-source data from the initial knowledge graph, each data source includes a confidence score and data timestamp; Based on the data timestamps, a time decay weight is calculated for each data source based on a time decay function.
[0095] Based on the above embodiment, the credit scoring unit is specifically used to: Based on the data type of each data source, calling a scoring model adapted to the data type to obtain a separate credit score for each scoring model; Adjust the dynamic weight of each scoring model; Based on the adjusted dynamic weights, the individual credit scores are weighted and integrated to obtain the comprehensive credit score.
[0096] Based on the above embodiment, the credit scoring unit is specifically used to: The dynamic weight of each scoring model is adjusted based on the following formula:
[0097] Where, For scoring model The dynamic weight of For scoring model of value, For all Scoring model The sum of values, is the characteristic drift indicator.
[0098] Figure 3 An example of a physical structure diagram of an electronic device is shown below. Figure 3As shown, the electronic device can include a processor 310, a communications interface 320, a memory 330, and a communications bus 340, wherein the processor 310, the communications interface 320, and the memory 330 communicate with each other through the communications bus 340. The processor 310 can invoke the logic instructions in the memory 330 to execute the credit business risk control decision method based on multi-source data fusion, which includes: acquiring multi-source data related to the subject identifier of the credit subject, the multi-source data including credit investigation data, Internet of Things device data, and social network data; performing entity analysis and alignment on the multi-source data based on the semantic features of each entity in the multi-source data to construct an initial knowledge graph for the credit subject; based on the confidence scores of each data source in the initial knowledge graph and the corresponding time decay weights, performing fusion and conflict resolution on the multi-source data through a spatiotemporal decay fusion model, performing entity screening on the initial knowledge graph based on the fusion result to obtain an optimized knowledge graph; for each entity node associated with the heterogeneous data source in the optimized knowledge graph, respectively invoke the scoring model adapted to its data characteristics, and generate a comprehensive credit score through a model integration mechanism; based on the comprehensive credit score, determine the risk control decision for the credit subject.
[0099] In addition, the logic instructions in the memory 330 described above can be implemented in the form of a software functional unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the part of the prior art that essentially contributes or the part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0100] In another aspect, the present application also provides a computer program product comprising a computer program, which can be stored on a non-transitory computer-readable storage medium, and the computer program can be executed by a processor to enable a computer to perform the credit risk control decision-making method based on multi-source data fusion provided by the above-mentioned methods, which comprises: acquiring multi-source data related to the subject identifier of a credit subject, wherein the multi-source data comprises credit investigation data, Internet of Things device data and social network data; performing entity analysis and alignment on the multi-source data based on the semantic features of each entity in the multi-source data to construct an initial knowledge graph for the credit subject; performing fusion and conflict resolution on the multi-source data based on the confidence scores of each data source in the initial knowledge graph and the corresponding time decay weights thereof through a spatio-temporal decay fusion model, performing entity screening on the initial knowledge graph based on the fusion result to obtain an optimized knowledge graph; calling a scoring model adapted to the data features of each heterogeneous data source associated with each entity node in the optimized knowledge graph respectively to generate a comprehensive credit score through a model integration mechanism; and determining a risk control decision for the credit subject based on the comprehensive credit score.
[0101] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and the computer program can be executed by a processor to implement the credit risk control decision-making method based on multi-source data fusion provided by the above-mentioned methods, which comprises: acquiring multi-source data related to the subject identifier of a credit subject, wherein the multi-source data comprises credit investigation data, Internet of Things device data and social network data; performing entity analysis and alignment on the multi-source data based on the semantic features of each entity in the multi-source data to construct an initial knowledge graph for the credit subject; performing fusion and conflict resolution on the multi-source data based on the confidence scores of each data source in the initial knowledge graph and the corresponding time decay weights thereof through a spatio-temporal decay fusion model, performing entity screening on the initial knowledge graph based on the fusion result to obtain an optimized knowledge graph; calling a scoring model adapted to the data features of each heterogeneous data source associated with each entity node in the optimized knowledge graph respectively to generate a comprehensive credit score through a model integration mechanism; and determining a risk control decision for the credit subject based on the comprehensive credit score.
[0102] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment scheme. Those skilled in the art can understand and implement it without creative labor.
[0103] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A credit risk control decision method based on multi-source data fusion, characterized in that, The method comprises the following steps: acquiring multi-source data related to the subject identification of the credit subject, wherein the multi-source data comprises credit investigation data, Internet of Things device data and social network data; performing entity analysis and alignment on the multi-source data based on the semantic features of each entity in the multi-source data, so as to construct an initial knowledge graph for the credit subject; performing fusion and conflict resolution on the multi-source data based on the confidence scores of each data source in the initial knowledge graph and the corresponding time decay weights, performing entity screening on the initial knowledge graph based on the fusion result, and obtaining an optimized knowledge graph; for each entity node associated with a heterogeneous data source in the optimized knowledge graph, calling a scoring model adapted to the data characteristics of the data source, and generating a comprehensive credit score through a model integration mechanism; determining a risk control decision for the credit subject based on the comprehensive credit score. 2.The credit risk control decision-making method based on multi-source data fusion according to claim 1, characterized in that, The method comprises the following steps: extracting the semantic features of each entity in the multi-source data; taking each entity as a node, defining the association edges between nodes according to business rules, and calculating the topological distance weights of the edges; calculating the similarity between each two nodes based on the semantic features of each node and the topological distance weights of the edges; merging each node based on the similarity between each two nodes and a preset similarity threshold. 3.The credit risk control decision-making method based on multi-source data fusion according to claim 2, characterized in that, The method comprises the following steps: determining the similarity between each two nodes based on the following formula: wherein, is a node is a node between the nodes is a semantic feature of a node is a semantic feature of a node is a semantic feature of a node is a semantic feature of a node is a similarity calculation function is a node is a node is a topological distance weight of an edge between the nodes 4. The credit business risk control decision-making method based on multi-source data fusion according to claim 1 is characterized in that: The method comprises the following steps: performing time-weighted fusion on the multi-source data based on the confidence scores of each data source and the corresponding time decay weights, and obtaining a time fusion confidence; calculating a conflict resolution confidence based on the differences between the time decay weights of each data source and the time fusion confidence; fusing the time fusion confidence and the conflict resolution confidence to obtain a final fusion confidence. 5.The credit risk control decision-making method based on multi-source data fusion according to claim 4, characterized in that, The method further comprises the following steps before performing time-weighted fusion on the multi-source data based on the confidence scores of each data source and the corresponding time decay weights: extracting multi-source data from the initial knowledge graph, wherein each data source comprises a confidence score and a data timestamp; calculating the time decay weight of each data source based on the data timestamp and a time decay function. 6.The credit risk control decision-making method based on multi-source data fusion according to claim 1, characterized in that, The method comprises the following steps: based on the data type of each data source, calling a scoring model adapted to the data type to obtain an individual credit score of each scoring model; adjusting the dynamic weight of each scoring model; performing weighted fusion on the individual credit scores based on the adjusted dynamic weight to obtain the comprehensive credit score.
7. The credit risk control decision-making method based on multi-source data fusion according to claim 6, characterized in that, The method comprises the following steps: The dynamic weight of each scoring model is adjusted based on the following formula: wherein is a dynamic weight of the scoring model , is a value of the scoring model , , is a sum of values of all scoring models, , is a feature drift indicator.
8. A credit business risk control decision device based on multi-source data fusion, characterized in that, Comprise: The data acquisition unit is configured to acquire multi-source data related to the subject identifier of the credit subject, wherein the multi-source data comprises credit investigation data, Internet of Things device data, and social network data; The graph construction unit is configured to perform entity analysis and alignment on the multi-source data based on semantic features of entities in the multi-source data, so as to construct an initial knowledge graph for the credit subject; The data fusion unit is configured to fuse and resolve conflicts of the multi-source data based on confidence scores of each data source in the initial knowledge graph and corresponding time decay weights, to perform entity screening on the initial knowledge graph based on a fusion result, and to obtain an optimized knowledge graph; The credit scoring unit is configured to call a scoring model adapted to the data features of each entity node associated with a heterogeneous data source in the optimized knowledge graph, and to generate a comprehensive credit score through a model integration mechanism; The decision determination unit is configured to determine a risk control decision for the credit subject based on the comprehensive credit score.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, The processor executes the computer program to implement the credit business risk control decision method based on multi-source data fusion according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the credit business risk control decision method based on multi-source data fusion according to any one of claims 1 to 7.
Citation Information
Patent Citations
Defective asset case retrieval analysis method and system
CN120147016A
Risk prediction method and apparatus, and device and storage medium
WO2023065545A1
Cited By
Credit investigation data processing method and device
CN121120237A
Government affair approval intelligent aid decision-making method and system based on multi-source data fusion
CN121707516A