Financial risk identification method and device based on dynamic knowledge graph, equipment and medium

By dynamically updating the financial knowledge graph through a distributed data acquisition architecture and incremental learning algorithms, and by mining the relationships between entities and constructing risk weights, combined with a risk transmission calculation model, the timeliness and accuracy of financial risk identification are solved, adapting to the risk prevention and control needs of multiple scenarios and improving the accuracy and interpretability of risk identification.

CN122264946APending Publication Date: 2026-06-23THE BANK OF CHONGQING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE BANK OF CHONGQING CO LTD
Filing Date
2026-03-25
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing financial knowledge graphs are mostly statically constructed, relying on manual updates or batch updates at fixed intervals, resulting in insufficient timeliness and difficulty in integrating multi-source heterogeneous data. Traditional risk transmission calculations cannot accurately locate risk transmission paths and default probabilities, leading to poor accuracy and interpretability in risk identification, low efficiency, and difficulty in supporting the needs of precise risk prevention and control.

Method used

Financial data is collected in real time through a pre-set distributed data acquisition architecture, deduplicated and standardized, and the financial knowledge graph is updated using incremental learning algorithms. Inter-entity correlation mining and risk weight construction are carried out, and risk identification is performed by combining a risk transmission calculation model to generate a risk identification report.

Benefits of technology

It improves the accuracy and interpretability of financial risk identification, adapts to risk prevention and control needs in multiple scenarios, supports the real-time requirements of high-frequency scenarios such as credit approval and risk monitoring, adapts to risk prevention for SMEs and micro-enterprises and cross-institutional risk joint prevention, and meets the traceability and audit requirements of regulatory compliance scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122264946A_ABST
    Figure CN122264946A_ABST
Patent Text Reader

Abstract

The application discloses a financial risk identification method and device based on a dynamic knowledge graph, equipment and a medium, relates to the technical field of financial risk prevention and control, and comprises the following steps: collecting target financial data in real time through a preset distributed data collection architecture and performing deduplication filtering to obtain an initial financial data set; performing standardization processing on the initial financial data set and performing knowledge fusion through a target entity disambiguation algorithm to obtain a standardized financial data set; after it is determined that a preset update condition is met, updating a preset financial knowledge graph based on the standardized financial data set and using an incremental learning algorithm; performing correlation mining between entities for the updated financial knowledge graph and determining a correlation risk weight to construct a correlation risk network; performing risk transmission calculation on the correlation risk network by using a preset risk transmission calculation model, performing risk identification based on the risk transmission calculation result and through a preset risk identification model, and outputting a risk identification report.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of financial risk prevention and control technology, and in particular to a method, apparatus, equipment and medium for identifying financial risks based on dynamic knowledge graphs. Background Technology

[0002] Existing financial knowledge graphs are mostly statically constructed, relying on manual updates or fixed-period batch updates, resulting in insufficient timeliness. Furthermore, current technologies struggle to integrate heterogeneous data from multiple sources such as business registration, credit reporting, news, and transactions, leading to the omission of relevant risk nodes. In addition, traditional risk transmission calculations are often based on simple topological analysis, easily ignoring node heterogeneity and risk decay, failing to accurately pinpoint risk transmission paths and default probabilities, and prone to misjudgment or omission of risks. This results in poor accuracy and interpretability in financial risk identification, low efficiency, and an inability to support the needs of precise risk prevention and control.

[0003] In conclusion, improving the accuracy and interpretability of financial risk identification to support the need for precise risk prevention and control is an urgent problem that needs to be solved. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a method, apparatus, device, and medium for financial risk identification based on dynamic knowledge graphs, which can improve the accuracy and interpretability of financial risk identification to support the needs of precise risk prevention and control. The specific solution is as follows: Firstly, this application provides a financial risk identification method based on dynamic knowledge graphs, applied to a target financial cloud platform, including: By using a pre-defined distributed data acquisition architecture, target financial data is collected in real time, and the target financial data is deduplicated to obtain an initial financial dataset. The initial financial dataset is standardized, and knowledge fusion is performed using a target entity disambiguation algorithm to obtain a standardized financial dataset. After determining that the preset update conditions are met, the preset financial knowledge graph is updated based on the standardized financial dataset and using an incremental learning algorithm to obtain the updated financial knowledge graph. The updated financial knowledge graph is used to mine the relationships between entities and determine the association risk weights, so as to construct a corresponding association risk network based on the association risk weights. The risk transmission calculation is performed on the associated risk network using a preset risk transmission calculation model. Based on the obtained risk transmission calculation results, the risk is identified through a preset risk identification model, and a corresponding risk identification report is generated and output.

[0005] Optionally, the step of collecting target financial data in real time through a preset distributed data acquisition architecture and deduplicating the target financial data to obtain an initial financial dataset includes: First Financial Data was scraped using a multi-threaded crawler configured based on the Scrapy framework. Call the preset API interface and use the preset encrypted transmission protocol to obtain the second financial data, so as to obtain the target financial data including the first financial data and the second financial data; The financial data in the target financial data that meets the preset invalidity conditions are filtered according to preset filtering rules, and duplicate data are removed based on a hash algorithm to obtain the initial financial dataset.

[0006] Optionally, the initial financial dataset includes structured financial data, semi-structured financial data, and unstructured financial data; Accordingly, the standardization of the initial financial dataset and the knowledge fusion through a target entity disambiguation algorithm to obtain a standardized financial dataset include: The fields in the structured financial data are mapped to standard fields in the target financial field using the field mapping method, and the numerical data in the structured financial data are normalized to obtain the first standardized data. The semi-structured financial data is parsed using XML path language to convert it into a target structured format, thus obtaining second standardized data. The unstructured financial data is segmented and part-of-speech tagging is performed based on a preset natural language processing model to obtain third standardized data. The target similarity between entities in the target standardized data is determined, and the target similarity is compared with a preset similarity threshold. Knowledge fusion is performed based on the comparison results to obtain a standardized financial dataset. The target similarity is determined based on the entity's name similarity, attribute similarity, and contextual similarity.

[0007] Optionally, after determining that the preset update conditions are met, updating the preset financial knowledge graph based on the standardized financial dataset and using an incremental learning algorithm to obtain the updated financial knowledge graph includes: Determine the current financial data increment based on the standardized financial dataset; If the incremental financial data exceeds the preset data increment threshold, or if the target key data is updated, or if the preset update time period is reached, then the newly added entity set and the newly added relationship set between the financial data increment are determined. An incremental learning algorithm is used to update the preset financial knowledge graph based on the newly added entity set and the newly added set of relationships between entities, so as to obtain the updated financial knowledge graph; the preset financial knowledge graph is constructed based on a preset standardized financial knowledge dataset, including an entity set and a set of relationships between entities.

[0008] Optionally, the step of mining inter-entity associations in the updated financial knowledge graph and determining association risk weights to construct a corresponding association risk network based on the association risk weights includes: The risk association types between entities in the updated financial knowledge graph are determined, and the type weight coefficients of each risk association type and the entity risk score of each entity are determined by the analytic hierarchy process. Based on the type weight coefficients and the entity risk scores, the association risk weights between entities are determined; The corresponding associated risk network is constructed using the associated risk weights and the updated financial knowledge graph.

[0009] Optionally, the step of performing risk transmission calculation on the associated risk network using a preset risk transmission calculation model includes: The risk source node in the associated risk network is determined according to preset risk identification rules, and the risk transmission parameters of each node in the associated risk network are determined. The risk transmission parameters include risk attenuation coefficient, node risk carrying capacity, and risk transmission probability. The risk attenuation coefficient is determined based on the distance to the risk source node. The node risk carrying capacity is determined based on the size of the entity assets and credit status. The risk transmission probability is determined based on the associated risk weight. By using a preset risk transmission calculation model and based on the risk transmission parameters, the node risk value of each node in the associated risk network is iteratively calculated until a preset iteration termination condition is met, thereby obtaining the target node risk value of each node; the preset risk transmission calculation model is constructed based on a graph neural network; the preset iteration termination condition is that the change in the node risk value of each node is lower than a preset change threshold. Using the target shortest path algorithm and based on the risk value of the target node, the target risk transmission path in the associated risk network is identified, and the risk contribution of each node is determined to obtain the risk transmission calculation results.

[0010] Optionally, the step of identifying risks based on the obtained risk transmission calculation results and using a preset risk identification model to generate and output a corresponding risk identification report includes: Based on the target node risk value and the target risk transmission path, and by using a preset enterprise-related risk identification model, related risks are identified to obtain a first risk identification result; Using the target node risk value, the risk contribution rate, and the enterprise's historical default records, and through a preset credit default risk identification model, default risk is identified to obtain a second risk identification result; Generate and output a risk identification report that includes the first risk identification result, the second risk identification result, the target risk transmission path, and risk warning suggestions.

[0011] Secondly, this application provides a financial risk identification device based on dynamic knowledge graphs, applied to a target financial cloud platform, comprising: The data filtering module is used to collect target financial data in real time through a preset distributed data acquisition architecture, and to perform deduplication filtering on the target financial data to obtain an initial financial dataset. The knowledge fusion module is used to standardize the initial financial dataset and perform knowledge fusion through a target entity disambiguation algorithm to obtain a standardized financial dataset. The graph update module is used to update the preset financial knowledge graph based on the standardized financial dataset and using an incremental learning algorithm after determining that the preset update conditions are met, so as to obtain the updated financial knowledge graph. The weight determination module is used to mine the relationships between entities in the updated financial knowledge graph and determine the association risk weights, so as to construct a corresponding association risk network based on the association risk weights. The risk transmission calculation module is used to perform risk transmission calculation on the associated risk network using a preset risk transmission calculation model, and to identify risks based on the obtained risk transmission calculation results and a preset risk identification model, thereby generating and outputting a corresponding risk identification report.

[0012] Thirdly, this application provides an electronic device, comprising: Memory, used to store computer programs; A processor is used to execute the computer program to implement the aforementioned financial risk identification method based on dynamic knowledge graphs.

[0013] Fourthly, this application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned financial risk identification method based on dynamic knowledge graphs.

[0014] In this application, a pre-defined distributed data acquisition architecture is used to collect target financial data in real time, and the target financial data is deduplicated to obtain an initial financial dataset. The initial financial dataset is then standardized, and knowledge fusion is performed using a target entity disambiguation algorithm to obtain a standardized financial dataset. After determining that a pre-defined update condition is met, a pre-defined financial knowledge graph is updated based on the standardized financial dataset using an incremental learning algorithm to obtain an updated financial knowledge graph. For the updated financial knowledge graph, inter-entity association mining is performed, and association risk weights are determined to construct a corresponding association risk network based on these weights. A pre-defined risk transmission calculation model is used to perform risk transmission calculations on the association risk network. Based on the obtained risk transmission calculation results, a pre-defined risk identification model is used to identify risks, generating and outputting a corresponding risk identification report. As can be seen from the above, this application collects target financial data in real time through a pre-set distributed data acquisition architecture, obtains an initial financial dataset after deduplication and filtering, then performs standardization processing and knowledge fusion with the target entity disambiguation algorithm to form a standardized financial dataset. When the pre-set update conditions are met, the pre-set financial knowledge graph is updated using an incremental learning algorithm. Based on the updated financial knowledge graph, entity association mining and association risk weight determination are performed to construct an association risk network. The pre-set risk transmission calculation model is used to complete the risk transmission calculation, and the pre-set risk identification model is used to identify risks. Finally, a risk identification report is generated and output. In this way, through the above-described process of this application, the real-time collection and deduplication filtering using a pre-defined distributed data acquisition architecture can ensure the comprehensiveness and accuracy of financial data collection; knowledge fusion through standardized processing and entity disambiguation algorithms can improve data consistency and knowledge credibility, reducing redundancy and ambiguity; updating the knowledge graph using incremental learning algorithms can efficiently complete graph iteration without destroying the original knowledge structure, adapting to the dynamic changes of financial data; constructing a risk network through association mining and risk weighting can accurately depict the risk relationships between financial entities; and using risk transmission models and risk identification models for calculation and identification can quantify the risk propagation path and impact, improving the timeliness and reliability of financial risk identification, thereby improving the accuracy and interpretability of financial risk identification to support the needs of precise risk prevention and control. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0016] Figure 1This is a flowchart of a financial risk identification method based on dynamic knowledge graphs disclosed in this application; Figure 2 This is a flowchart illustrating a financial risk identification method based on dynamic knowledge graphs disclosed in this application. Figure 3 This is a flowchart illustrating a target entity disambiguation algorithm disclosed in this application; Figure 4 This is a flowchart illustrating the dynamic updating process of a financial knowledge graph as disclosed in this application. Figure 5 This is a flowchart illustrating a risk transmission calculation algorithm disclosed in this application; Figure 6 This is a schematic diagram of the structure of a financial risk identification device based on dynamic knowledge graph disclosed in this application; Figure 7 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Existing financial knowledge graphs are mostly statically constructed, relying on manual updates or fixed-period batch updates, resulting in insufficient timeliness. Furthermore, current technologies struggle to integrate heterogeneous data from multiple sources such as business registration, credit reporting, news, and transactions, leading to the omission of relevant risk nodes. In addition, traditional risk transmission calculations are often based on simple topological analysis, easily ignoring node heterogeneity and risk decay, failing to accurately pinpoint risk transmission paths and default probabilities, and prone to misjudgment or omission of risks. This results in poor accuracy and interpretability in financial risk identification, low efficiency, and an inability to support the needs of precise risk prevention and control.

[0019] To overcome the aforementioned technical problems, this application provides a financial risk identification method based on dynamic knowledge graphs, which can improve the accuracy and interpretability of financial risk identification to support the needs of precise risk prevention and control.

[0020] See Figure 1 As shown, this embodiment of the invention discloses a financial risk identification method based on dynamic knowledge graphs, applied to a target financial cloud platform, including: Step S11: Collect target financial data in real time through a preset distributed data acquisition architecture, and perform deduplication filtering on the target financial data to obtain an initial financial dataset.

[0021] In this embodiment, a preset distributed data acquisition architecture is used to collect target financial data in real time, and the collected target financial data is deduplicated to obtain an initial financial dataset.

[0022] It should be noted that this application provides a financial risk identification method based on dynamic knowledge graphs to address the pain points of lacking dynamic update algorithms, insufficient cross-domain data fusion capabilities, and low adaptability of risk transmission models to financial scenarios, making it difficult to support the needs of precise risk prevention and control, and failing to meet the real-time requirements of high-frequency scenarios such as credit approval and risk monitoring. Figure 2The diagram illustrates the flowchart of a financial risk identification method based on dynamic knowledge graphs provided in this application. The technical solution includes a cross-domain financial data acquisition module: This module integrates multi-source data crawlers and API calls, covering structured, semi-structured, and unstructured data to achieve real-time data capture and incremental acquisition. By using preset data filtering rules, invalid data is filtered out, and initial data deduplication is completed, providing a high-quality data source for subsequent processing. A data standardization and knowledge fusion module is also included: This module designs data mapping rules specific to the financial field, normalizes multi-source data, and unifies the naming conventions for entities and relationships. An improved entity disambiguation algorithm is used to resolve conflicts between entities with the same name, achieving accurate fusion of knowledge from cross-data sources and constructing a basic knowledge network. Finally, a dynamic update algorithm module is proposed: This module proposes a dynamic update algorithm based on time-series triggering and incremental learning, setting data update thresholds and trigger conditions to achieve real-time incremental updates of entities, relationships, and attributes, while retaining historical versions and supporting graph backtracking analysis, thus addressing the problem of insufficient timeliness in static graphs. The module for constructing an associated risk network, based on a fused knowledge graph, mines potential relationships between entities, defines rules for calculating risk association weights, and constructs an associated risk network including nodes such as enterprises, institutions, and individuals. It quantifies the risk transmission capacity between nodes, providing a foundation for risk calculation. The module for calculating risk transmission paths designs a multi-dimensional risk transmission calculation model, introducing parameters such as risk attenuation coefficients and node risk carrying capacity. Combined with graph neural network algorithms, it accurately calculates the risk transmission path, intensity, and scope of impact, enabling visualized tracking of risk transmission. The module for risk identification and prevention output builds enterprise-related risk and credit default risk identification models based on the risk transmission calculation results. It outputs risk levels and early warning signals, connects to the risk control system, and supports precise prevention and control decisions in scenarios such as credit approval and risk monitoring. These modules are deployed on a financial cloud platform, employing a distributed architecture to ensure system stability and supporting a processing capacity of over 1000 data entries per second. An operation and maintenance monitoring module is set up to monitor the operating status of each module in real time, automatically triggering alarms when problems such as data collection failures or algorithm iteration anomalies occur. Data mapping rules, risk weight coefficients, and algorithm model parameters are regularly updated to adapt to changes in financial regulatory policies and market dynamics, ensuring the long-term applicability of the method. Breaking away from the limitations of traditional financial knowledge graphs, this technology achieves deep adaptation and innovative expansion across multiple scenarios. First, it expands from a single credit risk control scenario to a full-process risk prevention and control scenario, covering the entire chain from credit approval, loan monitoring, post-loan handling, and news risk warnings, solving the problem that traditional technologies are only applicable to a single stage. Second, it adapts to the risk prevention and control scenario for SMEs. Addressing the characteristics of SMEs' dispersed data and complex relationships, it achieves accurate identification of SME-related risks and default risks through implicit correlation mining and lightweight algorithm design, filling the gap in the insufficient adaptation of existing technologies for SME risk control.Third, it adds cross-institutional risk prevention scenarios, supports knowledge graph data sharing among multiple financial institutions (based on privacy computing technology), builds a cross-institutional risk network, and realizes cross-institutional risk transmission monitoring, solving the problem that traditional technologies are limited to risk control within a single institution. Fourth, it adapts to regulatory compliance scenarios, retains full-link data on risk transmission and the history of graph updates, meets the traceability and auditing requirements of regulatory authorities, and is more in line with financial regulatory needs than traditional technologies, improving risk control compliance.

[0023] Specifically, a first set of financial data is crawled using a multi-threaded crawler configured based on the Scrapy framework (an application framework for web crawling that quickly captures website data and extracts structured data from pages); a second set of financial data is obtained by calling a preset API (Application Programming Interface) and using a preset encrypted transmission protocol, thus obtaining target financial data including the first and second sets of financial data; financial data in the target financial data that meets preset invalidity conditions is filtered according to preset filtering rules, and duplicate data is removed based on a hash algorithm to obtain an initial financial dataset. The first set of financial data includes unstructured / semi-structured data such as business registration information (company shareholders, executives, business scope), news data (financial news, negative information about companies), and industry dynamics; the second set of financial data includes structured data such as bank credit data, credit reports, and transaction records. That is, this embodiment establishes a preset distributed data acquisition architecture, including two data acquisition methods: crawling public financial data using crawler technology and calling private financial data through API interfaces. Specifically, a multi-threaded crawler is configured based on the Scrapy framework, and an IP (Internet Protocol) proxy pool is set up to avoid anti-crawling measures. The crawling frequency can be dynamically adjusted according to the data update frequency (e.g., news data is crawled once every 10 minutes, and business data is crawled once a day) to crawl publicly available first financial data. Pre-set API interfaces are called to connect to the central bank's credit reporting system, bank core systems, etc., and data security is ensured through encrypted transmission protocols to obtain private second financial data from the connected systems. The target financial data is then integrated, and invalid data (such as blank values, duplicate data, and non-financial data) is filtered according to preset filtering rules. Duplicate data is removed using a hash algorithm, and finally, the initial financial dataset D is obtained, where D={D1,D2,D3}, D1 is structured financial data, D2 is semi-structured financial data, and D3 is unstructured financial data.

[0024] It's important to note that the core principle of using hash algorithms to remove duplicate data is to map data to a fixed-length hash value using a hash function, and then compare these hash values ​​to determine if the data is duplicated. The basic process is as follows: calculate a hash value for each piece of data and store it in a hash table (or a similar structure); when new data arrives, calculate its hash value; if the value already exists in the hash table, it is considered a duplicate; otherwise, it is added to the hash table and treated as new data. In this way, the distributed data acquisition architecture built in this embodiment enables parallel acquisition of financial data from multiple sources, improving data acquisition efficiency and coverage. Deduplication filtering of financial data effectively removes duplicate and invalid data, ensuring the simplicity and reliability of the initial financial dataset and providing a high-quality data foundation for subsequent data processing. The use of the Scrapy framework to configure multi-threaded crawlers provides base classes for various types of crawlers, facilitating modification according to needs and improving the efficiency and stability of financial data crawling, covering more data sources. Utilizing encrypted transmission protocols to obtain API interface data ensures the security and integrity of financial data transmission. By filtering invalid data through preset filtering rules and removing duplicate data using hash algorithms, redundant and invalid information can be accurately cleaned up, resulting in high processing efficiency, simple implementation, and ensuring the accuracy and purity of the initial financial dataset.

[0025] Step S12: Standardize the initial financial dataset and perform knowledge fusion using a target entity disambiguation algorithm to obtain a standardized financial dataset.

[0026] In this embodiment, the initial financial dataset is standardized, and a target entity disambiguation algorithm is used to complete knowledge fusion to obtain a standardized financial dataset.

[0027] It is understood that the initial financial dataset includes structured financial data, semi-structured financial data, and unstructured financial data. Correspondingly, the processing flow for obtaining a standardized financial dataset is as follows: Field mapping is used to map the fields in the structured financial data to standard fields in the target financial domain, and the numerical data in the structured financial data is normalized to obtain first standardized data; the semi-structured financial data is parsed using XML path language to convert it into a target structured format to obtain second standardized data; text segmentation and part-of-speech tagging are performed on the unstructured financial data based on a preset natural language processing model to obtain third standardized data; the target similarity between entities in the target standardized data is determined, and the target similarity is compared with a preset similarity threshold. Knowledge fusion is performed based on the comparison results to obtain the standardized financial dataset; the target similarity is determined based on entity name similarity, attribute similarity, and contextual similarity. That is, the initial financial datasets containing structured, semi-structured, and unstructured data are processed separately. For the structured financial data D1, fields from different data sources are uniformly mapped to standard fields in the target financial domain through field mapping (e.g., "loan amount" and "loan limit" are uniformly mapped to "credit limit"). Then, numerical normalization is performed using the Z-score (Z score / standard score) standardization formula. The standard deviation of the data itself measures the distance of each data point from its mean, thereby eliminating the influence of the original data's units and making it comparable, resulting in the first standardized data. The Z-score standardization formula is as follows: ; in, Here, x represents the original value, μ is the field mean, and σ is the field standard deviation. The basic idea of ​​Z-score standardization is to determine how many standard deviations a given value is from the mean of its dataset. Through this transformation, the original data is mapped to a new distribution with a mean of 0 and a standard deviation of 1. Specifically, the arithmetic mean (μ) and standard deviation (σ) of the data column to be processed need to be calculated. It's important to note that theoretically, population parameters should be used, but in practice, the sample mean and (sample) standard deviation are usually used for estimation. Then, for each original value (x) in the dataset, the Z-score standardization formula is substituted into the above formula for calculation. After the calculation, the original dataset is transformed into a new Z-score dataset. For the semi-structured financial data D2 (such as business registration information in HTML format), key information is extracted using XPath (XML Path Language, a declarative query language). Target nodes in the HTML / XML document tree (DOM tree, Document Object Model) are precisely located using path expressions, and their text or attribute values ​​are extracted. Finally, the extracted semi-structured information is organized into a regular table or database record, converted to a structured format, and the second standardized data is obtained. The core principle of this processing is as follows: a parser is used to convert the data model from HTML to the DOM tree. By writing path expressions, rules for locating nodes in the DOM tree are specified for declarative access and path navigation. The parser automatically traverses the DOM tree according to the specified expressions, finding all matching node sets. During the processing, it can flexibly handle semi-structured data by simultaneously satisfying both fuzzy search and precise positioning requirements. For the unstructured financial data D3 (such as news text), a pre-set natural language processing model, such as a pre-trained BERT (Bidirectional Encoder Representation from...), is used. Transformers (a type of pre-trained language model) are powerful semantic encoders that output context-aware vector representations. These vector representations are then fed into downstream task layers (such as fully connected layers) to map the original text to structured labels (word segmentation boundaries, parts of speech, entity categories). This process involves text segmentation, part-of-speech tagging, extraction of key entities and attributes to obtain third-order standardized data, and finally, entity matching and knowledge fusion based on entity names, attributes, and contextual similarity to resolve conflicts between entities with the same name and obtain a standardized financial dataset.

[0028] It should be noted that, for D3, this embodiment uses BERT as a feature extractor: by performing pre-training tasks such as masked language modeling and next-sentence prediction on a large-scale corpus, it learns deep bidirectional language representations, meaning that for each character or subword in the input text, BERT can generate a dense vector (embedding) that incorporates its entire context information. This vector captures semantic and syntactic information better than traditional word vectors (such as Word2Vec) or rule-based models. Sequence labeling task framework: Word segmentation, part-of-speech tagging, and named entity recognition (used to extract key entities and attributes) are essentially sequence labeling tasks. Their goal is to predict a label for each unit in the input sequence (such as a sequence of Chinese characters). For example, in Chinese word segmentation, the label can be "B" (word beginning), "I" (word middle), "E" (word ending), "S" (single character word); in part-of-speech tagging, the label is various parts of speech (such as nouns, verbs); in entity recognition, the label is entity type (such as person name, place, organization) and location information (such as B-PER, I-PER). The "pre-training + fine-tuning" paradigm: In practice, instead of training a model from scratch, a task-specific output layer (usually a linear transformation layer) is added to a pre-trained BERT model, and then fine-tuning is performed on labeled target domain data (such as labeled news text). In this way, the model can utilize the general language knowledge learned by BERT while adapting to the labeling system of a specific task.

[0029] It should be further noted that an entity may have multiple referents in text or different data sources, and these referents may point to multiple candidate entities in the knowledge base (i.e., conflicting entities with the same name). To resolve this ambiguity, the system needs to calculate the similarity between the referent and the candidate entity on multi-dimensional features, and determine whether they point to the same entity based on the comprehensive similarity, thereby resolving data redundancy and conflicts and forming a standardized knowledge base. This embodiment proposes an improved entity disambiguation algorithm that integrates three key features: entity name similarity: based on the string itself, such as edit distance, Jaccard coefficient (an index that measures the similarity between two sets), to solve problems such as spelling differences and abbreviations; attribute similarity: comparing the structured attributes of entities (such as birthday, location), commonly using methods such as edit distance, set similarity, or cosine similarity; context similarity: judging by analyzing the semantic relationship between the text surrounding the entity referent (context) and the candidate entity description text. Figure 3 The diagram shown is a flowchart of a target entity disambiguation algorithm provided in this application. Similarity is calculated using the cosine similarity formula, as follows: ; Where a and b are the feature vectors of the entities, and ||·|| is the norm operation. Before calculating similarity, data preprocessing and normalization are performed, cleaning and normalizing data from different sources, including: removing redundant symbols, correcting input errors, replacing nicknames and abbreviations with formal names, and standardizing the format of attributes such as dates / addresses. For the entity references to be disambiguated, a set of possible candidate entities is retrieved from the target knowledge base. Common methods include: using name dictionaries (such as disambiguation pages), string matching-based search, or using search engine APIs. For calculating entity name similarity, edit distance or the Jaccard coefficient (comparing the overlap of character n-gram (a statistical language model) sets) can also be used. For example, calculating the similarity between Steve Jobs and Steve_Jobs. An example formula (Jaccard) is shown below: Similarity = |A∩B| / |A∪B|; Here, A and B are two sets of words or n-grams after name segmentation, ∩ represents the intersection, and |·| represents the number of elements in the set. For attribute similarity calculation: for numerical or categorical attributes, direct comparison is possible; for text attributes (such as descriptions), they can be converted into word vectors, and then cosine similarity is calculated. The closer this value is to 1, the more consistent the directions of the two vectors, and the more semantically similar they are. For context similarity calculation: the context text referring to the entity (such as the sentence or paragraph it belongs to) and the descriptive text of the candidate entity (such as a summary in a knowledge base) are vectorized separately. Vectorization tools can use traditional TF-IDF (Term Frequency-Inverse Document Frequency, a weighted technique used for information retrieval and text mining), or deep learning models such as Word2Vec (a method for generating word embeddings) and BERT, and then the cosine similarity of the two text vectors is calculated. For example, using the BERT model to encode two text segments into fixed-dimensional vectors, and then calculating the cosine value as a measure of contextual semantic similarity. Assign weights (w1, w2, w3) to name, attribute, and context similarity, then sum them using a weighted average to obtain the final overall similarity score: Overall Similarity = w1 Name similarity + w2 Attribute similarity + w3 Contextual similarity. When the overall similarity exceeds a preset threshold, such as 0.8, the entities are considered to be the same. Alternatively, classifiers (such as logistic regression or neural networks) can be used to learn how to combine these features and determine whether they are the same entity. After determining that they are the same entity, their attributes and relationships are merged. This may involve conflict resolution. Attribute conflict resolution strategies include: confidence weight allocation (preferring authoritative data sources), majority voting, or retaining all values ​​and noting their sources. Relationship merging includes merging relationships from different sources to enrich the entity's association information. Finally, a standardized financial dataset S is obtained, with a unified dataset structure and consistent semantics, providing a foundation for knowledge graph construction. In this way, the standardization of the initial financial dataset in this embodiment can unify the data format and dimensions, and eliminate the calculation bias caused by data heterogeneity; the use of target entity disambiguation algorithm for knowledge fusion can solve the problems of synonymy and heteronymy of financial entities, integrate scattered financial information, form a standardized and unified dataset, and improve the standardization and knowledge consistency of the dataset; the use of categorized differentiated processing can adapt to the characteristics of financial data with different structures, achieve unified standardization of heterogeneous data, and improve data compatibility; and the combination of multi-dimensional similarity calculation based on name, attribute, and context can improve the accuracy of entity matching and reduce ambiguity and redundancy.

[0030] Step S13: After determining that the preset update conditions are met, update the preset financial knowledge graph based on the standardized financial dataset and using an incremental learning algorithm to obtain the updated financial knowledge graph.

[0031] In this embodiment, after the preset update conditions are met, the original preset financial knowledge graph is updated using an incremental learning algorithm based on the standardized financial dataset to obtain the updated financial knowledge graph.

[0032] Specifically, the current financial data increment is determined based on the standardized financial dataset. If the financial data increment exceeds a preset data increment threshold, or if the target key data is updated, or if a preset update time period is reached, then the newly added entity set and the newly added entity relationship set of the financial data increment are determined. An incremental learning algorithm is used, and based on the newly added entity set and the newly added entity relationship set, the preset financial knowledge graph is updated to obtain the updated financial knowledge graph. The preset financial knowledge graph is constructed based on the preset standardized financial knowledge dataset and includes an entity set and an entity relationship set. That is, in this embodiment, a basic preset financial knowledge graph G=(V,E) is first constructed based on the preset standardized financial knowledge dataset, where V is the entity set (enterprises, individuals, financial products, institutions, etc.), and E is the entity relationship set (equity relationships, credit relationships, guarantee relationships, cooperative relationships, etc.). It is stored using the Neo4j graph database (a native graph database) to improve graph query and calculation efficiency. Subsequently, based on the standardized financial dataset, the current financial data increment is determined. When the data increment exceeds a threshold (e.g., ≥50 new records for a single data category), target key data is updated (e.g., changes in corporate equity, new credit default records), or a preset time period is reached (e.g., full verification is performed daily at midnight), the set of newly added entities and the set of relationships between these entities are extracted. An incremental learning algorithm is then used to quickly process the new data. This eliminates the need to reconstruct the entire graph; only the newly added entities, relationships, and attributes are updated, thus updating the original financial knowledge graph and obtaining the updated financial knowledge graph. For example... Figure 4 The diagram illustrates a process for dynamically updating a financial knowledge graph as provided in this application. The specific update formula is as follows: ; in, The knowledge graph at time t. The graph at time t-1 To add a new entity set, This involves adding a set of relationships between new entities. In addition, this embodiment establishes a graph version management mechanism, retaining historical versions of each update and supporting time-based backtracking queries for easy risk tracing analysis. Thus, this embodiment updates the financial knowledge graph through incremental learning algorithms, efficiently integrating new data while preserving the original knowledge structure, avoiding full reconstruction, achieving real-time dynamic updates of the knowledge graph, improving update efficiency and stability, and ensuring the real-time and continuous nature of financial knowledge, ensuring the graph data is synchronized with changes in the financial market; multi-condition triggered updates can balance data volume, importance, and timeliness, making knowledge graph updates more intelligent and reasonable; using the Neo4j graph database to store the knowledge graph allows graph traversal operations to be executed efficiently at the storage level, avoiding complex join operations in relational databases. Furthermore, Neo4j uses an index-free neighbor node traversal query algorithm, so the search process is not affected by the data size, resulting in high query performance. Simultaneously, Neo4j's graph computing engine can directly execute algorithms on the stored graph structure, reducing the overhead of data transformation and movement.

[0033] Step S14: Perform entity association mining on the updated financial knowledge graph and determine association risk weights to construct a corresponding association risk network based on the association risk weights.

[0034] In this embodiment, the updated financial knowledge graph is dynamically updated to mine direct and potential relationships between entities, determine the associated risk weights corresponding to each relationship, and construct an associated risk network based on the associated risk weights.

[0035] Specifically, the risk association types between entities in the updated financial knowledge graph are determined, and the type weight coefficients of each risk association type and the entity risk score of each entity are determined using the Analytic Hierarchy Process (AHP). Based on the type weight coefficients and the entity risk scores, the association risk weights between entities are determined. A corresponding association risk network is constructed using the association risk weights and the updated financial knowledge graph. The risk association types include direct associations (explicit relationships such as equity, guarantees, and credit) and potential associations (implicit relationships such as shared executives, related-party transactions, and operations in the same region). That is, the risk association types between entities in the updated financial knowledge graph are first determined, and then the type weight coefficients of each risk association type and the basic entity risk scores of each entity are obtained using the Analytic Hierarchy Process (AHP). It is understood that different types of associations (such as equity associations, transaction associations, and guarantee associations) have different degrees of impact on risk transmission. AHP can systematically quantify the judgments of decision-makers (such as risk control experts) on the importance of different association types, thereby assigning more reasonable weights to the "edges" in the network. Specifically, the vague objective of assessing the importance of association types to risk transmission is decomposed into a clear structure of target layer, criterion layer, and alternative layer, forcing analysts to systematically consider all relevant association types. AHP transforms the complex global ranking problem into a series of simple local judgments by requiring experts to compare the importance of any two association types. These qualitative comparisons are then converted into a quantitative judgment matrix, and the weight coefficients of each association type are obtained by calculating eigenvectors. The built-in consistency check is a key quality control step, verifying whether there are logical contradictions in expert judgments (e.g., believing A is more important than B, B is more important than C, but also believing C is more important than A), ensuring the scientific validity and credibility of the weight coefficients. Based on the type weight coefficients and the entity risk scores, the association risk weights between entities are calculated. The specific formula for calculating the associated risk weight is as follows: ; in, The weight coefficient for the k-th class association. , These are the basic entity risk scores for entities i and j (score range 0-100, the higher the score, the higher the risk). The association risk weights characterize the relationship between entities i and j. Finally, combining these association risk weights with the updated financial knowledge graph, an association risk network G'=(V,E,W) is constructed, where W is the set of association risk weights. This completes the quantitative construction of the risk network, providing a carrier for risk transmission calculation. In this way, this embodiment can accurately capture potential relationships between financial entities by mining the relationships between entities in the updated financial knowledge graph, breaking down information silos; determining association risk weights and constructing an association risk network can quantify abstract relationships into analyzable risk indicators, forming a quantifiable and resolvable risk transmission structure, clearly presenting the risk transmission path between financial entities, and improving the accuracy and credibility of subsequent financial risk analysis; using the analytic hierarchy process (AHP) to quantify the risk association type weights and entity risk scores makes risk assessment more logical and objective; calculating association risk weights based on type weight coefficients and entity risk scores can accurately characterize the strength of risk relationships between financial entities.

[0036] Step S15: Perform risk transmission calculation on the associated risk network using a preset risk transmission calculation model, and based on the obtained risk transmission calculation results, perform risk identification through a preset risk identification model to generate and output a corresponding risk identification report.

[0037] In this embodiment, a preset risk transmission calculation model is used to perform risk transmission calculation on the associated risk network, and a preset risk identification model is used to identify risks based on the risk transmission calculation results, thereby generating and outputting a risk identification report.

[0038] It should be noted that the processing flow for calculating risk transmission in the associated risk network using a preset risk transmission calculation model is as follows: Based on preset risk identification rules, the risk source node in the associated risk network is determined, and the risk transmission parameters of each node in the associated risk network are determined. The risk transmission parameters include a risk attenuation coefficient, node risk carrying capacity, and risk transmission probability. The risk attenuation coefficient is determined based on the distance to the risk source node. The node risk carrying capacity is determined based on the size of the entity's assets and credit status. The risk transmission probability is determined based on the associated risk weight. Using the preset risk transmission calculation model and based on the risk transmission parameters, the node risk value of each node in the associated risk network is iteratively calculated until a preset iteration termination condition is met, obtaining the target node risk value for each node. The preset risk transmission calculation model is constructed based on a graph neural network. The preset iteration termination condition is that the change in the node risk value of each node is lower than a preset change threshold. Using a target shortest path algorithm and based on the target node risk value, the target risk transmission path in the associated risk network is identified, and the risk contribution of each node is determined to obtain the risk transmission calculation result. That is, as... Figure 5The diagram shown is a flowchart of a risk transmission calculation algorithm provided in this application. Based on preset risk identification rules (such as entities experiencing credit defaults or negative news outbreaks), the algorithm locates the risk source nodes in the associated risk network. Determine the risk attenuation coefficient (λ∈(0,1), the attenuation is more obvious with increasing distance), node risk carrying capacity (Determined based on the size of the entity's assets and credit history) ∈(0,100)), risk transmission probability The risk transmission parameters are composed of associated risk weights, where the risk transmission probability is determined by the associated risk weights. The conversion is as follows: ; It is understandable that the essence of risk transmission is the spread of risk events through the relationships between entities. Graph Neural Networks (GNNs) are ideal tools for processing such graph-structured data. They can abstract complex risk transmission networks (such as financial networks and supply chain networks) into graph-structured data. Utilizing the graph structure learning and information transmission capabilities of GNNs, the propagation process of risk along the network path can be accurately quantified. The transmission path, intensity, and scope of impact of risk from the source node to the target node can be calculated. Each node in the network (such as enterprises and financial institutions) aggregates information from its neighboring nodes to update its own state representation. Therefore, through a pre-defined risk transmission calculation model based on graph neural networks, using multi-layer graph convolution or graph attention operations, the multi-hop propagation of risk information in the network is simulated, and the node risk value of each node is iteratively calculated. At each layer, the node receives information from its direct neighbors and updates it based on its own characteristics (such as financial status and risk level). The specific iterative calculation formula is as follows: ; in, Let i be the risk value of node i in iteration t. Let N(i) be the risk value of node j in iteration t, and N(i) be the set of neighboring nodes of node i, with the initial risk value... (Risk saturation at the source node), with initial risk values ​​of 0 for other nodes. Iteration stops when the change in risk value for all nodes is ≤0.01, yielding the target node risk value for each node. Finally, a shortest path algorithm, such as Dijkstra's algorithm, is used to identify the core risk transmission path and calculate the node risk contribution. The complex risk network is abstracted into a weighted directed graph. By calculating the shortest or least-resistance path, the most likely and effective risk propagation channel is identified, and the risk contribution of each node on the path is quantified. Paths with a risk transmission strength ≥30 are retained, resulting in a risk transmission calculation result that includes the risk transmission link and the risk contribution of each node.

[0039] It should be further noted that the process of generating and outputting a corresponding risk identification report based on the obtained risk transmission calculation results and through a preset risk identification model is as follows: Based on the target node risk value and the target risk transmission path, related risks are identified through a preset enterprise-related risk identification model to obtain a first risk identification result; using the target node risk value, the risk contribution degree, and the enterprise's historical default records, default risks are identified through a preset credit default risk identification model to obtain a second risk identification result; a risk identification report including the first risk identification result, the second risk identification result, the target risk transmission path, and risk warning suggestions is generated and output. In other words, this embodiment constructs a dual risk identification model that includes enterprise-related risk identification and credit default risk identification. Based on the target node risk value and the target risk transmission path, a first risk identification result is obtained through a preset enterprise-related risk identification model. For example, when the number of risk nodes in an enterprise's related nodes is ≥3 or the total related risk value is ≥150, it is determined to be a high-related-risk enterprise. Combining the enterprise's target node risk value, risk contribution, and historical default records, a preset credit default risk identification model, such as a logistic regression model, is used. A non-linear transformation is performed through a logistic function (Sigmoid function) to output a probability value between 0 and 1, which is used to quantify the possibility of the enterprise defaulting, i.e., the default probability. The specific formula for calculating the default probability is as follows: ; Where y=1 represents default, x is the feature vector (risk value, associated risk weight, etc.), w is the feature coefficient, and b is the bias term. When the predicted probability is ≥0.6, it is judged as high default risk, and the second risk identification result is obtained. Finally, a risk identification report containing two types of identification results (list of risky enterprises, risk level (high / medium / low)), risk transmission path and early warning suggestions is generated and output. It is connected to the risk control system of financial institutions to realize precise prevention and control operations such as risk warning and credit approval interception. At the same time, the risk data is fed back to the data collection module to optimize data screening and risk identification rules, forming a closed loop.

[0040] It is understood that the performance data of the method in this application are all obtained through standard experimental comparisons. The experiments use real multi-source data from the financial field as samples and combine them with existing financial knowledge graph technology test benchmarks to ensure that the results are rigorous and reliable. The specific experimental comparison data and data sources are as follows: Experimental Testing Background: Three standard datasets in the financial field were selected (containing over 300,000 entities and over 100,000 relationships, covering various data categories including business registration, credit reporting, and news; the data scale closely matches the actual application scenarios of financial institutions). The method described in this application was compared with existing mainstream dynamic update algorithms (FastKGE incremental update algorithm) and traditional entity disambiguation algorithms (based solely on name similarity) for comparative testing. The test environment consisted of an Intel Xeon E5-2690 CPU, 64GB of memory, a 1TB hard drive, and CentOS 7.9 operating system. All algorithms were developed using Python 3.8 to ensure a consistent testing environment and eliminate the influence of environmental differences on the experimental results. The experimental data came from a self-built multi-source financial data test set and publicly available financial knowledge graph evaluation datasets. The data rationality was also verified by combining existing research results in related technologies in the financial field. Table 1 shows the experimental test comparison table.

[0041] Table 1 Comparison of Experimental Tests

[0042] Based on the experimental data above, the core technological innovations of this method are as follows: First, it proposes a time-triggered incremental dynamic update technology, which differs from traditional fixed-period updates or manual updates. By combining data increment thresholds and key event triggers, it achieves minute-level incremental updates of the graph while retaining historical versions, addressing the core pain point of insufficient timeliness of static graphs. Its update efficiency and resource consumption are superior to existing mainstream dynamic update algorithms. Second, it innovates a cross-domain knowledge fusion mechanism, designing financial-specific data mapping rules and an improved entity disambiguation algorithm. This breaks through the limitations of a single data source, achieving accurate fusion of structured and unstructured data, resolving conflicts between entities with the same name and the omission of related nodes. The fusion accuracy reaches over 95%, significantly better than the average level of existing financial entity disambiguation algorithms (approximately 92%). Third, it constructs a quantitative correlation risk network, introducing the analytic hierarchy process (AHP) to determine correlation weights, enabling the mining and quantification of implicit correlations, thus overcoming the shortcomings of existing technologies that only focus on explicit correlations. Fourth, a multi-parameter fusion risk transmission calculation model is proposed, which integrates parameters such as risk decay and node carrying capacity, and combines them with an improved GNN algorithm to achieve accurate calculation of risk transmission path and intensity. This solves the problem that traditional topology analysis ignores node heterogeneity and risk decay, and improves the accuracy and interpretability of risk identification.

[0043] In this way, this embodiment leverages a graph neural network-based risk transmission calculation model to iteratively converge and calculate node risk values, quantifying the risk diffusion effect. This allows the model to capture long-chain risk transmission effects beyond direct correlations, accurately simulating the propagation path and impact of risk among financial entities. Intelligent identification through a risk identification model automatically locates key risk points, improving the efficiency and accuracy of financial risk identification, ultimately generating standardized reports and providing a reliable basis for risk management. Employing multi-dimensional risk transmission parameters accurately reflects the actual patterns of risk transmission among financial entities, enhancing computational rationality. Using the shortest path algorithm to identify transmission paths and contributions clearly identifies key risk links and core impact nodes, efficiently and accurately pinpointing the key links with the greatest impact on systemic risk, providing precise support for subsequent risk identification and management. The separate models specifically identify associated risks and credit default risks, achieving refined and comprehensive risk identification and avoiding the limitations of single-model identification. Combining historical default records improves the accuracy of default risk identification. Based on the identification results and early warning suggestions output by the dual models, specific and implementable evidence can be provided for risk management, while clearly presenting the source and transmission path of risk, making risk management more targeted and forward-looking.

[0044] As can be seen from the above, the embodiments of this application collect target financial data in real time through a preset distributed data acquisition architecture, obtain an initial financial dataset after deduplication and filtering, and then perform standardization processing and knowledge fusion with the target entity disambiguation algorithm to form a standardized financial dataset. When the preset update conditions are met, the preset financial knowledge graph is updated using an incremental learning algorithm. Based on the updated financial knowledge graph, entity association mining and association risk weight determination are performed to construct an association risk network. The preset risk transmission calculation model is used to complete the risk transmission calculation, and the preset risk identification model is used to identify risks. Finally, a risk identification report is generated and output. In this way, through the above-described process of the embodiments of this application, the real-time collection and deduplication filtering using a pre-defined distributed data acquisition architecture can ensure the comprehensiveness and accuracy of financial data collection; knowledge fusion through standardized processing and entity disambiguation algorithms can improve data consistency and knowledge credibility, and reduce redundancy and ambiguity; updating the knowledge graph using incremental learning algorithms can efficiently complete graph iteration without destroying the original knowledge structure, adapting to the dynamic changes of financial data; constructing a risk network through association mining and risk weighting can accurately depict the risk relationships between financial entities; and using risk transmission models and risk identification models for calculation and identification can quantify the risk propagation path and impact, improve the timeliness and reliability of financial risk identification, and thus improve the accuracy and interpretability of financial risk identification to support the needs of precise risk prevention and control.

[0045] Accordingly, see Figure 6As shown, this application embodiment also provides a financial risk identification device based on a dynamic knowledge graph, applied to a target financial cloud platform, including: The data filtering module 11 is used to collect target financial data in real time through a preset distributed data acquisition architecture, and to perform deduplication filtering on the target financial data to obtain an initial financial dataset. The knowledge fusion module 12 is used to standardize the initial financial dataset and perform knowledge fusion through a target entity disambiguation algorithm to obtain a standardized financial dataset. The graph update module 13 is used to update the preset financial knowledge graph based on the standardized financial dataset and using an incremental learning algorithm after determining that the preset update conditions are met, so as to obtain the updated financial knowledge graph. The weight determination module 14 is used to mine the relationships between entities in the updated financial knowledge graph and determine the association risk weights, so as to construct a corresponding association risk network based on the association risk weights. The risk transmission calculation module 15 is used to perform risk transmission calculation on the associated risk network using a preset risk transmission calculation model, and to perform risk identification based on the obtained risk transmission calculation results and a preset risk identification model, thereby generating and outputting a corresponding risk identification report.

[0046] In some specific embodiments, the data filtering module 11 may specifically include: The data crawling unit is used to crawl first financial data using a multi-threaded crawler configured based on the Scrapy framework. The data acquisition unit is used to call a preset API interface and use a preset encrypted transmission protocol to acquire second financial data, so as to obtain target financial data including the first financial data and the second financial data; The data removal unit is used to filter financial data that meets the preset invalid conditions in the target financial data according to preset filtering rules, and to remove duplicate data based on a hash algorithm to obtain an initial financial dataset.

[0047] In some specific implementations, the initial financial dataset includes structured financial data, semi-structured financial data, and unstructured financial data; Accordingly, the knowledge fusion module 12 may specifically include: The normalization processing unit is used to map the fields in the structured financial data to standard fields in the target financial field using the field mapping method, and to normalize the numerical data in the structured financial data to obtain the first standardized data. The data parsing unit is used to parse the semi-structured financial data using XML path language, so as to convert the semi-structured financial data into a target structured format and obtain the second standardized data. The part-of-speech tagging unit is used to perform text segmentation and part-of-speech tagging on the unstructured financial data based on a preset natural language processing model to obtain third standardized data. The similarity comparison unit is used to determine the target similarity between entities in the target standardized data, and compare the target similarity with a preset similarity threshold to perform knowledge fusion based on the comparison results to obtain a standardized financial dataset; the target similarity is determined based on the entity's name similarity, attribute similarity and contextual similarity.

[0048] In some specific embodiments, the map update module 13 may specifically include: An incremental determination unit is used to determine the current financial data increment based on the standardized financial dataset. The set determination unit is used to determine the set of newly added entities and the set of relationships between newly added entities in the financial data increment if the increment of the financial data is greater than a preset data increment threshold, or if the target key data is updated, or if a preset update time period is reached. The graph update unit is used to update the preset financial knowledge graph using an incremental learning algorithm and based on the newly added entity set and the newly added entity relationship set, so as to obtain the updated financial knowledge graph; the preset financial knowledge graph is constructed based on a preset standardized financial knowledge dataset, including an entity set and an entity relationship set.

[0049] In some specific embodiments, the weight determination module 14 may specifically include: The scoring determination unit is used to determine the risk association types between entities in the updated financial knowledge graph, and to determine the type weight coefficients of each risk association type and the entity risk score of each entity through the analytic hierarchy process. The weight determination unit is used to determine the association risk weight between entities based on the type weight coefficient and the entity risk score; The network construction unit is used to construct a corresponding associated risk network using the associated risk weights and the updated financial knowledge graph.

[0050] In some specific embodiments, the risk transmission calculation module 15 may specifically include: The parameter determination unit is used to determine the risk source node in the associated risk network according to a preset risk identification rule, and to determine the risk transmission parameters of each node in the associated risk network; the risk transmission parameters include a risk attenuation coefficient, a node risk carrying capacity, and a risk transmission probability; the risk attenuation coefficient is determined based on the distance to the risk source node; the node risk carrying capacity is determined based on the size of the entity assets and credit status; the risk transmission probability is determined based on the associated risk weight. An iterative calculation unit is used to iteratively calculate the node risk value of each node in the associated risk network based on a preset risk transmission calculation model and the risk transmission parameters, until a preset iteration termination condition is met, thereby obtaining the target node risk value of each node; the preset risk transmission calculation model is constructed based on a graph neural network; the preset iteration termination condition is that the change in the node risk value of each node is lower than a preset change threshold. The path identification unit is used to identify the target risk transmission path in the associated risk network using the target shortest path algorithm and based on the risk value of the target node, and to determine the risk contribution of each node in order to obtain the risk transmission calculation result.

[0051] In some specific embodiments, the risk transmission calculation module 15 may specifically include: The first risk identification unit is used to identify associated risks based on the target node risk value and the target risk transmission path and through a preset enterprise associated risk identification model to obtain the first risk identification result; The second risk identification unit is used to identify default risk by using the target node risk value, the risk contribution degree and the enterprise's historical default records and by using a preset credit default risk identification model, so as to obtain the second risk identification result. The report output unit is used to generate and output a risk identification report that includes the first risk identification result, the second risk identification result, the target risk transmission path, and risk warning suggestions.

[0052] Furthermore, embodiments of this application also disclose an electronic device, Figure 7 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the financial risk identification method based on dynamic knowledge graphs disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0053] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0054] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0055] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the dynamic knowledge graph-based financial risk identification method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.

[0056] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned disclosed method for identifying financial risks based on dynamic knowledge graphs. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0057] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0058] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0059] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0060] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0061] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A financial risk identification method based on dynamic knowledge graphs, characterized in that, Applied to the target financial cloud platform, including: By using a pre-defined distributed data acquisition architecture, target financial data is collected in real time, and the target financial data is deduplicated to obtain an initial financial dataset. The initial financial dataset is standardized, and knowledge fusion is performed using a target entity disambiguation algorithm to obtain a standardized financial dataset. After determining that the preset update conditions are met, the preset financial knowledge graph is updated based on the standardized financial dataset and using an incremental learning algorithm to obtain the updated financial knowledge graph. The updated financial knowledge graph is used to mine the relationships between entities and determine the association risk weights, so as to construct a corresponding association risk network based on the association risk weights. The risk transmission calculation is performed on the associated risk network using a preset risk transmission calculation model. Based on the obtained risk transmission calculation results, the risk is identified through a preset risk identification model, and a corresponding risk identification report is generated and output.

2. The financial risk identification method based on dynamic knowledge graphs according to claim 1, characterized in that, The process involves acquiring target financial data in real time using a pre-defined distributed data acquisition architecture, and then filtering the target financial data for duplicates to obtain an initial financial dataset, including: First Financial Data was scraped using a multi-threaded crawler configured based on the Scrapy framework. Call the preset API interface and use the preset encrypted transmission protocol to obtain the second financial data, so as to obtain the target financial data including the first financial data and the second financial data; The financial data in the target financial data that meets the preset invalidity conditions are filtered according to preset filtering rules, and duplicate data are removed based on a hash algorithm to obtain the initial financial dataset.

3. The financial risk identification method based on dynamic knowledge graphs according to claim 1, characterized in that, The initial financial dataset includes structured financial data, semi-structured financial data, and unstructured financial data; Accordingly, the standardization of the initial financial dataset and the knowledge fusion through a target entity disambiguation algorithm to obtain a standardized financial dataset include: The fields in the structured financial data are mapped to standard fields in the target financial field using the field mapping method, and the numerical data in the structured financial data are normalized to obtain the first standardized data. The semi-structured financial data is parsed using XML path language to convert it into a target structured format, thus obtaining second standardized data. The unstructured financial data is segmented and part-of-speech tagging is performed based on a preset natural language processing model to obtain third standardized data. The target similarity between entities in the target standardized data is determined, and the target similarity is compared with a preset similarity threshold. Knowledge fusion is performed based on the comparison results to obtain a standardized financial dataset. The target similarity is determined based on the entity's name similarity, attribute similarity, and contextual similarity.

4. The financial risk identification method based on dynamic knowledge graphs according to claim 1, characterized in that, After determining that the preset update conditions are met, the preset financial knowledge graph is updated based on the standardized financial dataset and using an incremental learning algorithm to obtain the updated financial knowledge graph, including: Determine the current financial data increment based on the standardized financial dataset; If the incremental financial data exceeds the preset data increment threshold, or if the target key data is updated, or if the preset update time period is reached, then the newly added entity set and the newly added relationship set between the financial data increment are determined. An incremental learning algorithm is used to update the preset financial knowledge graph based on the newly added entity set and the newly added set of relationships between entities, so as to obtain the updated financial knowledge graph; the preset financial knowledge graph is constructed based on a preset standardized financial knowledge dataset, including an entity set and a set of relationships between entities.

5. The financial risk identification method based on dynamic knowledge graphs according to claim 1, characterized in that, The step of mining inter-entity relationships in the updated financial knowledge graph and determining relationship risk weights to construct a corresponding relationship risk network based on these weights includes: The risk association types between entities in the updated financial knowledge graph are determined, and the type weight coefficients of each risk association type and the entity risk score of each entity are determined by the analytic hierarchy process. Based on the type weight coefficients and the entity risk scores, the association risk weights between entities are determined; The corresponding associated risk network is constructed using the associated risk weights and the updated financial knowledge graph.

6. The financial risk identification method based on dynamic knowledge graphs according to any one of claims 1 to 5, characterized in that, The step of performing risk transmission calculation on the associated risk network using a preset risk transmission calculation model includes: The risk source node in the associated risk network is determined according to preset risk identification rules, and the risk transmission parameters of each node in the associated risk network are determined. The risk transmission parameters include risk attenuation coefficient, node risk carrying capacity, and risk transmission probability. The risk attenuation coefficient is determined based on the distance to the risk source node. The node risk carrying capacity is determined based on the size of the entity assets and credit status. The risk transmission probability is determined based on the associated risk weight. By using a preset risk transmission calculation model and based on the risk transmission parameters, the node risk value of each node in the associated risk network is iteratively calculated until a preset iteration termination condition is met, thereby obtaining the target node risk value of each node; the preset risk transmission calculation model is constructed based on a graph neural network; the preset iteration termination condition is that the change in the node risk value of each node is lower than a preset change threshold. Using the target shortest path algorithm and based on the risk value of the target node, the target risk transmission path in the associated risk network is identified, and the risk contribution of each node is determined to obtain the risk transmission calculation results.

7. The financial risk identification method based on dynamic knowledge graphs according to claim 6, characterized in that, The process involves identifying risks based on the obtained risk transmission calculation results and using a preset risk identification model to generate and output a corresponding risk identification report, including: Based on the target node risk value and the target risk transmission path, and by using a preset enterprise-related risk identification model, related risks are identified to obtain a first risk identification result; Using the target node risk value, the risk contribution rate, and the enterprise's historical default records, and through a preset credit default risk identification model, default risk is identified to obtain a second risk identification result; Generate and output a risk identification report that includes the first risk identification result, the second risk identification result, the target risk transmission path, and risk warning suggestions.

8. A financial risk identification device based on dynamic knowledge graph, characterized in that, Applied to the target financial cloud platform, including: The data filtering module is used to collect target financial data in real time through a preset distributed data acquisition architecture, and to perform deduplication filtering on the target financial data to obtain an initial financial dataset. The knowledge fusion module is used to standardize the initial financial dataset and perform knowledge fusion through a target entity disambiguation algorithm to obtain a standardized financial dataset. The graph update module is used to update the preset financial knowledge graph based on the standardized financial dataset and using an incremental learning algorithm after determining that the preset update conditions are met, so as to obtain the updated financial knowledge graph. The weight determination module is used to mine the relationships between entities in the updated financial knowledge graph and determine the association risk weights, so as to construct a corresponding association risk network based on the association risk weights. The risk transmission calculation module is used to perform risk transmission calculation on the associated risk network using a preset risk transmission calculation model, and to identify risks based on the obtained risk transmission calculation results and a preset risk identification model, thereby generating and outputting a corresponding risk identification report.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the financial risk identification method based on dynamic knowledge graph as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store computer programs; wherein, when the computer programs are executed by a processor, they implement the financial risk identification method based on dynamic knowledge graphs as described in any one of claims 1 to 7.