Enterprise archive intelligent risk control management method and system based on multi-source data
By constructing an enterprise knowledge graph and combining graph neural networks and time-series prediction models with a multi-source data fusion method, the problem of enterprise risk control systems being unable to integrate multi-source data has been solved. This enables comprehensive and forward-looking assessment and self-optimization of enterprise risks, improving the accuracy and adaptability of risk identification.
Patent Information
- Application Number
- CN202511706274.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-03-17
AI Technical Summary
Existing enterprise risk control systems cannot effectively integrate multi-source data, ignore the risks of enterprise-related networks and dynamic behaviors, resulting in incomplete risk assessments and a lack of self-updating mechanisms, making it difficult to adapt to new industry differences and new risks.
By constructing an enterprise knowledge graph, combining graph neural networks and time-series prediction models, risk quantification is achieved by integrating multi-source data, and the model is optimized through an adaptive learning mechanism driven by human feedback.
It enables multi-dimensional and dynamic quantitative assessment of enterprise risks, improves the comprehensiveness and foresight of risk identification, reduces false alarm rate, adapts to new risks in different industries, and has continuous optimization capabilities.
Smart Images

Figure CN121684601A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data management technology, specifically to an intelligent risk control management method and system for enterprise archives based on multi-source data. Background Technology
[0002] In the field of enterprise risk control management, existing technologies mainly focus on risk assessment based on a single data source. For example, financial institutions rely on credit reports, supply chain companies focus on performance records, and government regulators rely on business registration information. It is difficult to integrate cross-platform data such as judicial judgments, public opinion dynamics, and related companies, forming information silos. In terms of assessment dimensions, there is a focus on static indicators such as registered capital and debt-to-asset ratio, while neglecting the risk transmission of corporate networks and the trend risks of dynamic behaviors such as equity changes and litigation frequency, resulting in one-sided risk identification. From a technical perspective, most of them use fixed rules or single algorithms, which cannot adapt to the integrated analysis of structured financial data and unstructured public opinion texts, and lack a self-updating mechanism, making it difficult to cope with new industry differences and new risks. Existing enterprise risk control systems cannot simultaneously handle structural risks in enterprise networks and trend risks based on temporal behavioral characteristics, resulting in incomplete risk assessment. Therefore, there is an urgent need to develop an intelligent risk control management method for enterprise files based on multi-source data to solve the problems in existing technologies. Summary of the Invention
[0003] The purpose of this invention is to provide an intelligent risk control management method and system for enterprise files based on multi-source data, which can accurately quantify the network risks and behavioral trend risks of related enterprises, and has a simple structure and is easy to use, so as to solve the problems mentioned in the background art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: A method for intelligent risk control management of enterprise archives based on multi-source data includes the following steps: S1: Collect enterprise-related data from several heterogeneous data sources, process and integrate the data, and convert it into enterprise information in a unified format; S2: Based on the integrated enterprise information, construct an enterprise knowledge graph that includes enterprises, people and events, and reflect the relationships between the nodes; S3: Obtain the structural information of the subgraph centered on the target enterprise in the enterprise knowledge graph, and obtain the temporal behavioral feature information of the target enterprise. Input them together into a preset multi-factor risk fusion model for calculation, and obtain the comprehensive risk quantification result of the target enterprise. S4: Optimize and update the multi-factor risk fusion model based on human feedback information regarding the comprehensive risk quantification results.
[0005] By adopting the above technical solutions, a fundamental transformation of enterprise risk control has been achieved, from static, isolated, and post-event analysis to dynamic, integrated, and pre-event early warning. A unified enterprise information view has been constructed through multi-source data fusion, enterprise relationships have been deeply characterized using knowledge graphs, and risk quantification has been achieved by innovatively combining graph structure information with time-series behavioral characteristics. Finally, by introducing a closed loop of human feedback, the entire risk control system has the ability to continuously self-optimize, significantly improving the comprehensiveness, depth, and foresight of risk identification.
[0006] Preferably, S1 includes the following: S11: Collect structured and unstructured raw data from public data platforms and internal business systems, using any method including web crawlers and API interfaces; S12: Using natural language processing technology, name entity recognition and relation extraction are performed on the unstructured raw data to obtain standardized information on enterprises, people and events; S13: Based on a preset dynamic ontology model of enterprise information, the standardized information from different data sources is aligned and fused.
[0007] By adopting the above technical solutions, automated and standardized processing of multi-source heterogeneous data, especially unstructured data, has been achieved. Natural language processing technology is used to accurately extract key entities and relationships from text, and a dynamic ontology model is combined to solve the problem of fusion of semantic inconsistencies between data from different sources. This provides a reliable and unified data foundation for building high-quality knowledge graphs and subsequent accurate risk analysis, ensuring the accuracy of risk control decisions from the source.
[0008] As a preferred embodiment, the construction of the enterprise knowledge graph in S2 includes the following: storing the fused enterprise information in a graph database, wherein the event entities in the enterprise knowledge graph include any one or a combination of judicial events, public opinion events, equity change events, and financial events. Based on graph computing algorithms and preset rules, several dynamic risk labels are automatically assigned to enterprise entities in the graph.
[0009] By adopting the above technical solutions, various types of risk events are structurally incorporated into the knowledge graph, so that enterprise risks are no longer viewed in isolation. Furthermore, through graph computing and preset rules, dynamic risk labels are automatically assigned to enterprises, which can intelligently identify complex and implicit risk patterns.
[0010] Preferably, the multi-factor risk fusion model in S3 is a hybrid model based on neural networks, which includes: The graph neural network module is used to process the subgraph structure information and output a first feature vector representing the risk of the enterprise's associated network. The time-series prediction module is used to process the time-series behavioral feature information and output a second feature vector that represents the enterprise's own behavioral risk trend. The feature fusion module is used to fuse the first feature vector, the second feature vector, and the static feature vector of the target enterprise to generate the comprehensive risk quantification result.
[0011] By adopting the above technical solutions, a multi-dimensional and multi-perspective integrated quantitative assessment of enterprise risks is achieved. The graph neural network module mines structural and transmissive risks from the enterprise's interconnected network; the time series prediction module captures the risk evolution trend of the enterprise's own behavior; and finally, the feature fusion module cross-calculates network risks, behavioral trends, and static fundamental information, so that the final comprehensive risk quantification result has network insight, trend judgment, and fundamental stability, thereby improving accuracy.
[0012] Preferably, the time-series behavioral characteristic information includes at least one of the following within a preset time window: the number of litigation events involving the target enterprise, the number of negative public opinions, and the sequence of changes in financial indicators.
[0013] By adopting the above technical solution, the dynamic risk signals of enterprises can be quantitatively captured. By setting a preset time window and extracting key time-series indicators such as litigation, public opinion, and financial changes, abstract corporate behavior is transformed into specific and calculable feature sequences. This enables the model to keenly perceive the changing trends of corporate conditions, thereby providing key data input for forward-looking risk warning.
[0014] Preferably, step S4 optimizes and updates the multi-factor risk fusion model based on human feedback information regarding the comprehensive risk quantification results, specifically including the following: S41: Construct labeled training samples based on the early warning data generated by the system and identified as risk and false alarm labels; S42: When the number of labeled training samples reaches a preset threshold, a model update task is triggered; S43: Based on the currently running multi-factor risk fusion model, fine-tune it using the labeled training samples and update its model weight parameters; S44: Deploy the updated model to replace the old model.
[0015] By adopting the above technical solution, we collect expert confirmations and feedback on misjudgments of early warning results and transform them into high-quality labeled training samples. After the data accumulates to a certain scale, we trigger incremental learning of the model, enabling the multi-factor risk fusion model to continuously learn from the latest decisions of experts and maintain iterative progress.
[0016] A smart risk control management system for enterprise archives based on multi-source data, which applies the above-mentioned smart risk control management method for enterprise archives based on multi-source data, includes: a data acquisition and fusion module for automatically collecting enterprise-related data from multiple heterogeneous data sources and processing and integrating it; A knowledge graph construction and management module for building and maintaining enterprise knowledge graphs from the merged data; An intelligent risk quantification engine for running the multi-factor risk fusion model and outputting comprehensive risk quantification results; An adaptive learning engine for optimizing the multi-factor risk fusion model based on feedback information; This is a risk control business application platform used to display enterprise profiles, knowledge graphs, and risk warning information, and to receive the aforementioned manual feedback information.
[0017] By adopting the above technical solution, each module has a clear division of labor and works collaboratively. The data acquisition and fusion module ensures data supply, the knowledge graph module is responsible for knowledge expression and storage, the intelligent risk quantification engine is responsible for core calculations, the adaptive learning engine maintains system iteration and optimization, and the risk control business application platform provides a human-computer interaction interface. This not only ensures the efficient operation of the system, but also facilitates the independent upgrade and maintenance of each part in the future.
[0018] This application also discloses an intelligent risk control management device for enterprise files based on multi-source data, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned intelligent risk control management method for enterprise files based on multi-source data.
[0019] This application also discloses a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the above-described intelligent risk control management method for enterprise archives based on multi-source data.
[0020] Compared with the prior art, the beneficial effects of the present invention are: This application utilizes a hybrid algorithm combining graph neural networks and time-series prediction models, which breaks through the limitations of traditional methods. Compared with traditional single-algorithm models, it can improve the comprehensive identification accuracy of enterprise risks, predict risk trends in advance, and accurately quantify the network risks and behavioral trend risks of related enterprises, thus achieving a leap from static assessment to dynamic prediction. This application utilizes a feedback-driven adaptive learning mechanism to enable the system to continuously evolve and iterate, significantly improving the comprehensiveness, foresight, and automation of risk control, reducing the model's false alarm rate, and significantly enhancing its adaptability to different industries and new types of risks. Long-term use can continuously optimize the stability and accuracy of risk identification, reducing the cost of manual risk control.
[0021] Other features and advantages of the present invention will be disclosed in detail in the following detailed description and accompanying drawings. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating the steps of an intelligent risk control management method for enterprise archives based on multi-source data, as described in an embodiment of the present invention. Figure 2 This is a schematic diagram of an intelligent risk control management device for enterprise archives based on multi-source data, as described in an embodiment of the present invention. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] In this embodiment of the invention, a method for intelligent risk control management of enterprise archives based on multi-source data is described below. Figure 1 As shown, it includes the following: S1: Data Acquisition and Fusion S11: Multi-source data acquisition S111: Public Data Collection: A targeted crawler was developed using the Scrapy crawler framework, configured with an IP proxy pool, to crawl enterprise registration information from the National Enterprise Credit Information Publicity System, enterprise litigation judgments from the China Judgments Online website, and industry public opinion news from Sina Finance; supplementary data was obtained by connecting to third-party data platforms through API interfaces, with the collection frequency set to daily incremental collection and weekly full update. S112: Internal Data Acquisition: Through an interface program developed in Java, it connects to the enterprise's internal CRM system, ERP system, and contract management system to extract the client's cooperation records, financial transaction data, contract performance status, etc. The acquisition frequency is set to real-time acquisition and daily batch synchronization.
[0025] S12: Data Standardization Processing S121: Unstructured Data Parsing: Using the NLTK Natural Language Processing library, NER (Named Entity Recognition) and relation extraction are performed on unstructured data such as judgments and public opinion texts. For example, entities such as company names, personal names, event types, and time nodes are extracted from the text using a BERT pre-trained model; relations are extracted through dependency parsing. S122: Data cleaning: Based on the unified social credit code and event number, remove duplicate and invalid data, including incorrectly formatted dates and null fields, and correct erroneous data, such as typos and inconsistent numerical units.
[0026] S13: Data Alignment and Fusion S131: Load the preset dynamic ontology model of enterprise information, store it in the MySQL database, support manual maintenance and updates, and align the standardized data according to the semantic rules in the model; S132: Uses the Spark data processing engine to perform fusion operations, linking and integrating multi-source data from the same enterprise to generate a unified format enterprise information data table, including dimensions such as basic enterprise information, related personnel, historical events, and financial indicators.
[0027] S2: Enterprise Knowledge Graph Construction S21: Importing Spectral Data The merged enterprise information data table is converted into a node and edge data format supported by the graph database, and then imported into the Neo4j graph database using Cypher statements.
[0028] In one feasible embodiment, the following statements are included: CREATE (e:Enterprise {id: '91110105MA01Y23X7C', name: 'XX Technology Co., Ltd.', registerCapital: 500, establishDate: '2020-01-15'}) / / Create an enterprise node / / CREATE (p:Person {id: 'P001', name: 'Zhang San', position: 'Legal Representative'}) / / Create a person node / / MATCH (e:Enterprise), (p:Person) WHERE e.id = '91110105MA01Y23X7C' AND p.id = 'P001' CREATE (p)-[:IS_LEGAL_REPRESENTATIVE_OF]->(e) / / Create associated edges / / S22: Event Entity Entry For legal events, public opinion events, equity change events, and financial events, create event nodes and associate them with corresponding enterprise nodes.
[0029] In one feasible embodiment, the following statements are included: CREATE (event:JudicialEvent {id: 'E001', type: 'Sales Contract Dispute',startTime: '2023-03-20', status: 'Under Trial'}) MATCH (e:Enterprise {id: '91110105MA01Y23X7C'}) CREATE (e)-[:INVOLVES]->(event) S23: Dynamic Risk Label Generation The system deploys graph computation algorithms and a preset rule engine. These algorithms include PageRank and community detection algorithms, automatically tagging enterprise nodes. For example: Rule 1: If there are ≥3 lawsuits within 6 months, the case will be labeled "high legal risk"; Rule 2: If negative public opinion is exposed ≥ 5 times / month, it will be labeled "high public opinion risk"; Rule 3: If the number of equity changes is ≥ 2 times per quarter, then label it "Equity Unstable"; Tags are stored in the enterprise node attributes and can be updated in real time.
[0030] S3: Comprehensive Risk Quantification Calculation S31: Feature Data Extraction S311: Subgraph structure information extraction: Centered on the target enterprise, extract subgraph data with 1-3 degree association through Neo4j's graph query function, and convert it into an adjacency matrix and node feature matrix that can be processed by graph neural networks.
[0031] S312: Time-series behavioral feature extraction: Set a time window of 6 months to extract the time-series data sequence of the target company, including: Judicial dimension: Monthly number of litigation cases and cumulative amount of litigation; Public opinion dimensions: daily exposure frequency of negative public opinion, scope of public opinion dissemination; Financial dimension: The sequence of monthly changes in operating revenue, debt-to-equity ratio, and net cash flow.
[0032] S313: Static Feature Extraction: Extract fixed attribute information of the target enterprise, such as registered capital, years of establishment, industry category, qualification certification level, etc., and convert it into a standardized feature vector.
[0033] S32: Multifactor Risk Fusion Model Calculation S321: In one feasible embodiment, the model structure is as follows: Graph Neural Network Module: Adopting the GCN architecture, it takes the subgraph adjacency matrix and node feature matrix as input, and outputs a 128-dimensional first feature vector through two layers of convolution operations.
[0034] Temporal prediction module: It adopts an LSTM architecture, takes a temporal behavior feature sequence as input, sets the hidden layer dimension to 64, and outputs a 64-dimensional second feature vector.
[0035] Feature fusion module: The attention mechanism is used to perform weighted fusion of the first feature vector, the second feature vector, and the static feature vector, and the comprehensive risk quantification result of 0-100 points is output through a fully connected layer.
[0036] S33: Risk Outcome Output The comprehensive risk quantification score and the risk contribution ratio of each dimension are stored in the results database and synchronized to the risk control business application platform.
[0037] S4: Model Optimization and Update S41: Construction of labeled training samples S411: Manual Feedback Collection: Through the risk control business application platform, risk control experts annotate the risk quantification results, including: S4111: Confirmed Risk: If a risk event actually occurs and the model prediction score is ≥61, then it is marked as a true positive. S4112: False positive labeling: If the model predicts a score ≥ 61, but no risk event actually occurs, then label it as a false positive; S4113: Missed Reporting: If a risk event actually occurs, but the model prediction score is ≤30, then it is marked as a false negative. S412: Sample storage: Store the labeled samples in the model training sample library. The sample fields include subgraph features, temporal features, static features, and label.
[0038] S42: In one feasible embodiment, model update triggering and training include: S421: Triggering condition: Set the sample number threshold to 1000. When the cumulative number of newly added labeled samples reaches 1000, the model update task will be automatically triggered. S422: Fine-tuning training: Using the currently running model as the initial model, the Adam optimizer is used with a learning rate of 0.001, a batch size of 32, and 50 training iterations. Labeled samples are used to fine-tune the model weight parameters to minimize the cross-entropy loss function.
[0039] S43: Model Deployment and Replacement After training is completed, the model evaluation metrics are used to verify the effect. Once the target is met, the model is automatically deployed to the intelligent risk quantification engine to replace the old model. The deployment process supports canary releases.
[0040] The embodiment of the intelligent risk control management device for enterprise archives based on multi-source data of the present invention can be applied to any device with data processing capabilities, such as a computer. The device embodiment can be implemented through software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution. From a hardware perspective, such as... Figure 2 The diagram shown is a hardware structure diagram of any device with data processing capabilities, including the enterprise file intelligent risk control management device based on multi-source data according to the present invention. (Except for...) Figure 2 In addition to the processor, memory, network interface, and non-volatile memory shown, any data processing device in the embodiment may also include other hardware depending on the actual function of that data processing device, which will not be elaborated further. The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0041] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0042] This invention also provides a computer-readable storage medium storing a program thereon, which, when executed by a processor, implements the intelligent risk control management device for enterprise archives based on multi-source data described in the above embodiments.
[0043] The computer-readable storage medium can be an internal storage unit of any data processing device described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device of any data processing device, such as a plug-in hard disk, SmartMediaCard (SMC), SD card, or FlashCard equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of any data processing device. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.
[0044] Example 1 Using a commercial bank's credit risk control business as an example, the goal is to conduct a comprehensive risk assessment of companies applying for loans to avoid credit default risk. The target company is XX Intelligent Manufacturing Co., Ltd., which applied for a 5 million yuan working capital loan (hereinafter referred to as the target company).
[0045] Step 1: Data Acquisition and Fusion Public data collection: The target company's business registration information was obtained through web scraping: Unified Social Credit Code 9****************D, registered capital of 10 million yuan, established in 2018; 3 litigation records from March to August 2023 shown on the China Judgments Online website, all of which were sales contract disputes; 2 negative public opinion articles from June to August 2023 in industry news.
[0046] Internal data collection: The target company's past credit records were obtained through the bank's internal CRM system, and its financial data from March to August 2023 were obtained through the ERP system. The operating income was 8 million, 8.5 million, 7.8 million, 7.2 million, 6.8 million and 6.5 million respectively, showing a downward trend; the asset-liability ratio increased from 45% to 58%.
[0047] Data fusion: Through the dynamic ontology model of enterprise information, litigation records and disputes are uniformly classified as judicial events, product quality complaints are classified as public opinion events, and financial data is converted into monthly indicators in a unified format to generate a unified information table for the target enterprise.
[0048] Step 2: Knowledge Graph Construction Import graph database: Create target company node, legal representative Li Si node, 3 litigation event nodes, and 2 public opinion event nodes, and construct related edges. Li Si is the legal representative of the target company, and the target company is involved in 3 sales contract disputes and 2 quality complaint public opinion incidents.
[0049] Dynamic tag generation: Tag high legal risk if involved in 3 lawsuits within 6 months; tag medium public opinion risk if there are 2 negative public opinion events within 3 months; no unstable equity tag if there are no frequent equity changes.
[0050] Step 3: Comprehensive Risk Quantification Calculation Feature extraction: Subgraph structure: Extract two directly related companies of the target company, one of which has a litigation record. The subgraph is then converted into a 3×3 adjacency matrix.
[0051] Time series characteristics: 6-month litigation number sequence [0,1,1,0,0,1], negative public opinion sequence [0,0,1,0,1,0], operating revenue sequence [800,850,780,720,680,650], asset-liability ratio sequence [45%,48%,50%,53%,56%,58%].
[0052] Static characteristics: Registered capital of 10 million yuan, standardized value of 0.6; established for 5 years, standardized value of 0.5; manufacturing industry, standardized value of 0.7.
[0053] Model operation: The graph neural network module outputs the risk vector of the associated network, which represents the possibility of risk transmission among related enterprises. The LSTM module outputs the trend risk vector, which represents the risk of declining operating income and rising debt ratio. After feature fusion, the comprehensive risk score of 72 points is output, which is high risk.
[0054] Step 4: Human Feedback and Model Optimization Manual labeling: After verification by bank risk control experts, it was confirmed that the target company had a significant risk of credit default due to a continuous decline in operating income and unresolved litigation disputes. The labeled sample was a true positive.
[0055] Model update: When the cumulative number of such labeled samples reaches 1,000, the model is fine-tuned. After the update, the model adjusts the time-series trend risk weight of manufacturing enterprises to 0.45, the association network risk weight to 0.35, and the static risk weight to remain at 0.2.
[0056] In summary, the target company's overall risk score was 72, indicating high risk. Based on this result, the bank rejected its loan application. Subsequent follow-up showed that the company fell into a debt crisis three months later due to a broken cash flow, validating the effectiveness of the method. After model optimization, the accuracy of risk identification for manufacturing companies increased from 85% to 88%, and the false positive rate decreased from 12% to 9%.
[0057] This invention provides an intelligent risk control management method for enterprise files based on multi-source data, which can accurately quantify the network risks and behavioral trend risks of related enterprises and has high reliability.
[0058] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
[0059] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A multi-source data-based intelligent risk control management method for enterprise archives, characterized in that, Comprising the following steps: S1: Collecting data related to enterprises from a plurality of heterogeneous data sources, and processing and fusing the data to convert into unified format enterprise information; S2: Based on the fused enterprise information, constructing an enterprise knowledge graph containing enterprises, persons and events, and reflecting the association between nodes; S3: Obtaining the structural information of a subgraph centered on a target enterprise in the enterprise knowledge graph, and obtaining the time series behavior feature information of the target enterprise, and inputting them into a preset multi-factor risk fusion model for calculation to obtain the comprehensive risk quantification result of the target enterprise; S4: According to the artificial feedback information for the comprehensive risk quantification result, optimizing and updating the multi-factor risk fusion model.
2. The enterprise archives intelligent risk control management method based on multi-source data according to claim 1, characterized in that, The S1 includes the following contents: S11: Collecting structured and unstructured raw data from public data platforms and internal business systems, using any of the network crawler, API interface; S12: Through natural language processing technology, performing named entity recognition and relation extraction on the unstructured raw data to obtain standardized enterprise, person and event information; S13: Based on a preset enterprise information dynamic ontology model, aligning and fusing the standardized information from different data sources. 3.The enterprise archives intelligent risk control management method based on multi-source data according to claim 1, characterized in that, The enterprise knowledge graph construction in S2 includes the following contents: storing the fused enterprise information in a graph database, the event entities in the enterprise knowledge graph include any one and combination of judicial events, public opinion events, equity change events and financial events; Based on graph computing algorithm and preset rules, automatically labeling enterprise entities in the graph with a plurality of dynamic risk labels.
4. The enterprise archives intelligent risk control management method based on multi-source data according to claim 1, characterized in that, The multi-factor risk fusion model in S3 is a hybrid model based on neural network, which includes: A graph neural network module for processing the subgraph structure information and outputting a first feature vector representing the risk of enterprise association network; A time series prediction module for processing the time series behavior feature information and outputting a second feature vector representing the risk trend of enterprise behavior; A feature fusion module for fusing and calculating the first feature vector, the second feature vector and the static feature vector of the target enterprise to generate the comprehensive risk quantification result.
5. The method of claim 1, wherein the method further comprises: The time series behavior feature information includes at least one of the number of litigation events, the number of negative public opinions, and the sequence of financial index changes of the target enterprise within a preset time window.
6. The method of claim 1, wherein the method further comprises: The S4 optimizes and updates the multi-factor risk fusion model according to the artificial feedback information for the comprehensive risk quantification result, specifically including the following contents: S41: Based on the warning data generated by the system and confirmed as risk and false alarm labels, constructing labeled training samples; S42: When the number of labeled training samples reaches a preset threshold, trigger model updating task; S43: Based on the currently running multi-factor risk fusion model, fine-tune it using the labeled training samples to update its model weight parameters; S44: Deploy the updated model to replace the old model.
7. An enterprise archives intelligent risk control management system based on multi-source data, applying an enterprise archives intelligent risk control management method based on multi-source data according to claims 1-6, characterized in that, Comprising: A data collection and fusion module for automatically collecting and processing and fusing data related to an enterprise from multiple heterogeneous data sources; A knowledge graph construction and management module for constructing and maintaining an enterprise knowledge graph from the fused data; An intelligent risk quantification engine for running the multi-factor risk fusion model and outputting a comprehensive risk quantification result; An adaptive learning engine for optimizing the multi-factor risk fusion model according to feedback information; A risk control business application platform for displaying enterprise archives, knowledge graphs, and risk warning information and receiving the artificial feedback information.
8. An enterprise archives intelligent risk control management device based on multi-source data, characterized in that: A computer device comprising a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the method of claim 1-7.
9. A computer-readable storage medium, characterized in that: A computer program product, wherein a program is stored thereon, and the program is executed by a processor to implement the method of claim 1-7.
Citation Information
Cited By
AI-driven intelligent optimization decision-making method and system for bank asset allocation
CN122115117A