A multi-source data fusion enterprise credit report intelligent generation method and product

By using the MCP protocol and the entity alignment and time decay factor weighting mechanism of graph neural networks, the problem of accessing and integrating multi-source data in the generation of corporate credit reports is solved, thereby improving data consistency and accuracy. The generated reports meet financial compliance requirements and are suitable for corporate background checks, credit approval, and supply chain risk assessment.

CN122113870AActive Publication Date: 2026-05-29知呱呱(天津)大数据技术有限公司 +3
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
知呱呱(天津)大数据技术有限公司
Filing Date
2026-04-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies for generating corporate credit reports suffer from bottlenecks in the coupling and expansion of multi-source data access systems, a lack of deep semantic-level integration of multi-source heterogeneous data, and a lack of natural language interaction and uncontrollable generation. These issues lead to logical contradictions and factual errors in the reports, making it difficult to meet the compliance and security requirements of financial scenarios.

Method used

A multi-source data fusion approach is adopted, which uses the MCP protocol to standardize the encapsulation and dynamic scheduling of data sources, uses graph neural networks for entity alignment and a credibility weighting mechanism for time decay factors to resolve conflicts, and combines graph and deep learning models for feature calculation and report generation to ensure data consistency and traceability.

Benefits of technology

It achieves a unified model for multi-source data, improving data consistency and accuracy. The generated credit reports, while ensuring the flexibility of natural language interaction, meet the accuracy and auditability requirements of commercial compliance, and are suitable for scenarios such as corporate background checks, credit approval, and supply chain risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113870A_ABST
    Figure CN122113870A_ABST
Patent Text Reader

Abstract

The application discloses a multi-source data fusion enterprise credit report intelligent generation method and product. The method comprises the following steps: constructing and registering a heterogeneous data tool library based on an MCP protocol; an LLM analyzes a natural language instruction input by a user; a multi-source API is dynamically routed and scheduled through an MCP server; unified structured data is obtained by performing entity alignment on the obtained multi-source heterogeneous data, and then performing conflict resolution based on a time decay factor credibility weighting mechanism; feature calculation and association mining are performed based on fusion true values, and then a dynamic Prompt is constructed by combining analysis features and user intentions; a large language model generates a credit report and adds a data source anchor point. The application fundamentally alleviates the calculation deviation caused by the cross-modal semantic gap, effectively suppresses the "data illusion", and generates a report that guarantees the flexibility of natural language interaction while achieving commercial compliance levels in terms of accuracy and auditability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of artificial intelligence and financial data processing, specifically involving a method for processing and generating reports of multi-source heterogeneous enterprise credit data, applicable to scenarios such as enterprise background checks, credit approval, and supply chain risk assessment. Background Technology

[0002] Currently, in the field of corporate credit investigation and due diligence, the generation of credit reports mostly relies on traditional microservice architectures or manual data compilation. Existing technologies generally require business systems to first write hard-coded scripts to call the corresponding independent API interfaces of various data sources to obtain structured data. After acquisition, a static rule engine is used to perform simple data concatenation, and finally, business personnel manually analyze the data or apply fixed templates to generate reports.

[0003] In recent years, some companies have also begun to try using Large Language Models (LLM) for assisted generation. The usual practice is to directly concatenate the raw JSON data returned by various APIs and input it into the large model, allowing it to perform end-to-end summarization and text generation.

[0004] Existing microservice API calls and basic large-scale model generation solutions perform poorly in credit score retrieval and report generation, mainly for the following reasons:

[0005] I. Severe system coupling and expansion bottlenecks exist in multi-source data access: Commonly used single or point-to-point API integration methods lack standardized scheduling. Each new data source (such as enforcement information from a local court) requires re-interface integration and code deployment, lacking dynamic routing and orchestration capabilities.

[0006] Second, the lack of deep semantic-level integration of multi-source heterogeneous data: data sources such as business registration, finance, and law are fragmented and have different granularities. Existing systems often only perform physical splicing, and when different data sources conflict on the same indicator (such as the update time of a company's paid-in capital), they cannot effectively resolve the conflict, resulting in logical contradictions in the context provided to the large model.

[0007] Third, the lack of natural language interaction and uncontrollable generation (illusion problem): Traditional query rules are hard-coded in the front-end form, lacking flexibility. Directly feeding raw data to a large model to generate reports, due to the lack of a unified data specification layer (intermediate representation) constraint, the large model is prone to factual errors (data illusion), and it is difficult to trace the source of data, which cannot meet the stringent compliance and security requirements of financial scenarios.

[0008] Patent document CN120598693A discloses a risk assessment method based on a Model Context Protocol (MCP). This method integrates a large language model with external data sources and tools to establish a secure, bidirectional connection between the large model and the data sources. Through the MCP's roots mechanism and pre-built toolchain, it unifies the interfaces of multi-source heterogeneous data, reducing redundant development work and shortening the system integration cycle. By supporting cross-system data acquisition, it achieves dynamic data aggregation, eliminates data silos, reduces underwriting error rates, and improves the underwriting accuracy of financial insurance systems. While this solution overcomes the first problem to some extent, its subsequent processing methods for multi-source data are not suitable for the intelligent generation of corporate credit reports, and the latter two problems remain: conflicting observations of the same indicator from different data sources leading to logical contradictions in the context provided to the large model, and a lack of natural language interaction and uncontrollable generation. Summary of the Invention

[0009] To address the problems existing in the prior art, this application provides a method and product for intelligent generation of enterprise credit reports through multi-source data fusion.

[0010] The technical solution provided in this application is as follows:

[0011] A method for intelligently generating enterprise credit reports through multi-source data fusion includes:

[0012] S1. Receive natural language instructions input by the user, the natural language instructions containing the name of the target company and evaluation dimensions; perform intent parsing on the natural language instructions, and obtain multi-source heterogeneous data about the target company through a pre-built standardized toolset and dynamic discovery mechanism based on the MCP protocol.

[0013] S2. For the multi-source heterogeneous data, data cleaning and standardization are performed, and then a graph neural network model (GNN) is used to achieve entity alignment, mapping nodes from different data sources but pointing to the same enterprise entity to globally unique entity identifiers; after entity alignment, for conflicting observations of the same attribute from different data sources, a credibility weighting mechanism based on time decay factor is introduced to calculate the unique definite value of the attribute, and finally a unified data model for the target enterprise is obtained.

[0014] S3. Based on the evaluation dimensions analyzed in step S1, determine the applicable indicators based on the pre-set evaluation indicator system, extract a series of feature data corresponding to the applicable indicators from the unified data model obtained in step S2, and comprehensively calculate the score value of the evaluation dimension; then, based on the globally unique entity identifier, construct a business association graph as structured support data for use in the subsequent report generation stage.

[0015] S4. Based on the evaluation dimensions, recall the corresponding structured framework from the pre-built Prompt template library, and inject the score values ​​of the evaluation dimensions obtained in step S3 and the feature data into the Context area of ​​the reporting agent after format conversion; generate an enterprise credit report according to the constraint instructions set in the system prompt words; the constraint instructions require the reporting agent to reason only based on the data injected in the Context area, and each conclusion must be marked with the data source anchor point.

[0016] Optionally, in step S2, the confidence weighting mechanism based on the time decay factor is introduced, assuming that there are n data sources providing observations for the target enterprise attributes, and the observation value of the i-th data source is... The system presets static confidence weights for each data source. And introduce a time decay factor ,in This represents the time difference between the update time of the data record in this data source and the present time. The empirical decay coefficient is the uniquely determined value of the attribute. The calculation formula is as follows:

[0017] .

[0018] Optionally, in step S2, the graph neural network model constructs an enterprise entity graph with enterprise entities as nodes and common attributes between different data sources as edges; the graph neural network model aggregates neighbor node information through multi-layer graph convolution operations, learns the vector representation of nodes, and maps nodes from different data sources but pointing to the same real entity to globally unique entity identifiers based on vector similarity.

[0019] Optionally, in step S2, the data cleaning and standardization includes:

[0020] Complete the missing data;

[0021] Identify and process abnormal noise points in the data;

[0022] Scaling numerical features to a uniform dimension;

[0023] By referring to the pre-set evaluation index system, the raw feature data is transformed into structured feature data.

[0024] Optionally, in step S3, the comprehensive calculation specifically includes:

[0025] Suppose that the evaluation dimension mentioned in step S1 corresponds to a primary indicator in the pre-set evaluation indicator system. The applicable indicators of this evaluation dimension include multiple secondary indicators in the pre-set evaluation indicator system, and each secondary indicator is further subdivided into multiple tertiary indicators.

[0026] For each secondary indicator, the feature data of its multiple sub-tertiary indicators are converted into normalized feature values, and then weighted and fused to obtain the normalized feature fusion result.

[0027] The normalized feature fusion results of all secondary indicators under the primary indicator are combined to form a feature vector, which is then input into the trained deep learning model. During training, the deep learning model divides the score of the primary indicator into multiple scoring intervals as the training target. By outputting the probability corresponding to each scoring interval, the model calculates the final comprehensive score by weighting the score and the corresponding probability value of each scoring interval.

[0028] Optionally, in step S3, the construction of the business association graph is based on the globally unique entity identifier to identify the enterprise's industrial chain, industrial development prospects, enterprise equity penetration path, actual controller link, and associated risk transmission path.

[0029] Optionally, in step S4, the interceptor mechanism defined by the MCP protocol is used to monitor the model output stream in real time, identify text fragments involving specific values, and associate the specific value with the data injected into the Context area through matching or semantic recognition, and automatically insert data source anchors.

[0030] Optionally, in step S4, the function of generating enterprise credit reports is integrated into the user's instant messaging client as a program module. The structured report data is rendered into a visual page and displayed on the user's instant messaging client's interactive interface. In response to the user's document export command triggered in the interactive interface, the visual page is converted into a document in a preset format and sent to the instant messaging client's session window, or a download link for the document is provided.

[0031] This application also provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the above-described method steps for intelligent generation of enterprise credit reports based on multi-source data fusion.

[0032] This application also provides a computer program product, including a computer program / instruction that, when executed by a processor, implements the steps of the above-described intelligent generation method for enterprise credit reports based on multi-source data fusion.

[0033] Compared with the prior art, this application has at least the following beneficial effects:

[0034] The intelligent generation method for enterprise credit reports proposed in this application first uses a graph neural network model to align entities for multi-source heterogeneous data related to user intent. Then, for conflicting observations of the same attribute that may exist from different data sources, a credibility weighting mechanism based on time decay factor is introduced to resolve the conflict. This makes the multi-source data form a unique and definite value (Entity-ID) before being input into the large model, resulting in a unified data model (UDM) for the target enterprise (user intent). Compared with simple data splicing, the consistency of the underlying data is greatly improved, which fundamentally alleviates the computational bias caused by the cross-modal semantic gap and solves the pain point that multi-source fragmented data cannot be directly used for large model inference.

[0035] Based on the construction and acquisition of a unified data model for the target enterprise, this application extracts a series of feature data corresponding to applicable indicators, calculates a comprehensive score, and constructs a business relationship graph based on aligned entities (globally unique entity identifiers) to enrich the structured support data for use in the subsequent report generation stage. Then, the score and feature data are converted and injected into the Context area of ​​the report agent (large model). Explicit constraints are set in the system prompt, requiring the large model to reason only based on the injected Context data, and each conclusion must be labeled with the data source anchor. This implicitly ensures that the large model follows rigorous financial analysis logic, effectively suppressing "data illusion." This allows the automatically generated credit report to maintain the flexibility of natural language interaction while achieving commercial compliance levels in terms of accuracy and auditability, meeting regulatory requirements such as the "Credit Reporting Business Management Measures." Attached Figure Description

[0036] Figure 1 A schematic diagram of the principle (overall technical flow) of an intelligent generation method for enterprise credit reports based on multi-source data fusion provided in one embodiment of this application.

[0037] Figure 2 This is a timing diagram of parallel multi-source data invocation in one embodiment of this application;

[0038] Figure 3 This is a schematic diagram of a multi-source data fusion and conflict resolution process in one embodiment of this application. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments.

[0040] Based on further analysis and research into existing technical problems, this application recognizes that conflicting observations of the same indicator from different data sources can lead to logical contradictions in the context provided to large models. Furthermore, existing technologies directly feed raw collected data into large models to generate reports. Due to the lack of a unified data standardization layer (intermediate representation) constraint, large models are highly susceptible to factual errors (data illusions) and have difficulty tracing data sources, failing to meet the stringent compliance and security requirements of financial scenarios. Therefore, this application not only introduces the MCP protocol into enterprise credit reporting systems but also standardizes the underlying heterogeneous data sources by introducing standardized tool encapsulation based on the MCP protocol at the data access layer. Packaged as Tools, it enables dynamic scheduling and decoupling of the Large Language Model (LLM) and external data driven by natural language. It also constructs a multi-source data conflict resolution algorithm that integrates static weights and time decay factors. By outputting a unified attribute truth value (uniquely determined value) through a mathematical expectation formula, it solves the pain point that fragmented multi-source data cannot be directly used for large model inference. Finally, it proposes a restricted generation mechanism for large models of report agents with data traceability. By pre-injecting the calculated user intent evaluation dimensions (such as enterprise scientific and technological innovation capability scores) and cleaned structured data into the Prompt, it ensures the logical rigor and auditability of the generated report.

[0041] The main process of the entire plan is as follows: Figure 1 As shown, it includes:

[0042] S1. Pre-build and register a heterogeneous data tool library based on the MCP protocol; LLM parses the natural language commands input by the user; based on the parsed user intent, dynamically route and schedule multi-source APIs through the MCP server to obtain multi-source heterogeneous data;

[0043] S2. Perform entity alignment and format standardization on multi-source heterogeneous data, and then resolve conflicts based on the credibility weighting mechanism of time decay factor to obtain unified structured data.

[0044] S3. Based on the fusion of truth values, perform feature calculation and association mining (business association graph), and then combine the analysis features and user intent to construct a dynamic Prompt (the score value of the user intent evaluation dimension and the feature data after format conversion are injected into the Context area of ​​the large model).

[0045] S4. Through the constraint instructions set in the system prompt (requiring the large model to perform inference only based on the data injected from the Context region, and each conclusion must be marked with the data source anchor), the large model generates a credit report with data source anchors (source tracing anchors) under restrictions.

[0046] Specifically, in one embodiment, a method for intelligently generating enterprise credit reports through multi-source data fusion includes:

[0047] S1. Receive natural language instructions input by the user, which contain the name of the target company and evaluation dimensions; perform intent parsing on the natural language instructions, and obtain multi-source heterogeneous data about the target company through a pre-built standardized toolset and dynamic discovery mechanism based on the MCP protocol.

[0048] Here, users should enter the full name of the target company to avoid the system being unable to determine the user's true intentions. For example:

[0049] If you only enter "Generate Huawei's Science and Technology Innovation Capability Assessment Report", the system will find companies such as "Huawei Technologies Co., Ltd.", "Huawei Terminal Co., Ltd.", "Huawei HiSilicon Semiconductor Co., Ltd.", and "Huawei Cloud Computing Technology Co., Ltd." The system cannot determine whether the user's true intention is to search for "Huawei Technologies Co., Ltd." or other related companies such as "Huawei Terminal Co., Ltd.", "Huawei HiSilicon Semiconductor Co., Ltd.", and "Huawei Cloud Computing Technology Co., Ltd."

[0050] Of course, in this situation, a multi-round question-and-answer design can be used. For example, the system can return "Which of the following companies' science and technology innovation capability assessment reports do you want to query: ...", and the user can then enter the full name of the target company.

[0051] The core components of this step are the MCP Client (integrated into the large language model) and the MCP Server (data service bus). Distributed heterogeneous databases and APIs from sectors such as industry and commerce, law, and intellectual property are encapsulated into tools that conform to the MCP standard. Figure 2 As shown, when a user inputs a natural language command (such as "analyze the technological innovation risks of Company A"), the LLM (e.g., large language models such as Qianwen and Zhipu can be selected) acts as the MCP Client to parse the intent and send a standardized tool invocation request to the Server through the MCP protocol. This mechanism realizes dynamic discovery and invocation of data sources through standardized tool descriptions, eliminating the need to write hard-coded routing logic for each new data source, thus achieving flexible access to data sources at the application layer.

[0052] The system automatically collects diverse and heterogeneous data from various internal and external sources from enterprises. Since the user's intended evaluation dimension is scientific and technological innovation capability, data sources can include: enterprise business registration information, annual reports, patent data (applications, authorizations, transfers, etc.), R&D personnel composition, R&D expenditure, science and technology awards, investment and financing events (with a particular focus on financing in the science and technology innovation field), science and technology project initiation and acceptance information, etc. Currently, the raw data from these multiple sources is primarily in JSON and XML formats, but may also include data in other formats.

[0053] S2. For multi-source heterogeneous data, data cleaning and standardization are performed, and then a graph neural network model is used to achieve entity alignment, mapping nodes from different data sources but pointing to the same enterprise entity to a globally unique entity identifier. After entity alignment, for conflicting observations of the same attribute from different data sources, a credibility weighting mechanism based on time decay factor is introduced to calculate the unique definite value of the attribute, and finally a unified data model of the target enterprise of the user intent is obtained.

[0054] like Figure 3 As shown, this step mainly involves the fusion and conflict resolution calculations of multi-source heterogeneous data. Specifically, for the multi-source heterogeneous raw data obtained in the previous step, this step uses a graph neural network (GNN) model for entity alignment. This GNN model uses enterprise entities as nodes and common attributes (such as enterprise name, unified social credit code, and legal representative) across different data sources as edges to construct an enterprise entity graph. The model aggregates neighbor node information through multi-layer graph convolution operations, learns the vector representation of nodes, and maps nodes from different data sources but pointing to the same real entity to globally unique entity identifiers (Entity-IDs) based on vector similarity.

[0055] The training data for this GNN model consists of labeled enterprise entity matching pairs, including positive sample pairs (different representations of the same real-world entity) and negative sample pairs (entities of different entities). Taking enterprise names as an example, positive sample pairs include the same enterprise's registered name and the author's affiliation in a paper (e.g., the registered name "Xiaomi Technology Co., Ltd." and the author's affiliation "Xiaomi Technology", or the registered name "YOFC" and the author's affiliation "YOFC"). Negative sample pairs include the names of different enterprises (e.g., "Xiaomi Technology Co., Ltd." and "YOFC"). This GNN model is trained using a contrastive loss function. For positive sample pairs, the vector distance is calculated and minimized to make different representations of the same enterprise's name close in the vector space; for negative sample pairs, the vector distance is calculated and maximized (introducing boundary parameters) to separate different enterprise names in the vector space.

[0056] After entity alignment, to address potential conflicting observations from different data sources for the same attribute (e.g., company address), this invention introduces a reliability weighting mechanism based on a time decay factor. Suppose a certain company attribute has n data sources providing observations, and the observation value from the i-th data source is... The system presets static confidence weights for each data source. (For example, the weight for business registration is set at 0.9, the weight for the business data platform is set at 0.7, and the weight for self-reporting by enterprises is set at 0.5. These weights can be determined through historical data verification or expert scoring.) A time decay factor is also introduced. ,in This represents the time difference between the update time of this data record and the present time. This is an empirical attenuation coefficient (suggested value range: 0.01 to 0.1, used to adjust the impact of time on data reliability). Its unique truth value after fusion. The calculation formula is:

[0057] ;

[0058] for Extremely large (extremely old data) or In extreme cases with significant differences, such as when the time decay factors of all data sources approach zero, the result approximates a weighted average with static weights. Through the aforementioned weighted fusion calculation, data with higher reliability and more recent updates are given higher weights in the fusion result, thus obtaining a unique and definite value for that attribute. After cleaning and aligning conflicting heterogeneous data, a conflict-free, structurally unified Unified Data Model (UDM) is finally formed, providing high-quality input for subsequent steps.

[0059] S3. Based on the evaluation dimensions analyzed in step S1, determine the applicable indicators based on the pre-set evaluation indicator system, extract a series of feature data corresponding to the applicable indicators from the unified data model obtained in step S2, and calculate the score value of the user intent evaluation dimension. Then, based on the aforementioned globally unique entity identifier, construct a business association graph as structured support data for use in the subsequent report generation stage.

[0060] This step primarily involves feature calculation and association mining based on unified data. Specifically, this step uses the Unified Data Model (UDM) obtained in the previous step to perform secondary feature derivation, transforming basic attribute data into advanced indicators with business insights. For example, if the user's intent is to assess scientific and technological innovation capabilities, this step calculates the company's scientific and technological innovation capability score and extracts the feature data required for the assessment from the Unified Data Model (UDM). Table 1 below is a partial example of a pre-defined assessment indicator system (the scientific and technological innovation capability assessment report is a type of corporate credit report; therefore, scientific and technological innovation capability is used as a primary indicator here. A complete corporate credit assessment indicator system can also include other primary indicators; the assessment dimension in the user's intent usually corresponds to one or more primary indicators or next-level indicators in this assessment indicator system).

[0061] Table 1. Examples of some evaluation indicator systems

[0062]

[0063] For the scoring of secondary indicators, the feature data of multiple tertiary indicators subdivided under each secondary indicator are converted into a normalized feature fusion result. Depending on the situation, multimodal feature fusion technology can also be used to efficiently process the enterprise's numerical and textual data. The specific technical principles and optional implementation methods are conventional techniques and will not be elaborated here.

[0064] The following example uses Company A's "Research Platform" secondary indicator to illustrate how to convert its raw data into normalized feature values, and then obtain the normalized feature fusion result.

[0065] The “Research Platform” indicator consists of multiple tertiary indicators. Here, we simplify the example by selecting 6 tertiary indicators: number of manufacturing innovation centers, number of national-local joint engineering research centers, number of national-level key enterprise laboratories, number of enterprise R&D centers, number of academician workstations, and number of postdoctoral workstations. Assume that the feature data corresponding to these 6 tertiary indicators of Company A extracted from the Unified Data Model (UDM) are: x = (0, 2, 1, 5, 1, 2);

[0066] To eliminate the influence of different units of measurement, each indicator is normalized. The Min-Max standardization method is preferred.

[0067] ;

[0068] Assume the industry statistical scope is as shown in Table 2 below:

[0069] Table 2. Examples of industry statistical scope corresponding to the six tertiary indicators.

[0070]

[0071] The normalization result is: ;

[0072] The normalized tertiary indicators are then weighted and merged: ;

[0073] For example, set the weights w=(0.15,0.20,0.20,0.20,0.10,0.15), and F=0.34;

[0074] The normalized feature fusion result serves as the feature value of the secondary indicator "Scientific Research Platform," and together with the normalized feature fusion results (feature values) of other secondary indicators, it constitutes the input vector of the scoring model for the primary indicator "Scientific and Technological Innovation Capability." Specifically:

[0075] The normalized features of all the sub-indicators of the primary indicator are fused together to form a feature vector, which is then input into a trained deep learning model. The deep learning model outputs the classification probabilities in multiple pre-divided scoring intervals (categories). The comprehensive score of the primary indicator is obtained by weighted averaging.

[0076] During training, this deep learning model divides the scores of the primary indicators into multiple scoring intervals as the training objective. It inputs a feature vector composed of the normalized feature fusion results of all secondary indicators and outputs the probability corresponding to each scoring interval.

[0077] In this embodiment, the labels adopt a five-point system, dividing the comprehensive score into five scoring intervals (categories): [0-20], [20-40], [40-60], [60-80], and [80-100], which serve as the target for model learning.

[0078] Representative enterprises of different types were selected, and characteristic data for each enterprise were collected according to the evaluation index system in Table 1 above. Ground truth was obtained through expert scoring or other evaluation methods. For example, a training sample of no less than 5,000 enterprises covering multiple key industries such as manufacturing, information technology, and biomedicine was selected to ensure representativeness in terms of enterprise size (classified by registered capital) and geographical distribution. For each enterprise, characteristic data was collected according to the aforementioned evaluation index system, and no fewer than three experts with enterprise credit assessment qualifications were invited to score the enterprise independently. The average score was taken as the ground truth.

[0079] To enable the model to learn the uncertainty of the ratings, a soft label distribution is constructed (e.g., a Gaussian distribution centered on the ground truth is constructed and then mapped to the five rating intervals mentioned above).

[0080] In this way, the vector formed by the normalized feature fusion results (feature values) of all secondary indicators of each enterprise, the soft label (5 dimensions), and the corresponding ground truth value are used as a set of training sample data.

[0081] (1) Constructing the model network: An architecture consisting of a 4-layer deep learning network is adopted. This includes:

[0082] Input layer: Receives a feature vector composed of the fusion results of multiple normalized features (designed to have 128 neurons).

[0083] Two hidden layers: perform complex nonlinear feature transformation and learning, specifically using the Swish activation function and L2 regularization for nonlinear feature transformation.

[0084] Output layer: Generates the probabilities corresponding to the five rating intervals (categories).

[0085] (2) Application of training strategies: Multiple strategies are used during the training process to ensure the effectiveness of the model:

[0086] Regularization strategies: Use Dropout and L2 regularization to avoid overfitting and enhance the model's generalization ability.

[0087] Dynamic learning rate adjustment: Optimizes convergence efficiency during training.

[0088] Early stopping strategy: Terminate training early when model performance no longer improves to prevent overfitting.

[0089] In this way, the deep learning model outputs the classification probabilities of the above five scoring intervals (categories) for the current primary indicator, and obtains the score of the primary indicator by weighted average calculation.

[0090] The following example uses Company A to illustrate the specific calculation process of the enterprise science and technology innovation capability scoring method in this embodiment.

[0091] 1. Indicator Feature Vector: In this embodiment, the following eight secondary indicators are selected as input features: quality of results, innovation qualification, basic background, technological output, management capability, scientific research platform, talent reserve, and credit compliance. After standardization, normalization, and feature processing, the corresponding feature vectors are obtained as follows:

[0092] X = (0.88, 0.85, 0.90, 0.93, 0.87, 0.34, 0.91, 0.95).

[0093] 2. Probability Prediction

[0094] The deep learning model calculates the probability that a company falls into one of the five scoring intervals [0-20], [20-40], [40-60], [60-80], and [80-100] based on the input feature vector. The model outputs the probability value for each scoring interval, which is then normalized using the Softmax function. The calculation formula is as follows:

[0095] ;

[0096] in This represents the predicted probability of the i-th rating interval. This represents the original score (logits) corresponding to the i-th scoring interval output by the model, where K is the total number of scoring intervals. In this embodiment... , where e is the natural constant (approximately 2.71828).

[0097] The above feature vectors are input into a pre-trained deep learning model, which outputs the raw scores (logits) for each scoring interval, and then converts them into a probability distribution using the Softmax function.

[0098] The probability of Company A being in any of the five rating intervals is: P = (0.01, 0.03, 0.12, 0.39, 0.45).

[0099] 3. Weighted scoring

[0100] After obtaining the predicted probabilities for each scoring interval, the comprehensive score of the enterprise is obtained by weighting the representative values ​​of each interval. The calculation formula is as follows:

[0101] ;

[0102] Among them, Score: a comprehensive score of the enterprise's scientific and technological innovation capabilities, with a value range of [0, 100]; : The predicted probability of the i-th rating interval : The representative value corresponding to the i-th rating interval; K: The total number of rating intervals, in this embodiment .

[0103] The corresponding scoring intervals and their representative values ​​are shown in Table 3 below.

[0104] Table 3. Examples of scoring intervals and representative values ​​for each interval

[0105]

[0106] The overall score is calculated using probability weighting: Score = 0.01×10 + 0.03×30 + 0.12×50 + 0.39×70 + 0.45×90 = 74.8.

[0107] According to the calculation results, Company A's comprehensive score for scientific and technological innovation capability is 74.8 points, mainly distributed in the range of [60, 80] and [80, 100], with the highest probability in the high score range of [80–100], indicating that the company as a whole has strong scientific and technological innovation capability.

[0108] In this step, we take into account the diversity and uncertainty of the specific values ​​of the corresponding indicators of actual enterprises (such as numerical, textual, Boolean or empty), and the significant uncertainty and multi-dimensional coupling characteristics of user intent in actual business scenarios. Therefore, this embodiment can avoid the difficulty of effectively characterizing the boundary problems between different intervals by traditional deterministic evaluation methods based on rules or single regression models.

[0109] After the enterprise's science and technology innovation capability score is calculated, a business relationship graph is constructed based on the globally unique entity identifier (Entity-ID) stored in the UDM in step S2. This graph identifies the enterprise's industrial chain, industry development prospects, equity penetration path, actual controller chain, and related risk transmission path. Step S3 stores the calculated science and technology innovation capability score and relationship graph information as structured supporting data in the report agent context for use in the report generation stage in step S4.

[0110] S4. Based on the evaluation dimensions of user intent, recall the corresponding structured framework from the pre-built Prompt template library, and inject the score values ​​of the aforementioned evaluation dimensions obtained in step S3 and the feature data after format conversion into the Context area of ​​the reporting agent (large model); according to the constraint instructions set in the system prompt words, obtain the enterprise credit report; the constraint instructions require the large model to reason only based on the data injected in the Context area, and each conclusion must be labeled with the data source anchor point.

[0111] This step involves outputting a credit report based on dynamic templates and constraints. Specifically:

[0112] Based on the user's query intent, the system retrieves the corresponding structured framework (e.g., generating an enterprise science and technology innovation capability evaluation report) from the Prompt template library. The high-confidence UDM data and calculated features output in step S3 are converted to Markdown format and injected into the Context area of ​​the large model. Explicit constraints are set in the System Prompt, requiring the model to infer solely based on the injected Context data, and each conclusion must be annotated with its data source anchor.

[0113] To ensure data traceability, this embodiment utilizes the interceptor mechanism defined by the MCP protocol to add data citation hooks at the generated key conclusions, enabling audit trails: the interceptor monitors the model output stream, identifies text fragments involving specific values, associates these values ​​with the context-injected data through matching or semantic recognition, and automatically inserts citation hooks.

[0114] Taking a company address as an example: the original data sources are business registration (weight 0.9) and company address information from recruitment websites (weight 0.8). After conflict resolution in step S2, the fused value stored in UDM is "No. 77, Guanggu Avenue, East Lake High-tech Development Zone, Wuhan". When the report is generated, the interceptor detects this address value, automatically generates a traceability anchor, and outputs "Registered Address: No. 77, Guanggu Avenue, East Lake High-tech Development Zone, Wuhan [Anchor: Business Registration + Bidding Information]". Clicking this anchor will take you to the data fusion details page, which displays the original data source, weight, update time, and conflict resolution calculation process. In the final generated report, each numerical conclusion has a traceability anchor, enabling audit tracking and meeting the compliance requirements of the "Credit Reporting Business Management Measures" for traceable data sources.

[0115] The function of generating enterprise credit reports is integrated into the user's instant messaging client as a program module. It renders structured report data into a visual page and displays it on the user's instant messaging client interface. In response to the document export command triggered by the user in the interface, it converts the visual page into a document file in a preset format and sends the document file to the instant messaging client's chat window, or provides a download link for the document file.

[0116] For example, the enterprise credit report generation function is integrated into the user's WeChat Work client as a program module, such as WeChat Work skills or WeChat Work self-built applications.

[0117] In the instant messaging client's session window, generate and display a card message containing a visual page; in response to the user's triggering action on the card message, load and display the complete web page containing the visual page within the instant messaging client.

[0118] Then, in response to a user's document export command triggered in the interactive interface, the above-mentioned visual page can be converted into a document file in a preset format (such as PDF or Word format); the document file can be sent to the instant messaging client's chat window, or a download link for the document file can be provided.

[0119] In addition, the instant messaging client's interface can also display report subscription configuration options; in response to the user's subscription operation, the target company's credit information is automatically updated according to the preset monitoring cycle, and when information changes or risk events are detected, warning messages are pushed to the user through the instant messaging client.

[0120] This embodiment has at least the following advantages:

[0121] Improved data consistency and accuracy: Conflict resolution is achieved using a credibility weighting formula with a time decay factor, ensuring that multi-source data forms a unique truth value before being input into the large model. Compared to simple data concatenation, the consistency of the underlying data is significantly improved, fundamentally mitigating computational biases caused by cross-modal semantic gaps.

[0122] Compliance and controllability of generated content: The architecture of "pre-processing feature calculation + template constraint generation" implicitly ensures that the large model follows rigorous financial analysis logic, effectively suppressing "data illusions." Combined with data traceability anchoring technology, the automatically generated credit reports achieve commercial compliance levels in terms of accuracy and auditability while maintaining the flexibility of natural language interaction, meeting regulatory requirements such as the "Credit Reporting Business Management Measures."

[0123] In addition, by using standardized tool encapsulation of the MCP protocol to replace the traditional point-to-point hard-coded integration of APIs, the scheduling of multi-source data has achieved dynamic self-discovery, which greatly shortens the access and adaptation cycle of new data sources and realizes dynamic discovery and flexible scheduling of data sources.

[0124] Here is a simplified application example:

[0125] User input: "Generate an assessment report on the scientific and technological innovation capabilities of Company A".

[0126] Step 1: The LLM interprets the intent and uses the MCP protocol to call multiple MCP tools such as "Business Information", "Intellectual Property" and "Scientific Research Projects" to obtain the raw data of Company A in parallel.

[0127] Step Two: The system detects a discrepancy between the "company address" from the Administration for Industry and Commerce and a certain enterprise information platform. Entity alignment confirms that both address the same company (Company A). Based on the weights of the two data sources and their update times, a conflict resolution formula is applied to calculate a unique company address, which is then stored in the UDM (User Data Manager).

[0128] Step 3: The system reads the data corresponding to the set of secondary indicators related to the evaluation of scientific and technological innovation capabilities of Company A in UDM. First, it obtains the normalized feature fusion result by weighting and fusing the normalized feature values ​​of each of the subordinate tertiary indicators. Then, it forms a feature vector from the normalized feature fusion results of all secondary indicators and inputs it into the trained deep learning model. The deep learning algorithm outputs the probability of falling into each scoring interval. Then, it calculates the comprehensive score of the user's intention "scientific and technological innovation capability" based on the probability weighting, which is Score = 74.8.

[0129] Step 4: The system loads the "Science and Technology Innovation Capability Evaluation Report" template (structured framework), and injects Score = 74.8 and key supporting data (here, only "Number of Valid Patents" and "Provincial-Level Scientific Research Platform Qualification" are used as examples) into Prompt. The large model generates a report under constraints: "Company A's science and technology innovation capability score is 74.8, which is at a relatively high level [First Anchor Point]. This is mainly due to its possession of 20 valid patents [Second Anchor Point] and provincial-level scientific research platform qualification [Third Anchor Point]...".

[0130] It is important to note that in the above application examples, the evaluation of scientific and technological innovation capabilities is not solely based on the "number of valid patents" or "provincial-level scientific research platform qualifications." These are simplified examples provided for ease of understanding. The actual report will be based on the evaluation indicator system shown in Table 1 and various supporting data, resulting in a comprehensive report. For example, the supporting data also involves the business relationship graph constructed earlier. Therefore, the complete report will also cover the company's industry chain, equity penetration path, actual controller chain, and related risk transmission path, and the model's reasoning process is constrained by this specific information.

[0131] In one embodiment, a computer device is also provided, the processor of which provides computing and control capabilities, the computer device loading and running a computer program to implement the intelligent generation method for enterprise credit reports based on multi-source data fusion described in the above embodiments.

[0132] In one embodiment, a computer program product is also provided, including a computer program / instructions that, when executed by a processor, implement the steps of the intelligent generation method for enterprise credit reports based on multi-source data fusion described in the above embodiments.

[0133] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A method for intelligently generating enterprise credit reports through multi-source data fusion, characterized in that, include: S1. Receive natural language instructions input by the user, wherein the natural language instructions contain the name of the target enterprise and the evaluation dimensions; The natural language instructions are parsed for intent, and multi-source heterogeneous data about the target enterprise is obtained through a pre-built standardized toolset and dynamic discovery mechanism based on the MCP protocol. S2. For the multi-source heterogeneous data, data cleaning and standardization are performed, and then a graph neural network model is used to achieve entity alignment, mapping nodes from different data sources but pointing to the same enterprise entity to a globally unique entity identifier; after entity alignment, for conflicting observations of the same attribute from different data sources, a credibility weighting mechanism based on time decay factor is introduced to calculate the unique definite value of the attribute, and finally a unified data model for the target enterprise is obtained. S3. Based on the evaluation dimensions analyzed in step S1, determine the applicable indicators based on the pre-set evaluation indicator system, extract a series of feature data corresponding to the applicable indicators from the unified data model obtained in step S2, and comprehensively calculate the score value of the evaluation dimension; then, based on the globally unique entity identifier, construct a business association graph as structured support data for use in the subsequent report generation stage. S4. Based on the evaluation dimensions, recall the corresponding structured framework from the pre-built Prompt template library, and inject the score value of the evaluation dimensions obtained in step S3 and the feature data after format conversion into the Context area of ​​the reporting agent. Generate enterprise credit reports based on the constraint instructions set in the system prompts; the constraint instructions require the report agent to reason only based on data injected from the Context region, and each conclusion must be labeled with the data source anchor point.

2. The intelligent generation method for enterprise credit reports based on multi-source data fusion according to claim 1, characterized in that, In step S2, a credibility weighting mechanism based on a time decay factor is introduced. Assume there are n data sources providing observations for the target enterprise attributes, and the observation value from the i-th data source is... The system presets static confidence weights for each data source. And introduce a time decay factor ,in This represents the time difference between the update time of the data record in this data source and the present time. The empirical decay coefficient is the uniquely determined value of the attribute. The calculation formula is as follows: 。 3. The intelligent generation method for enterprise credit reports based on multi-source data fusion according to claim 1, characterized in that, In step S2, the graph neural network model constructs an enterprise entity graph with enterprise entities as nodes and common attributes between different data sources as edges. The graph neural network model aggregates neighbor node information through multi-layer graph convolution operations, learns the vector representation of nodes, and maps nodes from different data sources but pointing to the same real entity to globally unique entity identifiers based on vector similarity.

4. The intelligent generation method for enterprise credit reports based on multi-source data fusion according to claim 1, characterized in that, In step S2, the data cleaning and standardization includes: Complete the missing data; Identify and process abnormal noise points in the data; Scaling numerical features to a uniform dimension; By referring to the pre-set evaluation index system, the raw feature data is transformed into structured feature data.

5. The intelligent generation method for enterprise credit reports based on multi-source data fusion according to claim 1, characterized in that, In step S3, the comprehensive calculation specifically includes: Suppose that the evaluation dimension mentioned in step S1 corresponds to a primary indicator in the pre-set evaluation indicator system. The applicable indicators of this evaluation dimension include multiple secondary indicators in the pre-set evaluation indicator system, and each secondary indicator is further subdivided into multiple tertiary indicators. For each secondary indicator, the feature data of its multiple sub-tertiary indicators are converted into normalized feature values, and then weighted and fused to obtain the normalized feature fusion result. The normalized feature fusion results of all secondary indicators under the primary indicator are combined to form a feature vector, which is then input into the trained deep learning model. During training, the deep learning model divides the score of the primary indicator into multiple scoring intervals as the training target. By outputting the probability corresponding to each scoring interval, the model calculates the final comprehensive score by weighting the score and the corresponding probability value of each scoring interval.

6. The intelligent generation method for enterprise credit reports based on multi-source data fusion according to claim 1, characterized in that, In step S3, the construction of the business association graph is based on the globally unique entity identifier to identify the enterprise's industrial chain, industrial development prospects, enterprise equity penetration path, actual controller link and associated risk transmission path.

7. The intelligent generation method for enterprise credit reports based on multi-source data fusion according to claim 1, characterized in that, In step S4, the interceptor mechanism defined by the MCP protocol is used to monitor the model output stream in real time, identify text fragments involving specific values, and associate the specific value with the data injected into the Context area through matching or semantic recognition, and automatically insert data source anchors.

8. The intelligent generation method for enterprise credit reports based on multi-source data fusion according to claim 1, characterized in that, In step S4, the function of generating enterprise credit reports is integrated into the user's instant messaging client as a program module. The structured report data is rendered into a visual page and displayed on the user's instant messaging client's interactive interface. In response to a document export command triggered by the user in the interactive interface, the visualization page is converted into a document in a preset format, and the document is sent to the session window of the instant messaging client, or a download link for the document is provided.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the intelligent generation method for enterprise credit reports based on multi-source data fusion as described in claim 1.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the intelligent generation method for enterprise credit reports based on multi-source data fusion as described in claim 1.

Citation Information

Patent Citations

  • Risk assessment method and device based on model context protocol, equipment and medium

    CN120598693A

  • Intelligent marketing document generation device and method based on multi-source data fusion

    CN120822507A

  • Big data platform asset intelligent sensing method based on LLM and customizable MCP

    CN121579550A

  • Intelligent Data Fabric Query Engine

    US20250363085A1

  • Knowledge graph extraction

    WO2025085237A1