An enterprise credit intelligent analysis method and system fusing a knowledge graph and a dialogue engine

By constructing evidence chain data packages and introducing conflict grouping in corporate credit analysis, the problem of unclear relationship between conclusions and data sources in corporate credit analysis is solved, and structured credit analysis results are output, improving the reliability and consistency of the analysis.

CN122263903APending Publication Date: 2026-06-23HUACHENG CHITONG (BEIJING) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUACHENG CHITONG (BEIJING) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
Filing Date
2026-03-27
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

In existing corporate credit analysis technologies, corporate credit knowledge graphs lack a path or subgraph organization method to support facts. Multi-source credit data have differences in caliber and information conflicts, and the generative responses of large language models may produce inconsistent content, resulting in unclear relationships between conclusions and data sources and a lack of reliability.

Method used

By forming evidence chain data packages from credit conclusion units and paths or subgraphs, and introducing conflict grouping and confidence score ranking methods, an enterprise credit knowledge graph is constructed. Combined with a dialogue engine, a structured credit analysis result is generated, including the main conclusion, alternative conclusions, confidence scores, and evidence chain data packages.

Benefits of technology

It achieves a one-to-one correspondence between credit conclusions and supporting facts, improves the structure of conclusion expression, reduces the fluctuation of results caused by differences from a single source, and enhances the reliability and consistency of corporate credit analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122263903A_ABST
    Figure CN122263903A_ABST
Patent Text Reader

Abstract

The application provides an enterprise credit intelligent analysis method and system fusing a knowledge graph and a dialogue engine, and relates to the technical field of enterprise credit risk analysis.The application firstly acquires multi-source credit records of a target enterprise and associates source information fields to construct an enterprise credit knowledge graph; analyzes a credit query request to obtain a credit conclusion unit and a graph query condition; searches a candidate fact set and forms an evidence chain data package; groups according to a preset conflict rule and calculates a confidence score based on a preset scoring rule to determine a main conclusion and an alternative conclusion and generate a conflict explanation, which is output to a dialogue engine to generate a dialogue reply.The application associates the credit conclusion with the corresponding evidence chain, introduces a conflict grouping and confidence sorting mode, and improves the structured expression ability and conclusion reliability of the enterprise credit analysis result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of enterprise credit risk analysis technology, and in particular to an intelligent enterprise credit analysis method and system that integrates knowledge graphs and dialogue engines. Background Technology

[0002] Existing enterprise credit analysis products and methods typically rely on multi-source data such as business registration, judicial litigation, administrative penalties, bidding and tendering, and public opinion information. They organize the enterprise entity and its related relationships, including equity, guarantees, and transactions, and then output risk warnings or classification results. Publicly available patents have proposed solutions that construct a unified information model and domain ontology for enterprise credit big data, forming an enterprise credit knowledge graph. Based on this knowledge graph, feature data is input into a risk control model, which outputs classification results for enterprise risk detection. On the other hand, reviews of knowledge graphs generally consider knowledge extraction, knowledge fusion, and knowledge reasoning as crucial steps in graph construction and updating to improve the accuracy and usability of entity and relationship representations.

[0003] As large models and conversational interactions expand in enterprise intelligence applications, more and more solutions are trying to introduce external knowledge sources into the dialogue process, adopt retrieval enhancement generation ideas to improve the knowledge coverage and timeliness of answers, and various RAG variants (including GraphRAG that incorporates graph structure information) have emerged for knowledge-intensive tasks to support multi-hop relation retrieval, relation context organization, and more complex question answers.

[0004] However, significant shortcomings remain in corporate credit scenarios: First, some risk detection solutions based on corporate credit knowledge graphs focus on feature acquisition and model classification output. They often lack a complete organizational structure for supporting factual paths or subgraphs and their source fields, leading to unclear correspondence between conclusions and data sources. Second, multi-source credit data objectively exhibits differences in definitions and information conflicts. While knowledge graph reviews emphasize knowledge fusion for eliminating ambiguity and forming basic factual expressions, there is still a lack of an integrated process for credit question answering to uniformly group, score, and rank conflicting facts at the conclusion level to form readable conflict explanations. Third, large language models may still produce inconsistent content in generative responses. Related reviews indicate that retrieval enhancement and other methods can alleviate this problem, but further reductions in the risk of misleading output due to inconsistent outputs are needed in high-reliability business scenarios. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a method and system for intelligent enterprise credit analysis that integrates knowledge graphs and dialogue engines. By combining credit conclusion units with evidence chain data packages formed by paths or subgraphs, and introducing conflict grouping and confidence score ranking methods, the intelligent enterprise credit analysis output under structured support is realized.

[0006] To achieve the above objectives, the present invention provides the following solution: A method for intelligent enterprise credit analysis that integrates knowledge graphs and dialogue engines includes: Obtain the multi-source credit records of the target company and associate source information fields with each multi-source credit record; Based on multi-source credit records, entities and relationships are extracted and an enterprise credit knowledge graph is constructed, and the source information fields are written into the attributes of the corresponding nodes or edges. The system receives credit query requests through a dialogue engine, parses the credit query requests to obtain credit conclusion units and graph query conditions; Based on the graph query conditions, a set of candidate facts is retrieved in the enterprise credit knowledge graph. From the set of candidate facts, the path or subgraph corresponding to the credit conclusion unit is extracted to form an evidence chain data package. The evidence chain data package includes at least the node edge sequence and its source information field. The candidate fact set is grouped according to the preset conflict rules to obtain the consistent candidate set and the conflict candidate set, and the conflict type of the conflict candidate set is determined. The confidence score is calculated for the conflict candidate set based on the preset scoring rules. The main conclusion and alternative conclusions are determined and conflict descriptions are generated. The main conclusion, alternative conclusions, confidence scores, evidence chain data packages and conflict descriptions are output to the dialogue engine to generate a dialogue response. The dialogue response displays the confidence score and cites the evidence chain data packages for the credit conclusion unit.

[0007] Preferably, the source information field associated with each multi-source credit record includes: Generate a source information field for each multi-source credit record. The source information field shall include at least the source identifier. Establish a one-to-one correspondence between the source information field and the corresponding multi-source credit records, and store them together; When multi-source credit records are used to construct an enterprise credit knowledge graph, the source information field is called as the data source for subsequent writing of node or edge attributes.

[0008] Preferably, the process involves extracting entities and relationships from multi-source credit records and constructing an enterprise credit knowledge graph, including: Identify corporate entities, natural person entities, and event entities from multi-source credit records, and generate entity nodes; Identify at least one of the following relationships among entity nodes from multi-source credit records: equity relationship, employment relationship, guarantee relationship, litigation relationship, and penalty relationship, and generate relationship edges; A corporate credit knowledge graph is formed based on entity nodes and relation edges, and the type identifiers of entity nodes and relation edges are retained in the corporate credit knowledge graph.

[0009] Preferably, the credit query request is parsed to obtain the credit conclusion unit and graph query conditions, including: Determine the target company indication information from the credit inquiry request, and determine the scope of target companies corresponding to the credit conclusion unit accordingly; Determine the conclusion category indication information from the credit inquiry request, and determine the conclusion category of the credit conclusion unit accordingly; Based on the target enterprise scope and conclusion category, generate graph query conditions; graph query conditions shall include at least one or more of the following: node type restriction, relationship type restriction, and attribute filtering conditions.

[0010] Preferably, the evidence chain data package is formed by extracting the path or subgraph corresponding to the credit conclusion unit from the candidate fact set, including: The starting element is determined based on the credit conclusion unit, and the starting element is the target node or the target relationship edge; Centered on the initial element, the path set or subgraph set is obtained in the enterprise credit knowledge graph according to the graph query conditions to form a candidate fact set; According to the preset priority rules, at least one path or at least one subgraph is selected from the path set or subgraph set as the evidence chain data packet; wherein the preset priority rules include at least the priority rules for the number of nodes and edges of the path or subgraph and the priority rules for the source information field; The evidence chain data packet is organized into a data structure containing node edge sequences and their source information fields, and a correspondence description between the credit conclusion unit and the node edge sequence is established.

[0011] Preferably, the candidate fact set is grouped according to a preset conflict rule to obtain a consistent candidate set and a conflict candidate set, and the conflict type of the conflict candidate set is determined, including: For the same credit conclusion unit, determine the set of consistency fields used for comparison in the candidate fact set; Candidate facts whose consistency field set meets the preset consistency conditions are classified into the consistency candidate set; Candidate facts whose consistency field set in the candidate fact set does not meet the preset consistency conditions are classified into the conflict candidate set; Based on the field categories where differences occur in the consistency field set, the conflict type of the conflict candidate set is determined, and a description of the difference fields used for conflict explanation is generated.

[0012] Preferably, the conflict types include at least value conflicts, relational conflicts, temporal conflicts, and source difference conflicts, wherein: Value conflict occurs when candidate facts corresponding to the same credit assessment unit have different values ​​in the same attribute field. A relationship conflict occurs when the candidate facts corresponding to the same credit conclusion unit are different in the relationship edge type or the relationship edge endpoint entity. A time-series conflict occurs when the candidate facts corresponding to the same credit conclusion unit meet the preset conflict conditions in terms of the time or version information indicated by the source information field. Source discrepancy conflict occurs when the candidate facts corresponding to the same credit conclusion unit meet the preset conflict conditions in the source identifier indicated by the source information field.

[0013] Preferably, the confidence score of the conflict candidate set is calculated based on a preset scoring rule to determine the main conclusion and alternative conclusions and generate a conflict explanation, including: For each candidate fact in the conflict candidate set, determine the source authority level, the freshness of the collection time level, the fact consistency level, and the historical hit rate level; Each level is mapped to a level score, and the level scores are combined according to a preset scoring rule to obtain a confidence score. The conclusion corresponding to the candidate fact with the highest confidence score is determined as the main conclusion, and the conclusions corresponding to the candidate facts that meet the preset score conditions, other than the main conclusion, are determined as alternative conclusions. Based on the difference field description, the candidate facts corresponding to the main conclusion, and the candidate facts corresponding to the alternative conclusions, a conflict statement is generated.

[0014] Preferably, the main conclusion, alternative conclusions, confidence scores, evidence chain data packets, and conflict explanations are output to the dialogue engine to generate a dialogue response, including: Write the main conclusion and its corresponding confidence score into the corresponding position in the dialogue response for the credit conclusion unit; When there is a conflict candidate set, the alternative conclusions and conflict explanations are written into the corresponding position of the credit conclusion unit in the dialogue response; In the dialogue response, configure the reference identifier of the evidence chain data packet for the credit conclusion unit, so that the main conclusion and the alternative conclusion correspond to the node edge sequence in the evidence chain data packet respectively; The conflict description is associated with the reference identifier of the evidence chain data packet, so that the difference fields involved in the conflict description describe the source information fields in the corresponding evidence chain data packet.

[0015] A corporate credit intelligence analysis system integrating knowledge graphs and dialogue engines includes: The multi-source credit information collection unit is used to acquire multi-source credit information records of the target enterprise and associate source information fields with each multi-source credit information record; The knowledge graph construction unit is used to extract entities and relationships based on multi-source credit records and construct an enterprise credit knowledge graph, writing the source information fields into the attributes of the corresponding nodes or edges; The query parsing unit is used to receive credit query requests through the dialogue engine, parse the credit query requests to obtain credit conclusion units and graph query conditions; The evidence chain extraction unit is used to retrieve a set of candidate facts in the enterprise credit knowledge graph based on graph query conditions, and extract the path or subgraph corresponding to the credit conclusion unit from the set of candidate facts to form an evidence chain data package; the evidence chain data package includes at least the node edge sequence and its source information field; The conflict grouping unit is used to group the candidate fact set according to the preset conflict rules to obtain the consistent candidate set and the conflict candidate set, and to determine the conflict type of the conflict candidate set; The conclusion generation and output unit is used to calculate the confidence score of the conflict candidate set based on the preset scoring rules, determine the main conclusion and the alternative conclusions and generate conflict descriptions, and output the main conclusion, alternative conclusions, confidence scores, evidence chain data packets and conflict descriptions to the dialogue engine to generate a dialogue response; the dialogue response displays the confidence score and cites the evidence chain data packets to the credit conclusion unit.

[0016] The present invention discloses the following beneficial effects: This invention associates source information fields with each multi-source credit record during acquisition and writes these fields into the attributes of corresponding nodes or edges when constructing an enterprise credit knowledge graph. This allows subsequent credit conclusions to be directly linked to specific node and edge sequences in the graph, achieving a one-to-one correspondence between credit conclusions and supporting facts. Compared to existing solutions that only output risk classification results or numerical scores, this invention forms a data structure expression method centered on paths or subgraphs at the conclusion level, improving the structured nature of the conclusion expression.

[0017] After generating the credit conclusion unit, this invention retrieves a set of candidate facts based on graph query conditions, and extracts paths or subgraphs from the candidate fact set to form an evidence chain data package, thus establishing the credit conclusion on a clear graph structure. This technical solution avoids the problem of unclear evidence caused by outputting conclusions based solely on statistical features or implicit model parameters, ensuring that the credit analysis results have clear data sources and relationship chains at the structural level.

[0018] This invention performs grouping processing on the candidate fact set corresponding to the same credit assessment conclusion unit, distinguishing between consistent candidate sets and conflicting candidate sets, and further determining the conflict type. It introduces a method to differentiate between the main conclusion and alternative conclusions during the conclusion generation stage. This processing method can retain discrepancies when there are differences in multi-source data, rather than simply covering or discarding some data, thereby improving the adaptability of enterprise credit assessment in complex data environments.

[0019] This invention calculates confidence scores for conflict candidate sets based on preset scoring rules, and uses these scores to determine the main conclusion and alternative conclusions, providing a quantifiable basis for ranking the output conclusions. Compared to single-rule judgment or fixed-priority selection methods, this scheme can form a more stable ranking result under the combined influence of multiple factors, reducing result fluctuations caused by differences from a single source.

[0020] This invention outputs the main conclusion, alternative conclusions, confidence scores, evidence chain data packages, and conflict explanations to the dialogue engine, and displays the confidence scores and cites the evidence chain data packages in the dialogue response, ensuring data layer consistency between the dialogic interaction results and the enterprise credit knowledge graph. This technical solution maintains the dialogic interaction format while ensuring that the output content always corresponds to the structured graph data, thereby improving the reliability and consistency of enterprise credit analysis in practical business applications. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A flowchart of the method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the relationship between multi-source credit records provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the enterprise credit knowledge graph structure provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of a dialogue response structure based on a credit conclusion unit provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the system structure provided in an embodiment of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] The purpose of this invention is to provide a method and system for intelligent enterprise credit analysis that integrates knowledge graphs and dialogue engines. By simultaneously presenting the main conclusion, alternative conclusions, confidence scores, and evidence chain data packages in the dialogue response, the structured expression of enterprise credit analysis results and the retention of discrepancy information are achieved.

[0025] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0026] Figure 1 A flowchart of the method provided in the embodiments of the present invention, such as Figure 1 As shown, this invention provides a method for intelligent enterprise credit analysis that integrates knowledge graphs and dialogue engines, including: Step 100: Obtain the multi-source credit records of the target enterprise and associate the source information field with each multi-source credit record; Step 200: Extract entities and relationships based on multi-source credit records and construct an enterprise credit knowledge graph, writing the source information fields into the attributes of the corresponding nodes or edges; Step 300: Receive the credit query request through the dialogue engine, parse the credit query request to obtain the credit conclusion unit and graph query conditions; Step 400: Based on the graph query conditions, retrieve the candidate fact set in the enterprise credit knowledge graph, and extract the path or subgraph corresponding to the credit conclusion unit from the candidate fact set to form an evidence chain data package; the evidence chain data package shall at least include the node edge sequence and its source information field; Step 500: Group the candidate fact set according to the preset conflict rules to obtain the consistent candidate set and the conflict candidate set, and determine the conflict type of the conflict candidate set; Step 600: Calculate the confidence score for the conflict candidate set based on the preset scoring rules, determine the main conclusion and alternative conclusions and generate conflict descriptions, and output the main conclusion, alternative conclusions, confidence scores, evidence chain data packets and conflict descriptions to the dialogue engine to generate a dialogue response; the dialogue response displays the confidence score and cites the evidence chain data packets to the credit conclusion unit.

[0027] Specifically, in step 100 of this embodiment, the main identifier of the target enterprise is first determined, and multi-source credit records related to the target enterprise are obtained from at least two types of credit data sources. These credit data sources include, but are not limited to, business registration information, judicial rulings and enforcement information, administrative penalty information, bidding and winning bid announcements, and news and public opinion information. For each obtained multi-source credit record, this embodiment generates a source information field. The source information field is used to characterize the source attribute of the multi-source credit record and includes at least a source identifier. The source identifier is a preset discrete category value used to distinguish different data source types. For example, business registration is set as source identifier 1, judicial data as source identifier 2, administrative penalty data as source identifier 3, bidding data as source identifier 4, and public opinion information as source identifier 5. To ensure the repeatability of the source identifier generation process, this embodiment pre-establishes a correspondence table between source types and source identifiers, which contains at least 5 source type records. When a multi-source credit record is collected, this embodiment retrieves the corresponding source identifier from the correspondence table based on the source type of the multi-source credit record and writes the source identifier into the source information field of the multi-source credit record. When the same multi-source credit record matches multiple source types simultaneously, this embodiment determines a unique source type according to a preset priority, which is arranged from high to low in the order of judicial, administrative penalty, business registration, bidding, and public opinion information, to avoid source mismatches. The existence of multiple values ​​for the identifier leads to inconsistencies in subsequent processing. As shown in Table 1, the above-mentioned source types are all information categories publicly available under the current legal framework of my country and have objective existence. Judicial judgments and enforcement information, administrative penalty information, and business registration information are all legally disclosed information with legal effect. Bidding announcement information and public opinion information are legally disclosed or factually disclosed information, but do not directly produce legal judgment effect. Therefore, this embodiment sets judicial information as priority 1, administrative penalty information as priority 2, business registration information as priority 3, bidding information as priority 4, and public opinion information as priority 5, which is reasonable and has an objective basis.

[0028] Table 1. Example Table of Source Types and Priority Classification of Multi-Source Credit Records

[0029] In this embodiment, a record identifier is generated for each multi-source credit record, and the record identifier and the source information field are stored together as two fields of the same record item. The record identifier is used to uniquely indicate the multi-source credit record in subsequent steps. The record identifier consists of the target enterprise entity identifier, the source identifier, and a time sequence number. The time sequence number is an incrementing integer generated according to the order of collection, such as record 1, record 2, up to record 50, to ensure that multiple records of the same target enterprise under the same source identifier can still be distinguished. Subsequently, when the multi-source credit records are used to construct an enterprise credit knowledge graph, this embodiment uses the record identifier as an association field of nodes or edges during the extraction of entities and relationships, and writes the source information field into the attribute of the corresponding node or edge: when the multi-source credit record is extracted as an entity node, the source identifier is written into the source attribute of the entity node; when the multi-source credit record is extracted as a relationship edge, the source identifier is written into the source attribute of the relationship edge; thus, the nodes or edges obtained from the multi-source credit records all carry their corresponding source information fields. To improve data integrity, this embodiment performs a field integrity check before writing the source identifier. The check items include at least whether the record identifier is empty and whether the source identifier falls into a preset value set, which includes at least 1 to 5. When the source identifier is not in the preset value set, this embodiment marks the multi-source credit record as a record to be completed, and completes and writes it when data of the same source type is obtained again in the future.

[0030] Optionally, in step 200 of this embodiment, based on the multi-source credit records obtained in step 100, this embodiment performs structured parsing on each multi-source credit record to identify corporate entities, natural person entities, and event entities and generate entity nodes. Here, corporate entities refer to market entities that can be identified by a unified name or unified identifier; natural person entities refer to individual entities that can be identified by a name and at least one piece of auxiliary identity information; and event entities refer to records of matters related to corporate credit, which can be independently categorized and have time attributes. These records of matters include, for example, litigation cases, administrative penalty decisions, guarantee matters, or equity change matters. To ensure that the identification results are reusable, this embodiment establishes an entity type set, which includes at least three categories, corresponding to corporate entities, natural person entities, and event entities, respectively. At least two key fields are extracted for each multi-source credit record as the basis for entity identification. These key fields include a name field and an identifier field. The identifier field preferentially uses the publicly registered unified social credit code or registration number. If the identifier field cannot be obtained, a combination of the name field and the registered address field is used for differentiation. For records that cannot be directly matched as unique entities, this embodiment sets an entity merging threshold and performs entity merging processing. The entity merging threshold is that the number of identical key fields is not less than 2. For example, when two records have the same name field and the same unified social credit code, they are determined to be the same enterprise entity and merged into one entity node. When only the name field is consistent but the unified social credit code is missing, the registration address field is further compared. If the registration address field is consistent, they are merged into one entity node; otherwise, they are each retained as two entity nodes. After generating entity nodes, this embodiment writes an entity type identifier to each entity node. The entity type identifier takes the value 1, 2, or 3, representing an enterprise entity, a natural person entity, and an event entity, respectively. The source information field from step 100 is also written into the attributes of the corresponding entity node.

[0031] In this embodiment, as Figure 2As shown, after generating entity nodes, this embodiment further identifies relationships between entity nodes from multi-source credit records and generates relationship edges. These relationship edges include at least one of the following: equity relationship, employment relationship, guarantee relationship, litigation relationship, and penalty relationship. To ensure clear criteria for relationship identification, this embodiment pre-establishes a set of relationship types, which includes at least five categories, corresponding to equity relationships, employment relationships, guarantee relationships, litigation relationships, and penalty relationships. A set of relationship determination fields is configured for each type of relationship. For example, the set of relationship determination fields for equity relationships includes at least the investor field, the investee field, and the shareholding ratio field; the set of relationship determination fields for employment relationships includes at least the incumbent field, the employing company field, and the position field; the set of relationship determination fields for litigation relationships includes at least the party field, the case number field, and the case type field; and the set of relationship determination fields for penalty relationships includes at least the penalized entity field, the penalizing authority field, and the penalty decision field. When a multi-source credit record satisfies the condition that the complete number of fields in the relation determination field set of a certain relation type is not less than 3, this embodiment determines that the record corresponds to the relation type and generates a relation edge with two entity nodes as endpoints. Simultaneously, a relation type identifier is written to the relation edge, with values ​​ranging from 1 to 5, representing equity relationships, employment relationships, guarantee relationships, litigation relationships, and penalty relationships, respectively. Subsequently, this embodiment forms an enterprise credit knowledge graph based on the entity nodes and the relation edges, and retains the type identifiers of the entity nodes and relation edges in the enterprise credit knowledge graph. At the same time, the source information field from step 100 is written into the attributes of the corresponding relation edge, ensuring that entity nodes and relation edges formed from the same multi-source credit record carry consistent source information fields, thereby completing the construction process of step 200.

[0032] Figure 3 This diagram illustrates the structure of a corporate credit knowledge graph built upon multi-source credit records. The corporate credit knowledge graph uses the target company "Company A" as its core node, around which multiple entity nodes and relationship edges are constructed. These include natural person entities Zhang San, Li Si, and Wang Wu; corporate entities Company B and Company C; event entities such as Case No. 2021S123, a fine of 500,000 yuan, and equipment procurement contracts; and relationship edges including employment relationships, actual control relationships, equity relationships, litigation relationships, administrative penalty relationships, and external guarantee relationships. All entity nodes and relationship edges together constitute the graph structure, used to carry source information fields and provide a structured data foundation for subsequent graph queries, candidate fact set retrieval, and evidence chain data packet formation.

[0033] Further, in step 300 of this embodiment, after receiving the credit query request input by the user through the dialogue engine, this embodiment performs semantic element decomposition on the credit query request to obtain target enterprise indication information and conclusion category indication information. Here, target enterprise indication information refers to content fragments that can uniquely or nearly uniquely identify an enterprise entity, including at least one of the enterprise name and unified social credit code, or a combination of enterprise name and registered address; when the credit query request contains both the enterprise name and unified social credit code, this embodiment prioritizes using the unified social credit code to determine the scope of the target enterprise; when only the enterprise name is included, this embodiment retrieves the set of enterprise entity nodes matching the enterprise name in the enterprise credit knowledge graph, and uses the number of the enterprise entity node set as the ambiguity number; if the ambiguity number is greater than 1, this embodiment further extracts the regional indication words or address indication words appearing in the credit query request as auxiliary limiting conditions to converge the scope of the target enterprise from the set of enterprise entity nodes to no more than 3 enterprise entity nodes; if it is still greater than 3, this embodiment maintains the scope of the target enterprise as the set of enterprise entity nodes and presents a multi-entity result when generating a dialogue response in the subsequent process. The target enterprise scope here refers to the set of enterprise entity nodes covered by this query, which can be one or more enterprise entity nodes. Simultaneously, this embodiment extracts conclusion category indication information from the credit query request. This conclusion category indication information indicates the type of credit conclusion the user expects to obtain, such as equity structure, senior management appointments, external guarantees, litigation status, and administrative penalties. This embodiment pre-establishes a conclusion category thesaurus, which contains at least 10 conclusion category entries, and configures a set of synonym trigger words for each conclusion category entry. When a keyword in the credit query request matches the set of synonym trigger words, this embodiment determines the conclusion category of the credit conclusion unit as the matched conclusion category entry. When two or more conclusion category entries match simultaneously, this embodiment selects the first two as parallel credit conclusion units in order of appearance.

[0034] In this embodiment, after determining the target enterprise scope and conclusion category, graph query conditions are generated based on the target enterprise scope and conclusion category. Here, graph query conditions refer to a limited set used to locate candidate fact sets in the enterprise credit knowledge graph. The limited set includes at least one or more of node type limitations, relationship type limitations, and attribute filtering conditions. To ensure that the graph query conditions have clear construction rules, this embodiment pre-establishes a mapping table from conclusion categories to graph query conditions. This mapping table contains at least 10 mapping records, each containing at least two of the following three types of fields: conclusion category, node type limitation, relationship type limitation, and attribute filtering conditions. For example, when the conclusion category is equity structure, this embodiment limits the node type to corporate entities and natural person entities, limits the relationship type to equity relationships, and sets the attribute filtering condition to a non-empty shareholding ratio field. When the conclusion category is litigation, this embodiment limits the node type to corporate entities and event entities, limits the relationship type to litigation relationships, and sets the attribute filtering condition to a non-empty case number field and a case type field belonging to a preset case type set. The preset case type set contains at least three categories, corresponding to civil, administrative, and enforcement cases, respectively. For a target enterprise scope containing multiple enterprise entity nodes, this embodiment generates one set of graph query conditions for each enterprise entity node and combines multiple sets of graph query conditions into a set of parallel query conditions. For a case containing multiple credit conclusion units, this embodiment generates graph query conditions for each credit conclusion unit and forms a set of multiple conclusion query conditions, so that when retrieving candidate fact sets in the enterprise credit knowledge graph, the corresponding retrieval can be performed according to the target enterprise scope and the conclusion category.

[0035] Furthermore, in step 400 of this embodiment, after obtaining the credit conclusion unit and graph query conditions in step 300, this embodiment first determines the starting element based on the credit conclusion unit. Here, the starting element refers to the graph element that serves as the starting point for retrieval in the enterprise credit knowledge graph; the graph element is a target node or a target relationship edge. When the conclusion category of the credit conclusion unit corresponds to the enterprise entity's own attributes or a list of entities, this embodiment determines the enterprise entity node within the target enterprise scope as the target node as the starting element; for example, when the conclusion category is litigation, the enterprise entity node corresponding to the target enterprise is used as the starting element. When the conclusion category of the credit conclusion unit corresponds to the attributes of a specific relationship edge or a list of relationship edges, this embodiment determines the relationship edge indicated by the relationship type limitation corresponding to the conclusion category as the target relationship edge as the starting element; for example, when the conclusion category is external guarantee and the relationship type in the graph query conditions is limited to guarantee relationship, the guarantee relationship edge with the target enterprise as one endpoint is used as the starting element. To avoid an unbounded expansion of the candidate fact set due to too many starting elements, this embodiment sets an upper limit on the number of starting elements, which is 50. When the number of starting elements exceeds 50, this embodiment prioritizes retaining starting elements whose source identifiers corresponding to the source information field belong to the judicial or administrative penalty category, and truncates the remaining starting elements to 50 according to the order of collection.

[0036] In this embodiment, after determining the starting element, the embodiment uses the starting element as the center and obtains a set of paths or a set of subgraphs in the enterprise credit knowledge graph according to the graph query conditions, and forms a set of candidate facts accordingly. Here, the set of paths refers to the set of node edge sequences that start from the starting element and reach adjacent nodes sequentially along the relationship edges; the set of subgraphs refers to the set of graph elements centered on the starting element and containing several nodes and several relationship edges within a preset hop count range. The preset hop count range is determined by the conclusion category. In this embodiment, a correspondence table between the conclusion category and the hop count range is established in advance, and the correspondence table contains at least 10 records; for example, the hop count range is set to 3 for equity structure to cover the penetrating link of enterprise entity—equity relationship—enterprise entity or natural person entity; the hop count range is set to 2 for litigation situations to cover the direct association of enterprise entity—litigation relationship—event entity; and the hop count range is set to 2 for employment relationship to cover the association of natural person entity—employment relationship—enterprise entity. When obtaining the path set or subgraph set, this embodiment uses both the node type limitation and the relationship type limitation for filtering: only when the adjacent node type belongs to the node type limitation and the adjacent relationship edge type belongs to the relationship type limitation, is the corresponding node edge sequence included in the path set or the subgraph set. To control the size of the candidate fact set, this embodiment sets the upper limit of the candidate fact set capacity to 200; when the candidate fact set exceeds 200, this embodiment prioritizes retaining paths or subgraphs with no more than 8 nodes and edges, and truncates the remaining part to 200 nodes and edges in ascending order of node and edge count.

[0037] In this embodiment, after obtaining the path set or subgraph set, at least one path or at least one subgraph is selected from the path set or subgraph set as the evidence chain data packet according to a preset priority rule. The preset priority rule includes at least a priority rule based on the number of nodes and edges of the path or subgraph and a priority rule based on the source information field. Specifically, the priority rule based on the number of nodes and edges is as follows: under the premise of satisfying the graph query conditions, paths or subgraphs with fewer nodes and edges are selected first, so that the node and edge sequence corresponding to the evidence chain data packet is more compact. In this embodiment, the number of nodes and edges of candidate paths or candidate subgraphs is sorted from smallest to largest, and the first three are selected as candidates. The priority rule based on the source information field is as follows: when the number of nodes and edges is the same or the difference does not exceed 2, paths or subgraphs whose source identifiers corresponding to the source information field belong to the judicial or administrative penalty category are selected first; when there are still ties, paths or subgraphs with more source identifier entries are selected first, so that the evidence chain data packet covers more source information fields. This embodiment maps the source identifier of the source information field to a preset level score, which has 5 levels: level 5 for judicial categories, level 4 for administrative penalty categories, level 3 for business registration categories, level 2 for bidding and tendering categories, and level 1 for public opinion information categories. The maximum source identifier level score of each node or edge in the path or subgraph is used as the source priority score for that path or subgraph to determine the priority order among parallel candidates. To avoid over-limitation, this embodiment only uses the above-mentioned level scores for sorting and judgment, without performing further numerical calculations.

[0038] In this embodiment, after determining the path or subgraph corresponding to the evidence chain data packet, the evidence chain data packet is organized into a data structure containing node edge sequences and their source information fields, and a correspondence description is established between the credit conclusion unit and the node edge sequence. Here, the correspondence description refers to a descriptive item used to indicate the association between the conclusion category and target enterprise scope in the credit conclusion unit and the node edge sequence in the evidence chain data packet. Specifically, this embodiment generates a conclusion identifier for each credit conclusion unit and a sequence identifier for each node edge sequence. Both the conclusion identifier and the sequence identifier use incremental integer numbers, with the conclusion identifier and the sequence identifier both starting from at least 1. A corresponding record of conclusion identifier—sequence identifier—starting element identifier—conclusion category—target enterprise scope identifier is written into the correspondence description, and the corresponding record contains at least 5 fields. The source information field is written into the evidence chain data packet along with the node-edge sequence. When the node-edge sequence contains more than 8 nodes or edges, this embodiment splits the node-edge sequence into segmented sequences of no more than 8 nodes or edges, and writes the source information field into each segmented sequence to ensure the stability of the structure field length of the evidence chain data packet. Through the above method, this embodiment can obtain the evidence chain data packet corresponding to the credit conclusion unit and provide directly referenceable node-edge sequences and source information fields for subsequent steps 500 and 600.

[0039] Furthermore, in step 500 of this embodiment, after obtaining the credit conclusion unit and its corresponding candidate fact set in step 400, this embodiment first determines a set of consistency fields for comparison for the same credit conclusion unit. Here, the set of consistency fields refers to the set of fields used to determine whether different candidate facts express the same factual content, and its fields originate from the common fields of each candidate fact in the candidate fact set. This embodiment pre-establishes a correspondence table between conclusion categories and consistency field sets, the correspondence table containing at least 10 corresponding records; each corresponding record contains at least 2 consistency fields. For example, when the conclusion category is a business attribute category, the consistency field set at least includes the enterprise name and unified social credit code; when the conclusion category is a litigation category, the consistency field set at least includes the case number and the party identifier; when the conclusion category is an administrative penalty category, the consistency field set at least includes the penalty decision number and the identifier of the penalized entity; when the conclusion category is an external guarantee category, the consistency field set at least includes the endpoint entity identifier of the guarantee relationship edge and the guarantee amount field. Through the aforementioned correspondence table, this embodiment can determine a set of consistent fields for each credit conclusion unit under the condition of at least one table entry matching, so as to ensure that the source of fields for subsequent grouping determination is clear and reusable.

[0040] In this embodiment, after determining the set of consistent fields, a preset consistency condition is set, and the candidate fact set is judged for consistency accordingly. The preset consistency condition refers to a set of comparison rules for each field in the consistent field set, including at least field equivalence rules and field missing handling rules. The field equivalence rules are used to determine whether the values ​​of two candidate facts are consistent in the consistent field set. In this embodiment, all consistent fields are taken as the default consistency condition. When there are missing fields in the consistent field set, this embodiment uses field missing handling rules, that is, if the consistent field set has at least two fields, a maximum of one field is allowed to be missing without being directly judged as inconsistent, but at least one field must be consistent. To avoid the candidate fact set being too large and affecting grouping stability, this embodiment sets an upper limit on the number of candidate facts participating in the consistency judgment, with the upper limit being 200. When the number of candidate facts exceeds 200, this embodiment selects the first 200 facts according to the generation order of the candidate fact set to participate in the consistency judgment. Based on the aforementioned preset consistency conditions, this embodiment categorizes candidate facts that satisfy the preset consistency conditions into a consistency candidate set and generates a group identifier for each consistency candidate set. The group identifier uses an incrementing integer number and starts from at least 1.

[0041] In this embodiment, candidate facts that do not meet the preset consistency conditions are categorized into a conflict candidate set, and the conflict type of the conflict candidate set is further determined. Here, the conflict candidate set refers to a set of candidate facts that, for the same credit conclusion unit, have differences in the consistency field set or are triggered by preset conflict conditions; the conflict type refers to the category used to characterize the source of the conflict difference. This embodiment uses the field category in the consistency field set where differences occur as the first basis for determining the conflict type: when the difference field belongs to an attribute field, it is preferentially determined to be a value conflict; when the difference field involves a relationship edge type or a relationship edge endpoint entity, it is preferentially determined to be a relationship conflict. Simultaneously, this embodiment uses the time or version information indicated by the source information field, and the source identifier indicated by the source information field, as supplementary bases for determining the conflict type: when a candidate fact meets the preset conflict conditions in time or version information, it is determined to be a temporal conflict; when a candidate fact meets the preset conflict conditions in the source identifier, it is determined to be a source difference conflict. To avoid the same set of conflict candidates being classified into too many categories, this embodiment sets the upper limit of the number of conflict types to 2, and determines the final retained conflict types in the order of relational conflict, value conflict, time sequence conflict, and source difference conflict.

[0042] In this embodiment, specific judgment criteria are provided for value conflicts, relationship conflicts, temporal conflicts, and source difference conflicts. Value conflicts occur when candidate facts corresponding to the same credit assessment conclusion unit have different values ​​in the same attribute field. This embodiment limits the comparison scope of attribute fields to attribute fields in the consistency field set and key attribute fields corresponding to the conclusion category. The number of key attribute fields is not less than one; for example, in the external guarantee category, key attribute fields include the guarantee amount field, and in the litigation category, key attribute fields include the case type field. Relationship conflicts occur when candidate facts corresponding to the same credit assessment conclusion unit differ in relationship edge type or relationship edge endpoint entity. This embodiment uses the relationship type as the comparison benchmark; if the relationship edge types corresponding to the candidate facts are different or the relationship edge endpoint entity identifiers are different, it is determined to be a relationship conflict. The time-series conflict refers to the candidate facts corresponding to the same credit conclusion unit meeting preset conflict conditions in terms of time or version information indicated by the source information field. In this embodiment, the preset conflict conditions are set to trigger a time-series conflict when either the collection time difference exceeds 30 days or the version number difference exceeds two update batches. The 30-day difference and two update batches are example thresholds selected in this embodiment to distinguish between short-term update differences and significant time-series differences. The source difference conflict refers to the candidate facts corresponding to the same credit conclusion unit meeting preset conflict conditions in terms of source identifiers indicated by the source information field. In this embodiment, the preset conflict conditions are set to different source identifiers and a difference in category priority of at least two levels. The category priority of source identifiers can be exemplified by the five-level order shown in Table 1.

[0043] In this embodiment, after determining the conflict type, a difference field description is generated for conflict explanation, and a correspondence is established between the difference field description and the conflict candidate set. Here, the difference field description refers to a aggregated description of the names, difference values, and corresponding candidate fact identifiers of the difference fields in the conflict candidate set. This embodiment generates at least one difference field description for each conflict candidate set, and each difference field description contains at least four items: difference field name, primary difference value, alternative difference values, and a set of candidate fact identifiers; where the primary difference value and alternative difference values ​​are the two values ​​that appear most frequently in the conflict candidate set. When there are more than two difference values ​​in the conflict candidate set, this embodiment only retains the first two values ​​in the difference field description and assigns the remaining values ​​to other value items to ensure a stable description length; this embodiment sets the maximum number of difference field descriptions for each conflict candidate set to 5 to avoid excessive difference information in the dialogue response, which would affect readability. In this embodiment, the candidate fact set is grouped in the above manner to obtain a consistent candidate set and a conflict candidate set. The conflict type and difference field description of the conflict candidate set are determined, providing direct input for step 600 to generate the main conclusion, alternative conclusions and conflict description.

[0044] As an example, this embodiment quantifies the degree of consistency between any two candidate facts and uses whether the degree of consistency meets a preset consistency condition as the basis for grouping. Specifically, assuming that the consistency field set contains several consistency fields, this embodiment compares the values ​​of each consistency field separately and summarizes the comparison results into a degree of consistency; when the degree of consistency meets the preset consistency condition, the corresponding candidate fact is assigned to the consistent candidate set; otherwise, it is assigned to the conflict candidate set.

[0045] Consistent judgment: in, For the first Candidate Facts and the First The degree of consistency of candidate facts across the set of consistency fields; This represents the number of fields in the consistency field set. For the first The candidate facts in the first The values ​​of each consistent field; A function to indicate the comparison of field values; This indicates that the field is missing or has a null value; The consistency threshold corresponding to the preset consistency condition is determined by the following values: The determination is made based on whether or not missing fields are allowed.

[0046] In this embodiment, the time difference and version difference between candidate facts under the same credit conclusion unit are compared; when the time difference exceeds a preset number of days threshold or the version difference exceeds a preset batch threshold, it is determined that the timing conflict condition is met.

[0047] Timing conflict determination: in, For the first Candidate Facts and the First The time difference in the collection of candidate facts; For the first The time of collection of each candidate fact; In this embodiment, the time difference threshold is set to 30 days. For the first Candidate Facts and the First Version differences of candidate facts; For the first The update batch number corresponding to the version number of each candidate fact; To determine the version difference threshold, this embodiment uses two update batches. This indicates a logical OR.

[0048] In this embodiment, the source identifier is mapped to the source level, and the difference between the source levels is used for determination; when the source levels are different and the difference between the source levels reaches a preset level difference threshold, it is determined that the source difference conflict condition is met.

[0049] Determination of source discrepancies: .

[0050] in, Source identification Corresponding source level; For the first Source identifiers for each candidate fact; For the first The source level of each candidate fact; As the threshold for source grade difference, this embodiment uses 2 levels; This represents absolute value operations.

[0051] In this embodiment, to ensure that the primary difference value and the alternative difference value in the difference field description are the two values ​​that appear most frequently in the conflict candidate set, this embodiment counts each value of the same difference field in the conflict candidate set, selects the value with the largest count as the primary difference value, and selects the value with the second largest count as the alternative difference value; when there are ties in the count, this embodiment prioritizes the value supported by the candidate fact corresponding to the value with the higher source level, and if they are still tied, the value with the later collection time is given priority.

[0052] in, This refers to the field index of the difference field within the set of consistent fields. The number of candidate facts in the conflict candidate set; For the first The candidate facts in the first The values ​​that can be taken from the difference field; For the first in the conflict candidate set The set of possible values ​​for each difference field; This is an indicator function that takes the value 1 when the condition inside the parentheses is true, and 0 otherwise. For taking values In the The number of occurrences in each difference field; For the first The main difference values ​​for each difference field; For the first Each difference field has possible difference values; This represents the value selection operator that maximizes the count.

[0053] Furthermore, in step 600 of this embodiment, after obtaining the conflict candidate set and difference field description corresponding to the same credit conclusion unit in step 500, this embodiment determines the source authority level, collection time freshness level, fact consistency level, and historical hit rate level for each candidate fact in the conflict candidate set. The source authority level is determined by the source identifier in the source information field of the candidate fact, and can directly use the 5-level order shown in Table 1 as the level value, with level values ​​from 1 to 5. The collection time freshness level is determined by the interval between the collection time of the candidate fact and the latest collection time under the same credit conclusion unit. In this embodiment, an interval of no more than 7 days corresponds to level 5, no more than 30 days corresponds to level 4, no more than 90 days corresponds to level 3, no more than 180 days corresponds to level 2, and more than 180 days corresponds to level 1. The fact consistency level is used to characterize the degree of fit between the candidate fact and the consistent candidate set or mainstream value. In this embodiment, candidate facts supporting the same value are used as the basis for determining the level. The actual quantity serves as the consistency criterion. A quantity of at least 3 corresponds to Level 5, a quantity of 2 corresponds to Level 4, a quantity of 1 corresponds to Level 3, if the location can only be described by the difference field and the quantity cannot form an aggregation, it corresponds to Level 2, and if the missing key field makes comparison impossible, it corresponds to Level 1. The historical hit rate level is determined by the proportion of valid records with the same source identifier within a preset statistical window. The statistical window takes the most recent 200 candidate facts from the same source as an example. A hit rate of at least 90% corresponds to Level 5, at least 75% corresponds to Level 4, at least 60% corresponds to Level 3, at least 40% corresponds to Level 2, and below 40% corresponds to Level 1.

[0054] In this embodiment, after obtaining the above four levels, each level is mapped to a level score, and the level scores are combined according to a preset scoring rule to obtain a confidence score. The level score can be directly taken as the corresponding level value, i.e., level 1 corresponds to 1 point, level 2 to 2 points, up to level 5 to 5 points, to reduce unnecessary parameter items. The preset scoring rule uses a summation method to combine the four level scores, thereby making the confidence score range from 4 to 20, facilitating the ranking and comparison of different candidate facts. To avoid score anomalies due to the absence of a certain level, this embodiment uses a default level for missing levels, with the default level being level 2; for example, when the historical hit rate level cannot be stably obtained due to insufficient statistical windows (less than 200 records), the historical hit rate level is included in the combination as level 2. Through the above mapping and combination, this embodiment can form comparable confidence scores for each candidate fact in the conflict candidate set, while maintaining consistency between the meaning of the score and the meaning of the level.

[0055] In this embodiment, after obtaining the confidence scores of each candidate fact, the conclusion corresponding to the candidate fact with the highest confidence score is determined as the main conclusion, and the conclusions corresponding to candidate facts that meet the preset scoring conditions (excluding the main conclusion) are determined as alternative conclusions. The preset scoring conditions are used to control the quantity and quality of alternative conclusions. In this embodiment, a confidence score of not less than 12 points and a score difference of no more than 2 points from the main conclusion are used as example conditions, and the maximum number of alternative conclusions is set to 3. When multiple candidate facts have the same score and compete for the main conclusion, this embodiment prioritizes selecting the one with a higher source authority level as the main conclusion; if they are still the same, it prioritizes selecting the one with a higher collection time freshness level as the main conclusion. For cases where alternative conclusions have duplicate values, this embodiment only retains alternative conclusions with values ​​different from the main conclusion to ensure that the alternative conclusions reflect the differences in the conflicting candidate set rather than being repeatedly output.

[0056] In this embodiment, after determining the main conclusion and alternative conclusions, a conflict explanation is generated based on the description of the difference fields, the candidate facts corresponding to the main conclusion, and the candidate facts corresponding to the alternative conclusions. The conflict explanation is a description of the difference information for the credit conclusion unit, including at least the difference field name, the value of the main conclusion, the value of the alternative conclusion, and their corresponding source identifiers and collection time information, so that the difference information corresponds one-to-one with the candidate facts. When the difference field description contains multiple difference fields, this embodiment selects up to five difference fields to write into the conflict explanation in the following order: priority is given to differences in relationship edge type or relationship edge endpoint entity, followed by differences in key attribute fields, and then differences in other attribute fields. For example, when the difference field is the guarantee amount field, the conflict explanation can include two parallel descriptions: a guarantee amount of 3 million (source identifier 4, collection time is a certain day) and a guarantee amount of 2 million (source identifier 3, collection time is a certain day); when the difference field is the case type field, the conflict explanation can list values ​​such as enforcement civil cases and their source identifiers respectively. In this way, this embodiment ensures that the conflict description and the difference field description obtained in step 500 are in the same field system, and that the correspondence with the main conclusion and alternative conclusions is clear.

[0057] In this embodiment, the main conclusion, alternative conclusions, confidence scores, evidence chain data packets, and conflict explanations are output to the dialogue engine to generate a dialogue response. Specifically, this embodiment reserves a conclusion display position for each credit conclusion unit in the dialogue response and writes the main conclusion and its confidence score into the conclusion display position; when there is a conflict candidate set, the alternative conclusions and conflict explanations are written into the position corresponding to the same credit conclusion unit immediately after the main conclusion. At the same time, this embodiment configures a reference identifier for the evidence chain data packet for the credit conclusion unit. The reference identifier is generated using an incremental integer numbering method, starting from at least 1; and in the dialogue response, the reference identifier indicates the node edge sequence in the evidence chain data packets corresponding to the main conclusion and alternative conclusions, respectively. Further, this embodiment associates the conflict explanation with the reference identifier in the output, so that the description of the difference fields involved in the conflict explanation can correspond to the source information field in the evidence chain data packet; wherein, when the same credit conclusion unit corresponds to multiple evidence chain data packets, this embodiment sets the upper limit of the number of reference identifiers to 5, and prioritizes retaining the node edge sequence directly corresponding to the main conclusion and alternative conclusions to ensure that the dialogue response structure is compact and the reference relationship is clear.

[0058] Figure 4 This illustration demonstrates the structured output format of the present invention when generating dialogue responses. The dialogue response includes a dialogue engine interface, user query content, and a response content area corresponding to the credit conclusion unit. The response content area includes at least the credit conclusion unit identifier, the main conclusion and its confidence score, alternative conclusions and their confidence scores, conflict explanations, and corresponding citation identifiers. The citation identifiers correspond to the citation explanation area below, which lists the node edge sequences and their source information fields in the evidence chain data packet. By establishing a citation relationship between the credit conclusion unit and the evidence chain data packet, both the main conclusion and alternative conclusions have locatable structured supporting information.

[0059] As an example, in this embodiment, after mapping the source authority level, collection time freshness level, factual consistency level, and historical hit rate level to level scores, the confidence score of the candidate fact is obtained by summing: in, For the first Confidence score of each candidate fact; For the first The score corresponding to the authority level of the source of each candidate fact; For the first The score corresponding to the freshness level of the collection time of each candidate fact; For the first The score corresponding to the factual consistency level of each candidate fact; For the first The score corresponding to the historical hit rate level of each candidate fact; All values ​​are integers and range from 1 to 5.

[0060] In this embodiment, the main conclusion and alternative conclusions are determined based on the confidence scores, wherein the main conclusion corresponds to the candidate fact with the highest confidence score, and the alternative conclusions correspond to other candidate facts that meet preset score conditions. in, A set of candidate fact indexes for conflict candidate sets corresponding to the same credit conclusion unit; Index of candidate facts corresponding to the main conclusion; This is the set of candidate fact indices corresponding to the alternative conclusions; As the minimum score threshold for the candidate conclusions, this embodiment takes... The primary / secondary difference threshold is taken in this embodiment. The confidence score corresponding to the main conclusion.

[0061] In this embodiment, when multiple candidate facts satisfy... When there is a tie for the highest value, the unique primary conclusion is determined by prioritizing the source authority level, followed by the collection time freshness level. The corresponding selection expression is: in, This represents ordered triples compared lexicographically, i.e., compared first... If they are the same, then compare them. If they are still the same, then compare again. Thus determining the unique The meanings of the remaining symbols are the same as those described above.

[0062] Figure 5 This is a schematic diagram of the system structure provided in the embodiments of the present invention, such as... Figure 5 As shown, the present invention also provides an intelligent enterprise credit analysis system that integrates knowledge graphs and dialogue engines, comprising: The multi-source credit information collection unit is used to acquire multi-source credit information records of the target enterprise and associate source information fields with each multi-source credit information record; The knowledge graph construction unit is used to extract entities and relationships based on multi-source credit records and construct an enterprise credit knowledge graph, writing the source information fields into the attributes of the corresponding nodes or edges; The query parsing unit is used to receive credit query requests through the dialogue engine, parse the credit query requests to obtain credit conclusion units and graph query conditions; The evidence chain extraction unit is used to retrieve a set of candidate facts in the enterprise credit knowledge graph based on graph query conditions, and extract the path or subgraph corresponding to the credit conclusion unit from the set of candidate facts to form an evidence chain data package; the evidence chain data package includes at least the node edge sequence and its source information field; The conflict grouping unit is used to group the candidate fact set according to the preset conflict rules to obtain the consistent candidate set and the conflict candidate set, and to determine the conflict type of the conflict candidate set; The conclusion generation and output unit is used to calculate the confidence score of the conflict candidate set based on the preset scoring rules, determine the main conclusion and the alternative conclusions and generate conflict descriptions, and output the main conclusion, alternative conclusions, confidence scores, evidence chain data packets and conflict descriptions to the dialogue engine to generate a dialogue response; the dialogue response displays the confidence score and cites the evidence chain data packets to the credit conclusion unit.

[0063] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.

[0064] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for intelligent enterprise credit analysis that integrates knowledge graphs and dialogue engines, characterized in that, include: Obtain the multi-source credit records of the target enterprise, and associate source information fields with each of the multi-source credit records; Based on the multi-source credit records, entities and relationships are extracted and an enterprise credit knowledge graph is constructed. The source information fields are written into the attributes of the corresponding nodes or edges. The system receives credit query requests through a dialogue engine and parses the credit query requests to obtain credit conclusion units and graph query conditions. Based on the graph query conditions, a set of candidate facts is retrieved from the enterprise credit knowledge graph. From the set of candidate facts, a path or subgraph corresponding to the credit conclusion unit is extracted to form an evidence chain data packet. The evidence chain data packet includes at least a node edge sequence and its source information field. The candidate fact set is grouped according to a preset conflict rule to obtain a consistent candidate set and a conflict candidate set, and the conflict type of the conflict candidate set is determined. The confidence score is calculated for the conflict candidate set based on the preset scoring rules, the main conclusion and the alternative conclusions are determined and the conflict description is generated, and the main conclusion, the alternative conclusions, the confidence score, the evidence chain data package and the conflict description are output to the dialogue engine to generate a dialogue response; The dialogue response displays the confidence score to the credit conclusion unit and references the evidence chain data package.

2. The method according to claim 1, characterized in that, For each of the aforementioned multi-source credit records, the associated source information field includes: For each of the aforementioned multi-source credit records, a source information field is generated, wherein the source information field includes at least a source identifier; Establish a one-to-one correspondence between the source information field and the corresponding multi-source credit records, and store them together; When the multi-source credit records are used to construct the enterprise credit knowledge graph, the source information field is called as the data source for subsequent writing of node or edge attributes.

3. The method according to claim 1, characterized in that, Based on the aforementioned multi-source credit records, entities and relationships are extracted and a corporate credit knowledge graph is constructed, including: Identify enterprise entities, natural person entities, and event entities from the multi-source credit records, and generate entity nodes; Identify at least one of the following relationships among the entity nodes from the multi-source credit records: equity relationship, employment relationship, guarantee relationship, litigation relationship, and penalty relationship, and generate relationship edges; The enterprise credit knowledge graph is formed based on the entity nodes and the relationship edges, and the type identifiers of the entity nodes and the relationship edges are retained in the enterprise credit knowledge graph.

4. The method according to claim 1, characterized in that, The credit query request is parsed to obtain the credit conclusion unit and graph query conditions, including: The target enterprise indication information is determined from the credit inquiry request, and the target enterprise range corresponding to the credit conclusion unit is determined accordingly. Determine the conclusion category indication information from the credit inquiry request, and determine the conclusion category of the credit conclusion unit accordingly; Based on the target enterprise scope and the conclusion category, the graph query conditions are generated; the graph query conditions include at least one or more of the following: node type restriction, relationship type restriction, and attribute filtering conditions.

5. The method according to claim 1, characterized in that, Extracting the path or subgraph corresponding to the credit conclusion unit from the candidate fact set to form an evidence chain data packet includes: The starting element is determined based on the credit conclusion unit, and the starting element is the target node or the target relationship edge; Centered on the starting element, the path set or subgraph set is obtained in the enterprise credit knowledge graph according to the graph query conditions to form the candidate fact set; At least one path or at least one subgraph is selected from the path set or the subgraph set as the evidence chain data packet according to a preset priority rule; wherein the preset priority rule includes at least a priority rule for the number of nodes and edges of the path or subgraph and a priority rule for the source information field; The evidence chain data packet is organized into a data structure containing node edge sequences and their source information fields, and a correspondence description between the credit conclusion unit and the node edge sequences is established.

6. The method according to claim 1, characterized in that, The candidate fact set is grouped according to a preset conflict rule to obtain a consistent candidate set and a conflict candidate set, and the conflict type of the conflict candidate set is determined, including: For the same credit conclusion unit, determine the set of consistency fields used for comparison in the candidate fact set; Candidate facts whose consistency field set in the candidate fact set satisfies the preset consistency condition are included in the consistency candidate set; Candidate facts whose consistency field set in the candidate fact set does not meet the preset consistency condition are included in the conflict candidate set. Based on the field categories where differences occur in the consistency field set, the conflict type of the conflict candidate set is determined, and a difference field description for the conflict description is generated.

7. The method according to claim 6, characterized in that, The conflict types include at least value conflicts, relational conflicts, temporal conflicts, and source difference conflicts, wherein: The value conflict refers to the different values ​​of the candidate facts corresponding to the same credit conclusion unit on the same attribute field; The relationship conflict occurs when the candidate facts corresponding to the same credit conclusion unit are different in terms of relationship edge type or relationship edge endpoint entity. The time sequence conflict is when the candidate facts corresponding to the same credit conclusion unit meet the preset conflict conditions in the time or version information indicated by the source information field; The source difference conflict is when the candidate facts corresponding to the same credit conclusion unit meet the preset conflict conditions on the source identifier indicated by the source information field.

8. The method according to claim 6, characterized in that, Calculate confidence scores for the conflict candidate set based on preset scoring rules, determine the main conclusion and alternative conclusions, and generate conflict explanations, including: For each candidate fact in the conflict candidate set, determine the source authority level, the collection time freshness level, the fact consistency level, and the historical hit rate level; Each level is mapped to a level score, and the level scores are combined according to the preset scoring rules to obtain the confidence score; The conclusion corresponding to the candidate fact with the highest confidence score is determined as the main conclusion, and the conclusions corresponding to the candidate facts that meet the preset score conditions other than the main conclusion are determined as the alternative conclusions. The conflict description is generated based on the difference field description, the candidate facts corresponding to the main conclusion, and the candidate facts corresponding to the alternative conclusions.

9. The method according to claim 1, characterized in that, The main conclusion, the alternative conclusions, the confidence score, the evidence chain data package, and the conflict explanation are output to the dialogue engine to generate a dialogue response, including: Write the main conclusion and the confidence score corresponding to the main conclusion into the position corresponding to the credit conclusion unit in the dialogue response; When the conflict candidate set exists, the alternative conclusions and the conflict descriptions are written into the dialogue response at the position corresponding to the credit conclusion unit; In the dialogue response, the reference identifier of the evidence chain data packet is configured for the credit conclusion unit, so that the main conclusion and the alternative conclusion correspond to the node edge sequence in the evidence chain data packet, respectively. The conflict description is associated with the reference identifier of the evidence chain data packet and output, so that the difference field description involved in the conflict description corresponds to the source information field in the evidence chain data packet.

10. A corporate credit intelligent analysis system integrating knowledge graphs and dialogue engines, characterized in that, include: The multi-source credit information collection unit is used to acquire multi-source credit information records of the target enterprise and associate source information fields with each of the multi-source credit information records; The knowledge graph construction unit is used to extract entities and relationships based on the multi-source credit records and construct an enterprise credit knowledge graph, and write the source information fields into the attributes of the corresponding nodes or edges; The query parsing unit is used to receive credit query requests through the dialogue engine and parse the credit query requests to obtain credit conclusion units and graph query conditions. The evidence chain extraction unit is used to retrieve a set of candidate facts in the enterprise credit knowledge graph according to the graph query conditions, and extract the path or subgraph corresponding to the credit conclusion unit from the set of candidate facts to form an evidence chain data packet; the evidence chain data packet includes at least a node edge sequence and its source information field; The conflict grouping unit is used to group the candidate fact set according to a preset conflict rule to obtain a consistent candidate set and a conflict candidate set, and to determine the conflict type of the conflict candidate set. The conclusion generation and output unit is used to calculate the confidence score of the conflict candidate set based on the preset scoring rules, determine the main conclusion and the alternative conclusions and generate a conflict description, and output the main conclusion, the alternative conclusions, the confidence score, the evidence chain data package and the conflict description to the dialogue engine to generate a dialogue response. The dialogue response displays the confidence score to the credit conclusion unit and references the evidence chain data package.