Knowledge base data fusion and association method and system based on large ai model
Patent Information
- Application Number
- PCT/CN2026/107868
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-09-08
- Filing Date
- 2026-07-01
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026107868_01102026_PF_FP_ABST
Abstract
Description
A knowledge base data fusion and association method and system based on AI large model Technical Field
[0001] This application relates to the field of knowledge base fusion technology, specifically to a method and system for knowledge base data fusion and association based on an AI large model. Background Technology
[0002] In today's era, information is exploding, with massive amounts of data flooding our lives and work. The fragmentation, diversity, and rapid updating of information pose enormous challenges to people's acquisition and utilization of knowledge. Knowledge base data fusion has emerged to address this challenge, effectively eliminating data redundancy and inconsistency, filling data gaps, improving data accuracy, and thus enhancing data quality.
[0003] However, during data fusion, knowledge conflicts can lead to decreased data reliability, impaired accuracy, and reduced usability. They can also increase maintenance costs, worsen user experience, raise legal compliance risks, and even disrupt business processes, increasing the complexity of technical implementation. More challenging is that these conflicts may be hidden after data fusion, making them difficult to detect immediately and only gradually emerge during system operation. Currently, there is a lack of intelligent conflict detection, resolution, and optimization mechanisms, making it difficult to specifically identify the root causes and address complex fusion problems.
[0004] In existing technologies, knowledge base fusion methods for encyclopedia websites integrate the knowledge cards (infoboxes) of the most influential encyclopedias: Baidu Baike, Hudong Baike, and Chinese Wikipedia. These methods typically include the following steps: Step 1, obtaining and preprocessing query results for the same entity from the encyclopedia website; Step 2, establishing mapping relationships between entities in the encyclopedia website based on features of concept similarity, attribute similarity, and contextual similarity; Step 3, aligning the attributes of the knowledge cards of entities with established mapping relationships using an external dictionary; Step 4, designing single-truth-finding and multi-truth-finding schemes for attributes with conflicting values, depending on whether the attribute values are single-valued or multi-valued; Step 5, outputting the fused attribute-value pairs. The resulting highly reliable, redundant attribute-value pairs for the three major encyclopedias, while reducing redundancy and conflicts to some extent, still suffer from relatively simplistic conflict handling. It fails to provide targeted solutions or promptly identify the causes of conflicts and address the root causes, thus failing to truly resolve anomalies and reduce the probability of conflicts. Optimizing database fusion solutions, promptly identifying the main contradictions leading to data conflicts, and addressing the hidden dangers behind the data have become key research issues. Summary of the Invention
[0005] This application provides a knowledge base data fusion and association method and system based on AI large model, which overcomes the problem in existing database fusion methods that, when faced with data conflicts, it is difficult to specifically identify the causes of data conflicts and provide targeted solutions to the causes of data conflicts.
[0006] Firstly, this application provides a knowledge base data fusion and association method based on an AI large-scale model, including:
[0007] Data acquisition: Raw data is collected via connectors.
[0008] Data preprocessing involves cleaning the original data and removing duplicates from the cleaned data based on the Jaccard similarity threshold.
[0009] Field alignment involves obtaining the names of each field and generating corresponding vectorized semantics, and then aligning the vectorized semantics based on the field name similarity threshold.
[0010] Entity recognition and extraction: AI large model is used to extract entities and attributes from the cleaned data and record them as extracted entities and extracted attributes.
[0011] Entity disambiguation and linking: Generate corresponding vectorized semantics for each extracted entity; merge extracted entities based on density clustering algorithm; retrieve vectorized entities from the database based on entity similarity threshold; and record vectorized entities whose similarity with the extracted entity is higher than the entity similarity threshold as candidate entities. If there are several candidate entities, a link is constructed between the extracted entity with the highest similarity and the vectorized entity, and the corresponding extracted entity is recorded as a successfully linked entity. If there are no candidate entities, a new entity is constructed.
[0012] Relation extraction and triple construction: generating triples based on a large AI model;
[0013] Standardization and normalization assign unified identifiers to entities and attributes to obtain merged data.
[0014] Conflict detection and optimization: Based on the coefficient of variation of the original data and the fused data, determine whether the current data conflict meets the requirements; if the requirements are not met, determine the reason for the non-compliance based on the proportion of duplicate hash values in the total hash value of the fused data; and, based on the corresponding reason, correct the data preprocessing parameters, field alignment parameters, and entity disambiguation and linking parameters.
[0015] Vector index construction and output: The vector index is constructed using a graph database, and a graph download interface is provided.
[0016] Furthermore, the process of determining whether the current data conflict meets the requirements based on the coefficient of variation of the original data and the fused data includes:
[0017] The original data and the fused data are obtained and quantized, and the original data and the fused data are aligned.
[0018] The absolute difference value of each field is calculated one by one to generate a difference value sequence, and the coefficient of variation of the original data and the fused data is determined based on the difference sequence;
[0019] If the coefficient of variation is less than or equal to the preset coefficient of variation, it is determined that the current data conflict meets the requirements, and the parameters of the knowledge base data fusion and association method remain unchanged.
[0020] If the coefficient of variation is greater than the preset coefficient of variation, it is determined that the current data conflict does not meet the requirements, and the reason for not meeting the requirements is determined based on the proportion of the duplicate hash value of the merged data in the total hash value.
[0021] Furthermore, the process of determining the reason for dissatisfaction based on the proportion of duplicate hash values in the total hash value after fusion includes:
[0022] The fused data is then converted into hash values.
[0023] The hash collision rate is determined based on the proportion of duplicate hash values in the total number of hash values in the merged data.
[0024] If the hash collision rate is less than or equal to the preset hash collision rate, the reason for not meeting the requirement is determined based on the proportion of successfully linked entities in the extracted entities.
[0025] If the hash collision rate is greater than the preset hash collision rate, it is determined that the Jaccard similarity threshold setting is unreasonable, and the Jaccard similarity threshold is corrected based on the ratio of the hash collision rate to the preset hash collision rate.
[0026] The Jaccard similarity threshold refers to the minimum Jaccard similarity required between any two data attribute values when they are considered to be duplicated.
[0027] Furthermore, the process of adjusting the Jaccard similarity threshold based on the ratio of the hash collision rate to the preset hash collision rate includes:
[0028] The hash collision rate ratio is determined based on the ratio of the hash collision rate to the preset hash collision rate;
[0029] The Jaccard similarity threshold is reduced based on the hash collision rate ratio, and the reduction in the Jaccard similarity threshold is proportional to the hash collision rate ratio.
[0030] Furthermore, the process of determining the reason for dissatisfaction based on the proportion of successfully linked entities in the extracted entities includes:
[0031] The success rate of linking entities is determined based on the proportion of successfully linked entities in the extracted entities.
[0032] If the success rate of extracting entity links is less than or equal to the preset success rate of extracting entity links, it is determined that the entity similarity threshold setting is unreasonable, and the entity similarity threshold is corrected based on the ratio of the preset success rate of extracting entity links to the success rate of extracting entity links.
[0033] If the success rate of extracting entity links is greater than the preset success rate of extracting entity links, the reason for non-compliance is determined based on the proportion of the number of field name similarities distributed in the field name similarity range to the total number of field name similarities.
[0034] Wherein, the entity similarity threshold refers to the similarity that the extracted entity and the vectorized entity need to achieve when the vectorized entity is recorded as a candidate entity, and the field name similarity range refers to the similarity range that is evenly distributed on both sides of the field name similarity threshold and includes the field name similarity threshold.
[0035] Furthermore, the process of adjusting the entity similarity threshold based on the ratio of the preset entity linking success rate to the entity linking success rate includes:
[0036] The entity link success rate ratio is determined based on the ratio of the preset entity link extraction success rate to the entity link extraction success rate.
[0037] The entity similarity threshold is reduced based on the entity link success rate ratio, and the reduction in the entity similarity threshold is proportional to the entity link success rate ratio.
[0038] Furthermore, the process of determining the reason for non-compliance based on the proportion of the number of field name similarities distributed within the range of field name similarities to the total number of field name similarities includes:
[0039] The distribution ratio of field name similarity is determined based on the proportion of the number of field name similarities distributed within the range of field name similarities in the total number of field name similarities;
[0040] If the proportion of field name similarity distribution is less than or equal to the preset proportion of field name similarity distribution, it is determined that the density clustering algorithm parameter setting is unreasonable, and the density clustering algorithm neighborhood radius is corrected based on the ratio of the coefficient of variation to the preset coefficient of variation.
[0041] If the proportion of field name similarity distribution is greater than the proportion of preset field name similarity distribution, it is determined that the field name similarity threshold setting is unreasonable, and the field name similarity threshold is corrected based on the difference between the proportion of field name similarity distribution and the proportion of preset field name similarity distribution.
[0042] The field name similarity threshold refers to the minimum similarity required for the corresponding vectorized semantics when aligning the field names.
[0043] Furthermore, the process of correcting the neighborhood radius of the density clustering algorithm based on the ratio of the coefficient of variation to the preset coefficient of variation includes:
[0044] The coefficient of variation ratio is determined based on the ratio of the coefficient of variation to the preset coefficient of variation;
[0045] The radius of the density clustering algorithm is increased based on the coefficient of variation ratio, and the increase in the radius of the density clustering algorithm is proportional to the coefficient of variation ratio.
[0046] Furthermore, the process of adjusting the field name similarity threshold based on the difference between the field name similarity distribution ratio and the preset field name similarity distribution ratio includes:
[0047] The distribution ratio difference is determined based on the difference between the distribution ratio of the field name similarity and the preset distribution ratio of the field name similarity.
[0048] The field name similarity threshold is reduced based on the difference in distribution proportions, and the reduction in the field name similarity threshold is proportional to the difference in distribution proportions.
[0049] Secondly, this application provides a knowledge base data fusion and association system based on an AI large-scale model, including:
[0050] The acquisition module includes several connectors for acquiring the raw data;
[0051] A data preprocessing module, which is connected to the acquisition module, is used to clean the data and remove duplicates from the cleaned data based on the Jaccard similarity threshold.
[0052] The field alignment module, which is connected to the data preprocessing module, is used to obtain the field names and generate the corresponding vectorized semantics, and align the vectorized semantics based on the field name similarity threshold.
[0053] The entity recognition and extraction module, which is connected to the field alignment module, is used to extract entities and attributes from the cleaned data based on the AI large model, and denoted as the extracted entity and extracted attribute.
[0054] The entity disambiguation and linking module, which is connected to the entity recognition and extraction module, is used to generate vectorized semantics corresponding to each extracted entity, and to merge extracted entities based on density clustering algorithm, and to retrieve vectorized entities in the database based on entity similarity threshold. Vectorized entities whose similarity between extracted entities and vectorized entities is higher than the entity similarity threshold are recorded as candidate entities. If there are several candidate entities, a link is constructed between the extracted entity with the highest similarity and the vectorized entity, and the corresponding extracted entity is recorded as a successfully linked entity. If there are no candidate entities, a new entity is constructed.
[0055] The triplet construction module is connected to the entity disambiguation and linking module to generate triplets based on the AI large model;
[0056] The normalization module, which is connected to the triplet construction module, is used to assign unified identifiers to entities and attributes to obtain fused data;
[0057] The conflict detection and optimization module, which is connected to the normalization module, is used to determine whether the current data conflict meets the requirements based on the coefficient of variation of the original data and the fused data; when it is determined that the requirements are not met, the reason for the non-compliance is determined based on the proportion of duplicate hash values in the total hash value of the fused data; and, based on the corresponding reason, the data preprocessing parameters, field alignment parameters, and entity disambiguation and linking parameters are corrected.
[0058] The control module, which is connected to the conflict detection and optimization module, the data preprocessing module, the field alignment module, and the entity disambiguation and linking module, is used to schedule the corresponding modules to execute according to the instructions of the conflict detection and optimization module.
[0059] The output module, which is connected to the normalization module and the conflict detection and optimization module, is used to construct a vector index based on the graph database and provide a graph download interface.
[0060] Compared with existing technologies, the advantages of this application are as follows: It collects raw data through connectors, performs cleaning and deduplication to improve data quality, and then vectorizes and aligns field names to resolve inconsistencies. Subsequently, it uses an AI large-scale model to extract entities and attributes, and then performs entity disambiguation and linking using density clustering algorithms and entity similarity thresholds. Relationship triples are generated based on the AI large-scale model, and unified identifiers are assigned to entities and attributes to obtain the fused data. In the conflict detection and optimization stage, this application uses the coefficient of variation to evaluate fusion stability and accurately locates duplicate data problems through hash collision rate. Based on this, it dynamically adjusts data preprocessing parameters, field alignment parameters, and entity disambiguation and linking parameters to specifically correct data conflicts during the data fusion process. This effectively solves the problem of difficulty in identifying conflict causes and providing solutions in existing methods, significantly improving fusion quality and reliability. Finally, this solution uses a graph database to build a vector index and provides a graph download interface for convenient user querying and application. Through multi-dimensional quantitative analysis, this application accurately locates the root cause of data conflicts and dynamically adjusts relevant parameters to address the conflict problem in a targeted manner. It overcomes the shortcomings of existing methods, significantly improves the quality and reliability of knowledge base data fusion, and has important innovative significance and practical value.
[0061] Furthermore, this application provides a knowledge base data fusion and association system based on an AI large-scale model. Through the collaborative work of modules including data acquisition, data preprocessing, field alignment, entity recognition and extraction, entity disambiguation and linking, triple construction, normalization, conflict detection and optimization, control, and output, a complete knowledge base data fusion and association process is constructed. These modules are closely interconnected and cooperate with each other, efficiently completing a series of operations such as data acquisition, processing, fusion, conflict detection, and optimization. In particular, the conflict detection and optimization module can dynamically adjust the fusion strategy based on multiple parameters to ensure the quality and consistency of the fused data. In addition, the system uses a graph database to construct a vector index and provides a graph download interface, facilitating users to query and apply the fused knowledge base. This systematic solution not only improves the efficiency and quality of knowledge base data fusion. Attached Figure Description
[0062] Figure 1 is a block diagram of the knowledge base data fusion and association system based on AI large model in an embodiment of this application;
[0063] Figure 2 is a flowchart of the knowledge base data fusion and association method based on AI large model in the embodiments of this application;
[0064] Figure 3 is a flowchart of the process for determining the reasons for non-compliance based on hash collision rate in an embodiment of this application;
[0065] Figure 4 is a flowchart of the process of correcting the Jaccard similarity threshold based on the hash collision rate ratio in an embodiment of this application. Detailed Implementation
[0066] To make the objectives and advantages of this application clearer, the application will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining this application and are not intended to limit this application.
[0067] Preferred embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of this application and are not intended to limit the scope of protection of this application.
[0068] Please refer to Figure 1, which is a block diagram of the knowledge base data fusion and association system based on an AI large model in this embodiment of the application. The knowledge base data fusion and association system based on an AI large model described in this application includes: a data acquisition module, a data preprocessing module, a field alignment module, an entity recognition and extraction module, an entity disambiguation and linking module, a triplet construction module, a normalization module, a conflict detection and optimization module, a control module, and an output module.
[0069] The acquisition module includes several connectors for acquiring the raw data;
[0070] A data preprocessing module, which is connected to the acquisition module, is used to clean the data and remove duplicates from the cleaned data based on the Jaccard similarity threshold.
[0071] The field alignment module, which is connected to the data preprocessing module, is used to obtain the field names and generate the corresponding vectorized semantics, and align the vectorized semantics based on the field name similarity threshold.
[0072] The entity recognition and extraction module, which is connected to the field alignment module, is used to extract entities and attributes from the cleaned data based on the AI large model, and denoted as the extracted entity and extracted attribute.
[0073] The entity disambiguation and linking module, which is connected to the entity recognition and extraction module, is used to generate vectorized semantics corresponding to each extracted entity, and to merge extracted entities based on density clustering algorithm, and to retrieve vectorized entities in the database based on entity similarity threshold. Vectorized entities whose similarity between extracted entities and vectorized entities is higher than the entity similarity threshold are recorded as candidate entities. If there are several candidate entities, a link is constructed between the extracted entity with the highest similarity and the vectorized entity, and the corresponding extracted entity is recorded as a successfully linked entity. If there are no candidate entities, a new entity is constructed.
[0074] The triplet construction module is connected to the entity disambiguation and linking module to generate triplets based on the AI large model;
[0075] The normalization module, which is connected to the triplet construction module, is used to assign unified identifiers to entities and attributes to obtain fused data;
[0076] The conflict detection and optimization module, which is connected to the normalization module, is used to determine whether the current data conflict meets the requirements based on the coefficient of variation of the original data and the fused data, and, when it is determined that the requirements are not met, to determine the reason for the non-compliance based on the proportion of the duplicate hash value of the fused data in the total hash value, and to correct the data preprocessing parameters, field alignment parameters, entity disambiguation and linking parameters based on the corresponding reasons.
[0077] The control module, which is connected to the conflict detection and optimization module, the data preprocessing module, the field alignment module, and the entity disambiguation and linking module, is used to schedule the corresponding modules to execute according to the instructions of the conflict detection and optimization module.
[0078] The output module, which is connected to the normalization module and the conflict detection and optimization module, is used to construct a vector index based on the graph database and provide a graph download interface.
[0079] In principle, there are no restrictions on the methods of cleaning the raw data during the data preprocessing stage. Technical personnel can choose different data cleaning methods according to the characteristics of the data. For example, unifying the character encoding to UTF-8, removing non-printable characters, replacing special whitespace characters (multiple spaces, tabs) with single spaces, unifying the decimal point representation and date representation, etc., will not be elaborated here.
[0080] The data preprocessing stage, which involves deduplicating data attribute values based on the Jaccard similarity threshold, includes:
[0081] The AI large model is used to extract entities from the cleaned data, and each entity is treated as a set. Then, the intersection and union of each pair of entities are calculated.
[0082] The ratio of the size of the intersection (the number of elements in the set) to the size of the union is denoted as the Jaccard similarity.
[0083] If the Jaccard similarity is greater than the set Jaccard similarity threshold, the two sets are considered to be duplicates, and either set is deleted; otherwise, the calculation continues.
[0084] The Jaccard similarity threshold refers to the minimum similarity threshold that must be reached between any two data attribute values when they are found to be duplicated. In the fields of financial data fusion or medical data fusion, the Jaccard similarity threshold is usually set between 0.7 and 0.9. Therefore, in this solution, the Jaccard similarity threshold is set to a range of 0.7-0.9 to improve the accuracy, consistency and reliability of data fusion and reduce data conflicts.
[0085] In the field alignment stage and the entity disambiguation and linking stage, the method of generating corresponding vectorized semantics for each field name and extracted entity is not limited in principle. Technicians can use pre-trained language models (such as Word2Vec, GloVe, etc.) to convert each word in the field name into a vector, and then aggregate the vectors of all words in the field name to obtain the vector representation of the field name, or directly use natural language processing (NLP) technology to generate the corresponding vectorized semantics, which will not be elaborated here.
[0086] The field name similarity threshold refers to the minimum similarity required for the corresponding vectorized semantics when aligning the field names. In the existing data fusion process, the field name similarity threshold is usually set between 0.6 and 0.85 to balance the strictness and flexibility of similarity and adapt to different data characteristics and business needs.
[0087] The entity recognition and extraction process, which uses a large AI model to extract entities and attributes, includes:
[0088] Choose any AI model. Here, AI models include, but are not limited to, any one or more of the following: OpenAI series models, Anthropic Claude series models, Google Gemini series, xAI Grok-2, DeepSeek, Kimi, and Claude. This will not be elaborated further.
[0089] The preprocessed data is input into the selected AI model, and entities and attributes are extracted through the AI model.
[0090] In the entity disambiguation and linking stage, a density-based density clustering algorithm is used to merge and extract entities. When merging and extracting entities, the density-based density clustering algorithm has the advantages of strong ability to handle noise and outliers, discovery of clusters of arbitrary shapes, no need to pre-specify the number of clusters, adaptability to clusters of different densities, high computational efficiency, and relatively simple parameter selection. These advantages make the density-based density clustering algorithm perform well in handling complex data distributions and large-scale datasets, and can more accurately identify and link entities, thereby improving the overall quality and reliability of knowledge base data fusion and association. In principle, the type of density clustering algorithm is not limited, and technicians can choose any one of DBSCAN or OPTICS. In principle, the domain radius of the density clustering algorithm is not limited, and can be determined according to the actual situation or K-distance map.
[0091] The entity similarity threshold refers to the required similarity between the extracted entity and the vectorized entity when the vectorized entity is considered as a candidate entity. The value range of the entity similarity threshold is set to 0.8-0.95 to ensure that the link constructed between the extracted entity and the vectorized entity is sufficiently reasonable.
[0092] The process of generating triples based on a large AI model includes:
[0093] The data that has been completed for entity disambiguation and linking will be input into the large AI model;
[0094] Identify entities, relationships, and attributes using large-scale AI models;
[0095] AI large model output: triples of entity pairs and relations (format: entity1, relation, entity2) and triples of entities and attributes (format: entity, attribute, attribute value);
[0096] The vector index construction and output phase involves using a graph database to build the vector index and providing a graph download interface.
[0097] Choose a graph database. In principle, there are no restrictions on the type of graph database, including but not limited to Neo4j, OrientDB, ArangoDB, etc.
[0098] The merged data is converted into a format supported by the graph database and then imported into the database.
[0099] Generate vector representations for each entity and relation;
[0100] Use the indexing capabilities of graph databases or external indexing tools (such as FAISS, HNSW, etc.) to build vector indexes;
[0101] Develop an API that allows users to query and download map data.
[0102] Further, please refer to Figure 2, which is a flowchart of the knowledge base data fusion and association method based on an AI large model in this embodiment of the present application. The workflow of the knowledge base data fusion and association method based on an AI large model described in this application includes:
[0103] S1: The acquisition module acquires the raw data based on the connector;
[0104] S2: The data preprocessing module cleans the raw data and removes duplicate data attribute values based on the Jaccard similarity threshold;
[0105] S3: The field alignment module obtains the field names and generates the corresponding vectorized semantics, and aligns the vectorized semantics based on the field name similarity threshold;
[0106] S4: The entity recognition and extraction module uses an AI large model to extract entities and attributes from the cleaned data and records them as extracted entities and extracted attributes;
[0107] S5: The entity disambiguation and linking module generates vectorized semantics corresponding to each extracted entity, then merges the extracted entities based on the density clustering algorithm, and retrieves vectorized entities from the database based on the entity similarity threshold. Vectorized entities whose similarity between the extracted entity and the vectorized entity is higher than the entity similarity threshold are recorded as candidate entities. If there are several candidate entities, a link is constructed between the extracted entity with the highest similarity and the vectorized entity, and the corresponding extracted entity is recorded as the successfully linked entity. If there are no candidate entities, a new entity is constructed.
[0108] S6: The triplet construction module generates triplets based on the AI large model;
[0109] S7: The normalization module assigns unified identifiers to entities and attributes to obtain fused data;
[0110] S8: The conflict detection and optimization module determines whether the current data conflict meets the requirements based on the coefficient of variation of the original data and the fused data, and, when it is determined that the requirements are not met, it determines the reason for not meeting the requirements based on the proportion of the duplicate hash value of the fused data in the total hash value, and, based on the corresponding reason, it corrects the data preprocessing parameters, field alignment parameters, entity disambiguation and linking parameters.
[0111] S9: The control module schedules the corresponding module to execute according to the instructions of the conflict detection and optimization module;
[0112] S10: The output module uses a graph database to construct a vector index and provides a graph download interface.
[0113] Furthermore, the process of determining whether the current data conflict meets the requirements based on the coefficient of variation of the original data and the fused data includes:
[0114] The original data and the fused data are obtained and quantized, and the original data and the fused data are aligned.
[0115] The absolute difference value of each field is calculated one by one to generate a difference value sequence, and the coefficient of variation of the original data and the fused data is determined based on the difference sequence;
[0116] If the coefficient of variation is less than or equal to the preset coefficient of variation, it is determined that the current data conflict meets the requirements, and the parameters of the knowledge base data fusion and association method remain unchanged.
[0117] If the coefficient of variation is greater than the preset coefficient of variation, it is determined that the current data conflict does not meet the requirements, and the reason for not meeting the requirements is determined based on the proportion of the duplicate hash value of the merged data in the total hash value.
[0118] Cross-source data typically have different scales, units, and field distributions. The coefficient of variation (CV), which is the ratio of the standard deviation to the mean, is scale-independent and therefore suitable for comparing differences in multi-field, different-scale data across cross-source datasets. By calculating the CV of the difference sequence between the original data and the fused data, the overall fluctuation level of the fusion process can be quantified. A low CV usually indicates that the difference distribution of the fused data is uniform and stable, and the conflict resolution is reasonable. A high CV, on the other hand, suggests that there may be problems such as incorrect conflict handling, information loss, or mismatch during the fusion process, requiring further investigation. Combined with the proportion of duplicate hash values, the root cause of the conflict can also be located, such as incomplete deduplication or merging conflicts. Using the CV to determine whether the conflicts in the fused data are reasonable can provide a fast and quantifiable benchmark for assessing conflict stability, thereby quickly judging the quality of the fused result.
[0119] Specifically, the process of determining whether the current data conflict meets the requirements based on the coefficient of variation of the original data and the fused data includes:
[0120] The process of acquiring the original data and the fused data and quantizing them, and aligning the original data and the fused data, includes:
[0121] The conflict detection and optimization module obtains the field names of the original data and the fused data and generates corresponding vectorized semantics. It then aligns the vectorized semantics based on the field alignment threshold. The field alignment threshold AT is not limited in principle and can be adjusted by technicians according to their needs. In this solution, the field alignment threshold AT can be set to a range of 0.85-0.95 to avoid misalignment.
[0122] The conflict detection and optimization module uses an AI large model to extract the original data entities and attributes and quantize them to obtain original data entity vectors and original entity attribute vectors. The original data entity vectors and original entity attribute vectors are then aligned with the attributes corresponding to the extracted entities. The alignment threshold for entity vectors is set between 0.85 and 0.95 to strictly avoid misalignment, while the alignment threshold for attribute vectors is set between 0.75 and 0.85 to balance efficiency and accuracy.
[0123] Absolute difference value AD i =|Original Data Attributes RDA i -Data attributes after fusion FDA i |, where i = 1, 2, ..., n;
[0124] Calculate the absolute difference value for each attribute one by one, generating a difference value sequence AD1, AD2, ..., AD n .
[0125] The coefficient of variation between the original data and the fused data is determined based on the differential sequence.
[0126] The conflict detection and optimization module compares the coefficient of variation CV with the preset coefficient of variation CV1, wherein the preset coefficient of variation CV1 is set to [0.05, 0.15]. For the coefficient of variation CV, when the coefficient of variation is less than or equal to 0.15, the risk of data conflict is low. At the same time, for databases such as financial and medical data, there are higher requirements. When the coefficient of variation is 0.05, the embodiment of this application can avoid data conflict and better play the role of data fusion to provide decision support for users.
[0127] If the coefficient of variation CV is less than or equal to the preset coefficient of variation CV1, the conflict detection and optimization module determines that the current data conflict meets the requirements and keeps the parameters of the knowledge base data fusion and association method unchanged.
[0128] If the coefficient of variation CV is greater than the preset coefficient of variation CV1, the conflict detection and optimization module determines that the current data conflict does not meet the requirements. The conflict detection and optimization module determines the reason for the non-compliance based on the proportion of the duplicate hash value of the merged data in the total hash value.
[0129] Furthermore, the process by which this application determines the reasons for dissatisfaction based on the proportion of duplicate hash values in the total hash value after fusion includes:
[0130] The fused data is then converted into hash values.
[0131] The hash collision rate is determined based on the proportion of duplicate hash values in the total number of hash values in the merged data.
[0132] If the hash collision rate is less than or equal to the preset hash collision rate, the reason for not meeting the requirement is determined based on the proportion of successfully linked entities in the extracted entities.
[0133] If the hash collision rate is greater than the preset hash collision rate, it is determined that the Jaccard similarity threshold setting is unreasonable, and the Jaccard similarity threshold is corrected based on the ratio of the hash collision rate to the preset hash collision rate.
[0134] The Jaccard similarity threshold refers to the lowest Jaccard similarity between any two data attribute values when they are considered to be duplicated.
[0135] High redundancy indicates more inconsistencies and potential conflicts among similar or duplicated data, making it difficult to eliminate contradictions during the fusion process. The likelihood and severity of conflicts increase significantly. Therefore, controlling and properly handling data redundancy during data fusion is crucial for ensuring fusion quality and accuracy. In this solution, the proportion of duplicate hash values in the total hash values after fusion (hash collision rate) can effectively determine the cause of conflicts during data fusion. The hash collision rate reflects the degree of data duplication or similarity, helping this solution quickly locate the cause of anomalies. When the hash collision rate exceeds a preset threshold, it indicates that the Jaccard similarity threshold is set too high, resulting in insufficient deduplication. In this case, the Jaccard similarity threshold needs to be adjusted to optimize the deduplication strategy. This method has advantages such as efficient calculation, ease of quantification, and direct reflection of data duplication and fusion effects. It can efficiently, objectively, and automatically identify and handle fusion conflicts, improving the accuracy and robustness of knowledge base data fusion. It is a reasonable and effective technical solution.
[0136] Please refer to Figure 3, which is a flowchart of the process for determining the reasons for non-compliance based on the hash collision rate in this embodiment of the application. The process for determining the reasons for non-compliance based on the proportion of duplicate hash values in the total hash value after merging in this embodiment of the application includes:
[0137] The conflict detection and optimization module converts the fused data into hash values using a hash algorithm. In principle, the hash algorithm is not limited, and technicians can choose any hash algorithm such as MD5, SHA-1, or SHA-256 to convert the fused data into hash values.
[0138] The collision detection and optimization module determines the hash collision rate (HCR) based on the proportion of duplicate hash values in the fused data to the total number of hash values.
[0139] The collision detection and optimization module compares the hash collision rate HCR with the preset hash collision rate HCR1, where the preset hash collision rate HCR1 is set to [3%, 15%]. When determining the preset hash collision rate HCR1, the characteristics of the hash algorithm, practical application experience, and performance trade-offs need to be comprehensively considered. For high-precision knowledge bases (such as those in the medical and financial fields), a lower hash collision rate is often required. In this case, the preset hash collision rate HCR1 is usually guaranteed to be between 3% and 5% to ensure the accuracy of the results. For databases in the fields of text processing (5%-10%) and image processing (10%-15%, because visual features themselves have a certain degree of ambiguity), a slightly higher hash collision rate can be accepted. Therefore, in this scheme, the value range of the preset hash collision rate is set to 3%-15%.
[0140] If the hash collision rate HCR is less than or equal to the preset hash collision rate HCR1, the collision detection and optimization module determines the reason for non-compliance based on the proportion of successfully linked entities in the extracted entities.
[0141] If the hash collision rate HCR is greater than the preset hash collision rate HCR1, the collision detection and optimization module determines that the Jaccard similarity threshold setting is unreasonable, and corrects the Jaccard similarity threshold based on the ratio of the hash collision rate to the preset hash collision rate.
[0142] Furthermore, the process of adjusting the Jaccard similarity threshold based on the ratio of the hash collision rate to the preset hash collision rate includes:
[0143] The hash collision rate ratio is determined based on the ratio of the hash collision rate to the preset hash collision rate;
[0144] The Jaccard similarity threshold is reduced based on the hash collision rate ratio, and the reduction in the Jaccard similarity threshold is proportional to the hash collision rate ratio.
[0145] Hash collision rate refers to the proportion of duplicate hash values in the merged data to the total number of hash values. Hash collision rate ratio is the ratio of the current hash collision rate to the preset hash collision rate. The hash collision rate ratio reflects the degree of deviation between the current hash collision rate and the preset value. A higher ratio means that there is a lot of duplicate data that has not been processed correctly during the data fusion process. Therefore, by lowering the Jaccard similarity threshold, the probability of data being identified as duplicates can be increased, thereby reducing the hash collision rate and improving the quality of data fusion. At the same time, this dynamic adjustment method based on the hash collision rate ratio can achieve adaptive optimization of the system without frequent manual intervention, improving the automation and efficiency of the system, enabling the system to better adapt to dynamic changes in data and maintain the stability and reliability of data fusion.
[0146] Please refer to Figure 4, which is a flowchart of the process for correcting the Jaccard similarity threshold based on the hash collision rate ratio in an embodiment of this application. The process of correcting the Jaccard similarity threshold based on the ratio of the hash collision rate to the preset hash collision rate in this application includes:
[0147] The conflict detection and optimization module determines a hash conflict rate ratio A based on the ratio of the hash conflict rate to the preset hash conflict rate, and compares the hash conflict rate ratio A with two preset hash conflict rate ratios: a first preset hash conflict rate ratio A1 and a second preset hash conflict rate ratio A2. The first preset hash conflict rate ratio A1 ∈ (1, 2), and the second preset hash conflict rate ratio A2 ∈ [2, 3]. Setting the first preset hash conflict rate ratio A1 ∈ (1, 2) and the second preset hash conflict rate ratio A2 ∈ [2, 3] helps to implement a hierarchical adjustment mechanism while more precisely controlling the Jaccard similarity threshold according to the actual severity of hash conflicts. The adjustment range is as follows: when the actual calculated hash collision rate ratio A is between 1 and 2, only a slight adjustment to the Jaccard similarity threshold is needed. When the hash collision rate ratio A reaches 2, it means that the actual hash collision rate has become twice the ideal hash collision rate, and the adjustment mechanism needs to be further upgraded. When its value is 3, it may mean that there may be more serious data collisions in the system, and a larger adjustment is needed. This hierarchical setting can more accurately deal with data collisions of different degrees, avoid misjudgments due to over-adjustment, and provide greater flexibility and adaptability to deal with the collision characteristics of different datasets and application scenarios.
[0148] If the hash collision rate ratio A is less than or equal to the first preset hash collision rate ratio A1, the collision detection and optimization module uses the first Jaccard similarity correction threshold α1 to correct the Jaccard similarity threshold JST. The corrected Jaccard similarity threshold JST' = JST × α1, where the first Jaccard similarity correction threshold α1 is set to 0.98.
[0149] If the hash collision rate ratio A is greater than the first preset hash collision rate ratio A1 and less than or equal to the second preset hash collision rate ratio A2, then the collision detection and optimization module uses the second Jaccard similarity correction threshold α2 to correct the Jaccard similarity threshold JST. The corrected Jaccard similarity threshold JST' = JST × α2, where the second Jaccard similarity correction threshold α2 is set to 0.95.
[0150] If the hash collision rate ratio A is greater than the second preset hash collision rate ratio A2, the collision detection and optimization module uses a third Jaccard similarity correction threshold α3 to correct the Jaccard similarity threshold JST. The corrected Jaccard similarity threshold JST' = JST × α3, where the third Jaccard similarity correction threshold α3 is set to 0.9.
[0151] Furthermore, the process of determining the reason for dissatisfaction based on the proportion of successfully linked entities in the extracted entities includes:
[0152] The success rate of linking entities is determined based on the proportion of successfully linked entities in the extracted entities.
[0153] If the success rate of extracting entity links is less than or equal to the preset success rate of extracting entity links, it is determined that the entity similarity threshold setting is unreasonable, and the entity similarity threshold is corrected based on the ratio of the preset success rate of extracting entity links to the success rate of extracting entity links.
[0154] If the success rate of extracting entity links is greater than the preset success rate of extracting entity links, then the reason for non-compliance is determined based on the proportion of the number of field name similarities distributed in the field name similarity range to the total number of field name similarities.
[0155] The field name similarity range refers to the similarity range that is evenly distributed on both sides of the field name similarity threshold, including the field name similarity threshold.
[0156] The entity linking success rate refers to the proportion of extracted entities that are successfully linked to vectorized entities in the database. It reflects the degree of success in matching extracted entities with entities in the database during entity disambiguation and linking. If the entity linking success rate is low, it means that many extracted entities failed to link to vectorized entities in the database during entity disambiguation and linking. This may be because the entity similarity threshold is set too high, causing some actually similar entities to not be correctly identified as matching entities. The entity linking success rate provides a quantitative indicator that can intuitively evaluate the effectiveness of the entity disambiguation and linking process, accurately locate the causes of anomalies, and further improve the performance and reliability of the entire knowledge base data fusion and association method.
[0157] Specifically, the process of determining the reason for dissatisfaction based on the proportion of successfully linked entities in the extracted entities includes:
[0158] The conflict detection and optimization module determines the link success rate B of the extracted entities based on the proportion of successfully linked entities in the extracted entities.
[0159] The conflict detection and optimization module compares the entity linking success rate B with the preset entity linking success rate B1. The preset entity linking success rate B1 is set to [80%, 90%]. According to the KDD 2024 benchmark test, in the medical knowledge base fusion test, an 88% linking success rate can guarantee an entity disambiguation accuracy of >92%. For knowledge fusion tests in the social media field, a slightly lower entity linking success rate (80%-85%) is acceptable. Therefore, in this solution, the preset entity linking success rate B1 is set to [80%, 90%].
[0160] If the entity link extraction success rate B is less than or equal to the preset entity link extraction success rate B1, then the entity similarity threshold setting is determined to be unreasonable, and the conflict detection and optimization module corrects the entity similarity threshold based on the ratio of the preset entity link extraction success rate to the entity link extraction success rate.
[0161] If the success rate of extracting entity links B is greater than the preset success rate of extracting entity links B1, the conflict detection and optimization module determines the reason for non-compliance based on the proportion of the number of field name similarities distributed in the field name similarity range to the total number of field name similarities.
[0162] Furthermore, the process of adjusting the entity similarity threshold based on the ratio of the preset entity linking success rate to the entity linking success rate includes:
[0163] The entity link success rate ratio is determined based on the ratio of the preset entity link extraction success rate to the entity link extraction success rate.
[0164] The entity similarity threshold is reduced based on the entity link success rate ratio, and the reduction in the entity similarity threshold is proportional to the entity link success rate ratio.
[0165] The entity link success rate ratio refers to the ratio of the preset entity link success rate to the actual entity link success rate. The entity link success rate ratio reflects the degree of deviation of the current link success rate from the expected value. When the entity link success rate ratio is high, it indicates that the actual link success rate is much lower than the preset value, which means that the current entity similarity threshold is too high and needs to be lowered to improve the link success rate. By making the reduction proportional to the entity link success rate ratio, dynamic adjustment can be achieved, enabling the system to better adapt to changes in data, improve the link success rate, reduce false judgments, improve system performance, and has high interpretability and operability.
[0166] Specifically, the process of adjusting the entity similarity threshold based on the ratio of the preset entity linking success rate to the entity linking success rate includes:
[0167] The conflict detection and optimization module determines the entity link success rate ratio C based on the ratio of the preset entity link success rate to the entity link success rate.
[0168] The conflict detection and optimization module compares the entity link success rate ratio C with the set first preset entity link success rate ratio C1 and second preset entity link success rate ratio C2. The first preset entity link success rate ratio C1 is set to [1.05, 1.2), and the second preset entity link success rate ratio C2 is set to [1.2, 1.4]. According to the database system operation data monitoring, when the entity link success rate is lower than 75%, it will cause significant data conflict problems. When the link rate is lower than 70%, the conflict will become more serious. Since the entity link success rate ratio C is the preset ratio of the extracted entity link success rate to the extracted entity link success rate, in this solution, the first preset entity link success rate ratio C1 is set to [1.05, 1.2), and the second preset entity link success rate ratio C2 is set to [1.2, 1.4].
[0169] If the entity link success rate ratio C is less than or equal to the first preset entity link success rate ratio C1, then the conflict detection and optimization module uses the first entity similarity correction threshold β1 to correct the entity similarity threshold EST. The corrected entity similarity threshold EST' = β1 × EST, where the first entity similarity correction threshold β1 is set to 0.99.
[0170] If the entity link success rate ratio C is greater than the first preset entity link success rate ratio C1 and less than or equal to the second preset entity link success rate ratio C2, then the conflict detection and optimization module uses the second entity similarity correction threshold β2 to correct the entity similarity threshold EST. The corrected entity similarity threshold EST' = β2 × EST, where the second entity similarity correction threshold β2 is set to 0.96.
[0171] If the entity link success rate ratio C is greater than the second preset entity link success rate ratio C2, then the conflict detection and optimization module uses a third entity similarity correction threshold β3 to correct the entity similarity threshold EST. The corrected entity similarity threshold EST' = β3 × EST, where the third entity similarity correction threshold β3 is set to 0.92.
[0172] Furthermore, the process of determining the reason for non-compliance based on the proportion of the number of field name similarities distributed within the range of field name similarities to the total number of field name similarities includes:
[0173] The distribution ratio of field name similarity is determined based on the proportion of the number of field name similarities distributed within the range of field name similarities in the total number of field name similarities;
[0174] If the proportion of field name similarity distribution is less than or equal to the preset proportion of field name similarity distribution, it is determined that the density clustering algorithm parameter setting is unreasonable, and the density clustering algorithm neighborhood radius is corrected based on the ratio of the coefficient of variation to the preset coefficient of variation.
[0175] If the proportion of field name similarity distribution is greater than the proportion of preset field name similarity distribution, it is determined that the field name similarity threshold setting is unreasonable, and the field name similarity threshold is corrected based on the difference between the proportion of field name similarity distribution and the proportion of preset field name similarity distribution.
[0176] The field name similarity threshold refers to the minimum similarity required for the corresponding vectorized semantics when aligning the field names.
[0177] The field name similarity distribution ratio refers to the proportion of field name similarities within the similarity range to the total number of field name similarities. It reflects the distribution of field name similarities around a similarity threshold during field alignment. A high distribution ratio indicates that most fields are distributed around the similarity threshold, suggesting the threshold may be improperly set, leading to misalignment and affecting the fusion and association effect. Conversely, a low distribution ratio indicates that most field name pairs are either significantly similar or significantly dissimilar. This anomaly may stem from improper clustering parameter settings, such as cluster radius, which fails to properly group similar field names into the same cluster, resulting in decreased entity disambiguation. Identifying the cause of anomalies through the field name similarity distribution ratio enables data-driven adaptive parameter adjustments, accurately pinpointing problems, improving alignment quality and system stability, and is simple to operate and easy to implement.
[0178] Specifically, the process of determining the reason for non-compliance based on the proportion of the number of field name similarities distributed within the range of field name similarities to the total number of field name similarities includes:
[0179] The conflict detection and optimization module determines the field name similarity distribution ratio D based on the proportion of the number of field name similarities distributed in the field name similarity range to the total number of field name similarities, where the field name similarity range = field name similarity threshold FNST ± γ, γ∈[0.05,0.2];
[0180] The conflict detection and optimization module compares the field name similarity distribution ratio D with the preset field name similarity distribution ratio D1. The preset field name similarity distribution ratio D1 ∈ [40%, 50%]. If the field name similarity distribution is relatively uniform or concentrated at both ends (0 and 1), it indicates that the field name similarity threshold is set reasonably. However, if the field name similarity is concentrated on both sides of the field name similarity threshold, it may be that the field name similarity threshold is set unreasonably. When the field name similarity distribution ratio reaches 40%, the number of field names in the similarity range has already accounted for two-fifths of the total, indicating that the field name similarity threshold is set very unreasonably. Therefore, in order to ensure flexibility, this scheme sets the preset field name similarity distribution ratio D1 ∈ [40%, 50%].
[0181] If the proportion of field name similarity distribution D is less than or equal to the proportion of preset field name similarity distribution D1, the conflict detection and optimization module determines that the density clustering algorithm parameter settings are unreasonable, and corrects the neighborhood radius of the density clustering algorithm based on the ratio of the coefficient of variation to the preset coefficient of variation.
[0182] If the proportion of field name similarity distribution D is greater than the preset proportion of field name similarity distribution D1, the conflict detection and optimization module determines that the field name similarity threshold setting is unreasonable, and corrects the field name similarity threshold based on the difference between the proportion of field name similarity distribution and the preset proportion of field name similarity distribution.
[0183] Furthermore, the process of correcting the neighborhood radius of the density clustering algorithm based on the ratio of the coefficient of variation to the preset coefficient of variation includes:
[0184] The coefficient of variation ratio is determined based on the ratio of the coefficient of variation to the preset coefficient of variation;
[0185] The radius of the density clustering algorithm is increased based on the coefficient of variation ratio, and the increase in the radius of the density clustering algorithm is proportional to the coefficient of variation ratio.
[0186] The coefficient of variation ratio refers to the ratio of the coefficient of variation to the preset coefficient of variation. The coefficient of variation ratio reflects the degree of fluctuation of the current data. When the coefficient of variation ratio is greater than 1, it indicates that there are large fluctuations and conflicts in the data fusion process. This may be due to the small neighborhood radius of the density clustering algorithm. By increasing the neighborhood radius, the clustering effect of field names can be improved and the conflicts in the data fusion process can be reduced. By making the increase of the neighborhood radius proportional to the coefficient of variation ratio, dynamic adjustment can be achieved, so that the system can automatically adjust the clustering parameters according to the actual data fluctuation, thereby better adapting to different datasets and application scenarios.
[0187] Specifically, the process of adjusting the neighborhood radius of the density clustering algorithm based on the ratio of the coefficient of variation to the preset coefficient of variation includes:
[0188] The conflict detection and optimization module determines the coefficient of variation ratio E based on the ratio of the coefficient of variation to the preset coefficient of variation;
[0189] The conflict detection and optimization module compares the coefficient of variation ratio E with the set first preset coefficient of variation ratio E1 and second preset coefficient of variation ratio E2. The first preset coefficient of variation ratio E1 is set to (1, 2) and the second preset coefficient of variation ratio E2 is set to [2, 4]. According to relevant research in the field of data fusion, when the coefficient of variation between the original data and the fused data exceeds 0.2, there may be a large risk of data conflict. When the coefficient of variation exceeds 0.3, it may indicate a serious data conflict. Therefore, in this scheme, the first preset coefficient of variation ratio E1 is set to (1, 2) and the second preset coefficient of variation ratio E2 is set to [2, 4].
[0190] If the coefficient of variation ratio E is less than or equal to the first preset coefficient of variation ratio E1, the conflict detection and optimization module uses the first neighborhood radius correction threshold θ1 to correct the neighborhood radius R of the density clustering algorithm. The corrected neighborhood radius R' of the density clustering algorithm is R × θ1, where the first neighborhood radius correction threshold θ1 is set to 1.02.
[0191] If the coefficient of variation ratio E is greater than the first preset coefficient of variation ratio E1 and less than or equal to the second preset coefficient of variation ratio E2, then the conflict detection and optimization module uses the second neighborhood radius correction threshold θ2 to correct the neighborhood radius R of the density clustering algorithm. The corrected neighborhood radius R' of the density clustering algorithm is R × θ2, where the second neighborhood radius correction threshold θ2 is set to 1.05.
[0192] If the coefficient of variation ratio E is greater than the second preset coefficient of variation ratio E2, the conflict detection and optimization module uses the third neighborhood radius correction threshold θ3 to correct the neighborhood radius R of the density clustering algorithm. The corrected neighborhood radius R' = R × θ3, where the third neighborhood radius correction threshold θ3 is set to 1.09.
[0193] Furthermore, the process of adjusting the field name similarity threshold based on the difference between the field name similarity distribution ratio and the preset field name similarity distribution ratio includes:
[0194] The distribution ratio difference is determined based on the difference between the distribution ratio of the field name similarity and the preset distribution ratio of the field name similarity.
[0195] The field name similarity threshold is reduced based on the difference in distribution proportions, and the reduction in the field name similarity threshold is proportional to the difference in distribution proportions.
[0196] The distribution ratio difference refers to the difference between the distribution ratio of field name similarity and the preset distribution ratio of field name similarity. It reflects the degree of deviation between the distribution of boundary scenes and the preset target during the current field name alignment process. When the distribution ratio difference is positive, it indicates that the current field name similarity threshold is set too low, resulting in a large number of field name pairs being near the similarity threshold, leading to high uncertainty in alignment determination. Therefore, it is necessary to lower the field name similarity threshold, with the reduction being proportional to the distribution ratio difference, to achieve dynamic adjustment and improve the clarity and accuracy of alignment.
[0197] Specifically, the process of adjusting the field name similarity threshold based on the difference between the field name similarity distribution ratio and the preset field name similarity distribution ratio includes:
[0198] The conflict detection and optimization module determines the distribution ratio difference F based on the difference between the field name similarity distribution ratio and the preset field name similarity distribution ratio;
[0199] The conflict detection and optimization module compares the distribution ratio difference F with two preset distribution ratio differences: F1 and F2. The first preset distribution ratio difference F1 ∈ [0, 10%] and the second preset distribution ratio difference F2 ∈ (10%, 20%). Setting the first preset distribution ratio difference F1 ∈ [0, 10%] and the second preset distribution ratio difference F2 ∈ (10%, 20%) allows for more flexible adaptation to different data distributions, improving the system's adaptability and robustness. Through experimental verification and data distribution characteristic analysis, reasonable F1 and F2 values can be determined, ensuring that the system maintains good performance when facing different datasets.
[0200] If the distribution ratio difference F is less than or equal to the first preset distribution ratio difference F1, the conflict detection and optimization module uses the first field name correction threshold λ1 to correct the field name similarity threshold FNST. The corrected field name similarity threshold FNST' = λ1 × FNST, where the first field name correction threshold λ1 is set to 0.99.
[0201] If the distribution ratio difference F is greater than the first preset distribution ratio difference F1 and less than or equal to the second preset distribution ratio difference F2, then the conflict detection and optimization module uses the second field name correction threshold λ2 to correct the field name similarity threshold FNST. The corrected field name similarity threshold FNST' = λ2 × FNST, where the second field name correction threshold λ2 is set to 0.95.
[0202] If the distribution ratio difference F is greater than the second preset distribution ratio difference F2, the conflict detection and optimization module uses the third field name correction threshold λ3 to correct the field name similarity threshold FNST. The corrected field name similarity threshold FNST' = λ3 × FNST, where the third field name correction threshold λ3 is set to 0.9.
[0203] In summary, this application provides a method and system for knowledge base data fusion and association based on an AI large-scale model. Through a series of steps including data acquisition, preprocessing, field alignment, entity recognition and extraction, disambiguation and linking, relation extraction and triple construction, normalization, conflict detection, and optimization, it achieves efficient fusion of knowledge base data. This method effectively eliminates data redundancy and inconsistency, fills data gaps, and improves data accuracy, thereby enhancing data quality. Simultaneously, the introduction of an intelligent conflict detection and optimization mechanism can specifically identify the causes of data conflicts and provide corresponding solutions, overcoming the difficulty in resolving data conflicts in existing database fusion methods. Finally, through vector index construction and output, it provides users with standardized, easy-to-use, and shareable fused data, promoting the development of knowledge base data fusion technology.
[0204] The technical solutions of this application have been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of this application is obviously not limited to these specific embodiments. Without departing from the principles of this application, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of this application.
Claims
1. A knowledge base data fusion and association method based on an AI large-scale model, characterized in that, include: Data acquisition: Raw data is collected via connectors. Data preprocessing involves cleaning the original data and removing duplicates from the cleaned data based on the Jaccard similarity threshold. Field alignment involves obtaining the names of each field and generating corresponding vectorized semantics, and then aligning the vectorized semantics based on the field name similarity threshold. Entity recognition and extraction: AI large model is used to extract entities and attributes from the cleaned data and record them as extracted entities and extracted attributes. Entity disambiguation and linking: Generate corresponding vectorized semantics for each extracted entity; merge extracted entities based on density clustering algorithm; retrieve vectorized entities from the database based on entity similarity threshold; and record vectorized entities whose similarity with the extracted entity is higher than the entity similarity threshold as candidate entities. If there are several candidate entities, a link is constructed between the extracted entity with the highest similarity and the vectorized entity, and the corresponding extracted entity is recorded as a successfully linked entity. If there are no candidate entities, a new entity is constructed. Relation extraction and triple construction: generating triples based on a large AI model; Standardization and normalization assign unified identifiers to entities and attributes to obtain merged data. Conflict detection and optimization: Based on the coefficient of variation of the original data and the fused data, it is determined whether the current data conflict meets the requirements. When a requirement is not met, the reason for the non-compliance is determined based on the proportion of duplicate hash values in the total hash value after fusion. In addition, the data preprocessing parameters, field alignment parameters, and entity disambiguation and linking parameters are corrected based on the corresponding reasons; Vector index construction and output: The vector index is constructed using a graph database, and a graph download interface is provided.
2. The knowledge base data fusion and association method according to claim 1, characterized in that, The process of determining whether the current data conflict meets the requirements based on the coefficient of variation of the original data and the fused data includes: The original data and the fused data are obtained and quantized, and the original data and the fused data are aligned. The absolute difference value of each field is calculated one by one to generate a difference value sequence, and the coefficient of variation of the original data and the fused data is determined based on the difference sequence; If the coefficient of variation is less than or equal to the preset coefficient of variation, it is determined that the current data conflict meets the requirements, and the parameters of the knowledge base data fusion and association method remain unchanged. If the coefficient of variation is greater than the preset coefficient of variation, it is determined that the current data conflict does not meet the requirements, and the reason for not meeting the requirements is determined based on the proportion of the duplicate hash value of the merged data in the total hash value.
3. The knowledge base data fusion and association method according to claim 2, characterized in that, The process of determining the reasons for dissatisfaction based on the proportion of duplicate hash values in the total hash value after fusion includes: The fused data is then converted into hash values. The hash collision rate is determined based on the proportion of duplicate hash values in the total number of hash values in the merged data. If the hash collision rate is less than or equal to the preset hash collision rate, the reason for not meeting the requirement is determined based on the proportion of successfully linked entities in the extracted entities. If the hash collision rate is greater than the preset hash collision rate, it is determined that the Jaccard similarity threshold setting is unreasonable, and the Jaccard similarity threshold is corrected based on the ratio of the hash collision rate to the preset hash collision rate. The Jaccard similarity threshold refers to the minimum Jaccard similarity required between any two data attribute values when they are considered to be duplicated.
4. The knowledge base data fusion and association method according to claim 3, characterized in that, The process of adjusting the Jaccard similarity threshold based on the ratio of the hash collision rate to the preset hash collision rate includes: The hash collision rate ratio is determined based on the ratio of the hash collision rate to the preset hash collision rate; The Jaccard similarity threshold is reduced based on the hash collision rate ratio, and the reduction in the Jaccard similarity threshold is proportional to the hash collision rate ratio.
5. The knowledge base data fusion and association method according to claim 3, characterized in that, The process of determining the reason for dissatisfaction based on the proportion of successfully linked entities in the extracted entities includes: The success rate of linking entities is determined based on the proportion of successfully linked entities in the extracted entities. If the success rate of extracting entity links is less than or equal to the preset success rate of extracting entity links, it is determined that the entity similarity threshold setting is unreasonable, and the entity similarity threshold is corrected based on the ratio of the preset success rate of extracting entity links to the success rate of extracting entity links. If the success rate of extracting entity links is greater than the preset success rate of extracting entity links, the reason for non-compliance is determined based on the proportion of the number of field name similarities distributed in the field name similarity range to the total number of field name similarities. Wherein, the entity similarity threshold refers to the similarity that the extracted entity and the vectorized entity need to reach when the vectorized entity is recorded as a candidate entity, and the field name similarity range refers to the similarity range that is evenly distributed on both sides of the field name similarity threshold and includes the field name similarity threshold.
6. The knowledge base data fusion and association method according to claim 5, characterized in that, The process of adjusting the entity similarity threshold based on the ratio of the preset entity link success rate to the entity link success rate includes: The entity link success rate ratio is determined based on the ratio of the preset entity link extraction success rate to the entity link extraction success rate. The entity similarity threshold is reduced based on the entity link success rate ratio, and the reduction in the entity similarity threshold is proportional to the entity link success rate ratio.
7. The knowledge base data fusion and association method according to claim 5, characterized in that, The process of determining the reasons for non-compliance based on the proportion of the number of field name similarities distributed within the range of field name similarities to the total number of field name similarities includes: The distribution ratio of field name similarity is determined based on the proportion of the number of field name similarities distributed within the range of field name similarities in the total number of field name similarities; If the proportion of field name similarity distribution is less than or equal to the preset proportion of field name similarity distribution, it is determined that the density clustering algorithm parameter setting is unreasonable, and the density clustering algorithm neighborhood radius is corrected based on the ratio of the coefficient of variation to the preset coefficient of variation. If the proportion of field name similarity distribution is greater than the proportion of preset field name similarity distribution, it is determined that the field name similarity threshold setting is unreasonable, and the field name similarity threshold is corrected based on the difference between the proportion of field name similarity distribution and the proportion of preset field name similarity distribution. The field name similarity threshold refers to the minimum similarity required for the corresponding vectorized semantics when aligning the field names.
8. The knowledge base data fusion and association method according to claim 7, characterized in that, The process of correcting the neighborhood radius of the density clustering algorithm based on the ratio of the coefficient of variation to the preset coefficient of variation includes: The coefficient of variation ratio is determined based on the ratio of the coefficient of variation to the preset coefficient of variation; The radius of the density clustering algorithm is increased based on the coefficient of variation ratio, and the increase in the radius of the density clustering algorithm is proportional to the coefficient of variation ratio.
9. The knowledge base data fusion and association method according to claim 7, characterized in that, The process of adjusting the field name similarity threshold based on the difference between the field name similarity distribution ratio and the preset field name similarity distribution ratio includes: The distribution ratio difference is determined based on the difference between the distribution ratio of the field name similarity and the preset distribution ratio of the field name similarity. The field name similarity threshold is reduced based on the difference in distribution proportions, and the reduction in the field name similarity threshold is proportional to the difference in distribution proportions.
10. A knowledge base data fusion and association system based on an AI large-scale model, characterized in that, Used to perform the knowledge base data fusion and association method based on the AI large model as described in any one of claims 1-9 include, The acquisition module includes several connectors for acquiring the raw data; A data preprocessing module, which is connected to the acquisition module, is used to clean the data and remove duplicates from the cleaned data based on the Jaccard similarity threshold. The field alignment module, which is connected to the data preprocessing module, is used to obtain the field names and generate the corresponding vectorized semantics, and align the vectorized semantics based on the field name similarity threshold. The entity recognition and extraction module, which is connected to the field alignment module, is used to extract entities and attributes from the cleaned data based on the AI large model, and denoted as the extracted entity and extracted attribute. The entity disambiguation and linking module, which is connected to the entity recognition and extraction module, is used to generate vectorized semantics corresponding to each extracted entity, and to merge extracted entities based on density clustering algorithm, and to retrieve vectorized entities in the database based on entity similarity threshold. Vectorized entities whose similarity between extracted entities and vectorized entities is higher than the entity similarity threshold are recorded as candidate entities. If there are several candidate entities, a link is constructed between the extracted entity with the highest similarity and the vectorized entity, and the corresponding extracted entity is recorded as a successfully linked entity. If there are no candidate entities, a new entity is constructed. The triplet construction module is connected to the entity disambiguation and linking module to generate triplets based on the AI large model; The normalization module, which is connected to the triplet construction module, is used to assign unified identifiers to entities and attributes to obtain fused data; The conflict detection and optimization module, which is connected to the normalization module, is used to determine whether the current data conflict meets the requirements based on the coefficient of variation of the original data and the fused data. When a requirement is not met, the reason for the non-compliance is determined based on the proportion of duplicate hash values in the total hash value after fusion. In addition, the data preprocessing parameters, field alignment parameters, and entity disambiguation and linking parameters are corrected based on the corresponding reasons; The control module, which is connected to the conflict detection and optimization module, the data preprocessing module, the field alignment module, and the entity disambiguation and linking module, is used to schedule the corresponding modules to execute according to the instructions of the conflict detection and optimization module. The output module, which is connected to the normalization module and the conflict detection and optimization module, is used to construct a vector index based on the graph database and provide a graph download interface.