Retrieval processing method, device and equipment based on multi-source knowledge base data
By standardizing and identifying duplicate entities in multi-source heterogeneous data, a knowledge graph is constructed, which solves the problems of difficulty in integrating multi-source heterogeneous data and insufficient real-time response capability. It enables efficient and accurate cross-database retrieval and real-time updates, improving data management and query efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GTCOM INFORMATION TECH (SHANGHAI) CO LTD
- Filing Date
- 2025-12-05
- Publication Date
- 2026-05-08
AI Technical Summary
The existing strategic database suffers from difficulties in integrating multi-source heterogeneous data, high data redundancy, severe query response delays, insufficient knowledge retrieval accuracy, and lack of data real-time performance, making it impossible to achieve real-time response and efficient cross-database association.
By acquiring heterogeneous data from multiple sources, standardizing data formats and terminology, identifying and removing duplicate entities, constructing a knowledge graph, enabling cross-database retrieval and real-time updates, and utilizing the knowledge graph for knowledge-enhanced queries.
It improves the accuracy of cross-database retrieval, reduces data redundancy, enhances real-time response capabilities, and improves the accuracy and efficiency of knowledge retrieval, with a recall rate of over 95% and a latency of less than 2 seconds.
Smart Images

Figure CN121996799A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer information technology processing technology, and in particular to a retrieval and processing method, apparatus and equipment based on multi-source knowledge base data. Background Technology
[0002] Current strategic databases often employ an isolated architecture (such as independent talent pools or policy databases), which presents the following problems in data processing: The integration of multi-source heterogeneous data is difficult, and it is impossible to efficiently clean and standardize talent, institutional, and policy data in various formats (such as Excel and JSON), resulting in high data redundancy (>15%) and low quality (null value rate >10%). The talent pool, institutional pool, and policy pool operate independently, lacking cross-database linkage capabilities, and query response latency exceeds 5 seconds.
[0003] Knowledge retrieval accuracy is insufficient. Traditional keyword retrieval only supports single-field matching and cannot achieve fuzzy matching across multiple fields such as talent name and research field, resulting in a recall rate of less than 70%. Implicit relationships between policies, institutions, and talents (such as "the impact of policies on talents in specific fields") have not been explored due to a lack of knowledge graph modeling and reasoning capabilities.
[0004] The data suffers from real-time limitations, relying on manual batch processing for updates, resulting in delays exceeding one hour and an inability to respond promptly to talent mobility or policy changes. Incremental updates also exhibit a high conflict rate (>10%) due to the lack of real-time synchronization and conflict resolution mechanisms. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a retrieval and processing method, apparatus and equipment based on multi-source knowledge base data, to uniformly manage multi-source data, improve the accuracy of cross-database retrieval through knowledge graphs, and improve the real-time response capability of data.
[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: A retrieval and processing method based on multi-source knowledge base data includes: Acquire multi-source heterogeneous data; The data format and target terminology of the multi-source heterogeneous data are standardized to obtain standard data; The standard data is subjected to duplicate entity identification to remove duplicate entities from the standard data, thereby obtaining the target data; Based on the multi-source heterogeneous data, obtain the entity relationships in the multi-source heterogeneous data and construct a knowledge graph; The system retrieves the user's target question, queries the target data and knowledge graph, and obtains the target output result.
[0007] Optionally, the data format and target terminology of the multi-source heterogeneous data are standardized to obtain standard data, including: The data in different data formats in the multi-source heterogeneous data are converted into standardized data. When a target term is identified in the standardized data, the target term in the standardized data is mapped to a standard term according to the terminology mapping relationship in the preset domain dictionary, thus obtaining the standard data.
[0008] Optionally, duplicate entity identification is performed on the standard data to remove duplicate entities and obtain target data, including: The standard data is processed using natural language processing to generate multiple semantic vectors; The similarity of the multiple semantic vectors is calculated. When the similarity score is greater than a preset threshold, the two semantic vectors corresponding to the similarity score are determined to be the same entity. According to the preset coordination rules, one of the entities in the same entity is retained to obtain the target data.
[0009] Optionally, based on the multi-source heterogeneous data, entity relationships are obtained from the multi-source heterogeneous data, and a knowledge graph is constructed, including: Based on the multi-source heterogeneous data, extract the target entity from the multi-source heterogeneous data; The target entities are classified and aligned, and multiple triples are obtained based on the entity relationships between the target entities. Based on the triples, a knowledge graph is obtained.
[0010] Optionally, the user's target question is obtained, and the target data and knowledge graph are queried to obtain the target output results, including: The target question for obtaining user input; Based on the target question, a search is performed in the target data to obtain search results; Based on the search results, knowledge enhancement is performed using the knowledge graph to obtain the target output result.
[0011] Optionally, based on the target question, a search is performed in the target data to obtain search results, including: Based on the target question, a structured retrieval is performed in the target data using keyword queries to obtain the first retrieval result; The target question is vectorized, and semantic retrieval is performed on the target data to obtain a second retrieval result; The first and second search results are merged and deduplicated, and the results are fused according to a preset weight formula to obtain the search results.
[0012] Optionally, based on the search results, knowledge enhancement is performed using the knowledge graph to obtain the target output results, including: By performing extended queries on the knowledge graph using preset query statements, relevant entities can be obtained; The relevance of each relevant entity is calculated based on a preset weighting mechanism to obtain the target output result.
[0013] The present invention also provides a retrieval and processing device based on multi-source knowledge base data, comprising: The acquisition module is used to acquire heterogeneous data from multiple sources. The processing module is used to standardize the data format and target terms of the multi-source heterogeneous data to obtain standard data; to identify duplicate entities in the standard data and remove duplicate entities to obtain target data; to obtain entity relationships in the multi-source heterogeneous data and construct a knowledge graph based on the multi-source heterogeneous data; and to obtain the user's target question, query it in the target data and the knowledge graph to obtain the target output result.
[0014] The present invention also provides a computing device, comprising: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above.
[0015] The present invention also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above.
[0016] The above-described solution of the present invention has at least the following beneficial effects: The above-described solution of the present invention acquires multi-source heterogeneous data; standardizes the data format and target terms of the multi-source heterogeneous data to obtain standard data; identifies duplicate entities in the standard data and removes duplicate entities to obtain target data; obtains entity relationships in the multi-source heterogeneous data based on the multi-source heterogeneous data and constructs a knowledge graph; obtains the user's target question, queries it in the target data and the knowledge graph, and obtains the target output result. This allows for unified management of multi-source data, improves the accuracy of cross-database data retrieval through the knowledge graph, and enhances the real-time response capability of the data. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the retrieval and processing method based on multi-source knowledge base data according to an embodiment of the present invention. Figure 2 This is a structural diagram of a retrieval and processing device based on multi-source knowledge base data according to an embodiment of the present invention. Detailed Implementation
[0018] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.
[0019] like Figure 1 As shown, embodiments of the present invention propose a retrieval and processing method based on multi-source knowledge base data, including: Step 11: Obtain multi-source heterogeneous data; Step 12: Standardize the data format and target terminology of the multi-source heterogeneous data to obtain standard data; Step 13: Perform duplicate entity identification on the standard data, remove duplicate entities from the standard data, and obtain the target data; Step 14: Based on the multi-source heterogeneous data, obtain the entity relationships in the multi-source heterogeneous data and construct a knowledge graph; Step 15: Obtain the user's target question, query it in the target data and knowledge graph, and obtain the target output result.
[0020] In this embodiment, a complete end-to-end implementation process is formed, from data source collection, heterogeneous data cleaning, entity semantic alignment, knowledge graph construction, real-time updates to AI semantic retrieval.
[0021] This solution employs a five-layer distributed architecture (data source layer, processing layer, knowledge layer, service layer, and application layer) to uniformly govern data from multiple databases. It improves cross-database retrieval accuracy (recall rate >95%) based on knowledge graph-based association reasoning. A real-time stream processing engine enables second-level data updates (latency <2 seconds).
[0022] In an optional embodiment of the present invention, step 12, standardizing the data format and target terminology of the multi-source heterogeneous data to obtain standard data, may include: Step 121: Convert the data of different data formats in the multi-source heterogeneous data to unify the data of different data formats into standardized data; Here, a multi-threaded parsing module processes data in different formats. For Excel (spreadsheet) files, a custom event-driven parser reads data line by line. To ensure I / O concurrency and memory stability, a thread parallelism of twice the number of CPU (Central Processing Unit) cores is used. For JSON (a lightweight data exchange format), a streaming block parsing strategy is used, reading 4KB blocks of data at a time, with parsing speed controlled at >50MB / s. For non-UTF-8 (a variable-length character encoding) encoded files, an automatic transcoding process is performed to unify them to UTF-8 encoding. Abnormal data rows are isolated and analyzed in a separate queue without blocking the main process.
[0023] Step 122: When a target term is identified in the standardized data, the target term in the standardized data is mapped to a standard term according to the term comparison relationship in the preset domain dictionary, thus obtaining the standard data.
[0024] Here, a pre-loaded domain dictionary is invoked, with an update cycle of 10 minutes. When a mappable target term appears, the domain dictionary is invoked to map the target term to the standard term recorded in the dictionary. For example, when the term "AI" appears, the system automatically maps it to the standard term "artificial intelligence" and updates historical data synchronously. A dynamic hot-update mechanism is employed, allowing new rules to be loaded without restarting the service. After multiple tests, this mechanism improved the semantic consistency rate of data from the original 70% to 98%.
[0025] In an optional embodiment of the present invention, step 13, which involves identifying duplicate entities in the standard data to remove duplicate entities and obtain the target data, may include: Step 131: Perform natural language processing on the standard data to generate multiple semantic vectors; Here, a pre-defined natural language processing (NLP) model is used to process standard data and generate semantic vectors. The NLP model has a hidden layer dimension of 768 dimensions and a maximum input length of 512 characters; the output is a unit vector, which is normalized and used for subsequent similarity comparison; the processing time for a single sample is less than or equal to 120ms.
[0026] The natural language processing model uses 12 encoder layers, each containing 12 self-attention heads. During training, a large-scale unlabeled corpus is used to predict relationships and semantic information between sentences.
[0027] Specifically, the architecture of the natural language processing model includes: Embedding layer: Converts input text into vector representation, including word embedding, position embedding, and segment embedding; Multilayer encoder: Each layer contains a multi-head self-attention mechanism, a feedforward neural network, residual connections, and layer normalization; Task-specific output layer: An output layer added based on the requirements of downstream tasks.
[0028] Step 132: Calculate the similarity of the multiple semantic vectors. When the similarity score is greater than a preset threshold, determine the two semantic vectors corresponding to the similarity score as the same entity. Here, a hybrid similarity calculation method is used to calculate the simple matching coefficient and the word frequency-inverse document frequency weighted average of the two semantic vectors respectively.
[0029] Let the two semantic vectors be A and B, then the simple matching coefficient J is calculated as follows:
[0030] The formula for calculating the term frequency-inverse document frequency weighted average is as follows:
[0031] Where t represents the target word and d represents the semantic vector. Indicate word frequency, Indicates inverse document frequency.
[0032] The simple matching coefficient calculates the proportion of common characters between the two strings, while the term frequency-inverse document frequency weighting reflects the discriminative power of high-frequency words. Finally, the two scores are multiplied by their respective weight coefficients and summed to obtain the final similarity score. Preferably, the weight coefficient for the simple matching coefficient is 0.6, and the weight coefficient for the term frequency-inverse document frequency weighting is 0.4. When the final similarity score is greater than or equal to 0.85, the two semantic vectors are considered to be the same entity.
[0033] Step 133: According to the preset coordination rules, retain one of the entities in the same entity to obtain the target data.
[0034] When the similarity score is greater than 0.85 after similarity calculation, and two entities are determined to be the same entity (e.g., the same talent or organization), a coordination mechanism is initiated for distributed consensus decision-making. A three-node voting system participates in the decision-making process simultaneously. If the votes are consistent, the latest timestamp is used; if the votes are inconsistent, a manual review queue is triggered. Time synchronization accuracy is controlled within ±50ms. This coordination mechanism retains one entity among the identical entities.
[0035] In an optional embodiment of the present invention, step 14, which involves obtaining entity relationships from the multi-source heterogeneous data and constructing a knowledge graph, may include: Step 141: Extract the target entity from the multi-source heterogeneous data based on the multi-source heterogeneous data; Step 142: Perform entity classification and alignment on the target entities, and obtain multiple triples based on the entity relationships between the target entities; Step 143: Obtain the knowledge graph based on the triples.
[0036] In this embodiment, a pre-trained neural network model extracts target entities from multi-source heterogeneous data. These target entities can be names of people, organizations, or domain terms. The entities and their labels are then output. The target entities are classified and aligned to obtain multiple triples. A knowledge graph is then generated based on these triples. For example, the triples could be: Zhang San, domain, new energy. The pre-trained neural network model includes an input layer, a feature layer, an attention layer, and an output layer. The input layer has 32-dimensional character embeddings and 300-dimensional word vectors; the feature layer has 256 units with a dropout rate of 0.25. The attention layer uses a self-attention mechanism to calculate context weights. The output layer uses a CRF state transition matrix, including a label set of 8 categories, such as names of people, organizations, and places.
[0037] The training process employs a transfer learning workflow, using general Chinese corpus in the pre-training stage; policy and institutional domain samples are loaded in the domain fine-tuning stage; and high-frequency error samples are manually verified in the calibration stage; the final model achieves an F1 score of 0.924.
[0038] In an optional embodiment of the present invention, step 14 further includes: updating the target data and knowledge graph when the multi-source heterogeneous data is updated.
[0039] In this embodiment, an incremental pipeline combining change capture and streaming processing engines is used to achieve low-latency data updates. In the change capture layer, database change events (such as inserts, updates, and deletes) are monitored in real time, and data changes are sent to the messaging system as messages. This supports scenarios such as data synchronization, cache updates, and microservice decoupling, with parsing granularity down to the row level. The Kafka partitioning strategy uses table name hashing to ensure the sequential nature of data for the same entity, with a latency of ≤200ms.
[0040] In the stream processing engine, the checkpoint interval is 10 seconds; the rolling window aggregation period is 10 seconds; the fault tolerance mechanism is: the first failure is immediately retried, and subsequent failures are exponentially backed up by 500ms × 2ⁿ, with a maximum of 5 retries; the parallelism is set to 8, and the processing throughput is >50,000 records / second. The system guarantees no duplicate commits through Exactly-Once semantics, and after 7 consecutive days of testing, data consistency is 100%.
[0041] Once the processing layer detects a data change, the system triggers the graph synchronization process. First, entity change identification is performed, including: maintaining a buffer pool of the 100,000 most recently changed primary keys in memory; employing an LRU (Least Recently Used) eviction policy to ensure that active entities are updated first; and comparing node hash values to determine if an update is needed. Then, batch write optimization is performed, triggering a batch write every 1000 cumulative changes; the write process uses transaction commit, with a single batch latency of ≤800ms; if a write fails, the system rolls back and retryes; actual update latency is ≤2 seconds.
[0042] In an optional embodiment of the present invention, step 15 involves obtaining the user's target question, querying the target data and knowledge graph to obtain the target output result, including: Step 151: Obtain the target question input by the user; Step 152: Based on the target question, perform a search in the target data to obtain search results; Step 153: Based on the search results, perform knowledge enhancement through the knowledge graph to obtain the target output result.
[0043] Specifically, in step 152, based on the target question, a search is performed in the target data to obtain search results, which may include: Step 1521: Based on the target question, perform a structured search in the target data using keyword queries to obtain the first search result; Here, an inverted index is constructed to support multi-condition Boolean queries, such as "Research Field: Artificial Intelligence AND Nationality: US"; fuzzy matching uses the IK tokenizer with a segmentation granularity of 4–12 characters; the index update cycle is 30 seconds, and it supports tiered storage for hot and cold queries. Experiments show that the average response time for a single query is less than or equal to 180ms.
[0044] Step 1522: Vectorize the target question and perform semantic retrieval on the target data to obtain a second retrieval result; Here, the input text is segmented into words, then natural language processing is performed to generate vectors, and finally, the vectors are indexed and retrieved in a vector database. A nearest neighbor graph index is used, with parameters set to M=16 and efSearch=80, achieving an optimal balance between retrieval recall and latency. In this embodiment, the vector generation time is less than or equal to 120ms.
[0045] Step 1523: Merge and deduplicate the first and second search results, and fuse the results according to the preset weight formula to obtain the search results.
[0046] Here, the top 50 results from the structured search and the top 50 results from the semantic search are merged and deduplicated, then comprehensively ranked using a weighted formula: Comprehensive Score = 0.7 × Keyword Matching Degree + 0.3 × Time Decay Coefficient. The keyword matching degree is the matching score from the structured search stage, i.e., the relevance score between the document and the user-input keywords. The time decay coefficient reflects the timeliness of the data, determined by a decay function. Where age is the number of years since the data was created, and λ is the decay coefficient (usually set to 0.2). By setting the time decay coefficient, the weight of older data can be reduced, and the latest policies or talent achievements can be displayed first.
[0047] Specifically, in step 153, based on the retrieval results, knowledge enhancement is performed using the knowledge graph to obtain the target output result, which may include: Step 1531: Perform an extended query on the knowledge graph using a preset query statement to obtain relevant entities; Here, each item in the search results output contains at least one identifier corresponding to an entity node in the knowledge graph, for example: A unique identifier for a talent entity (e.g., person_id); A unique identifier for an organization (e.g., org_id); A unique identifier for a policy entity (e.g., policy_id); Domain category tags, etc.; Each triplet node in the knowledge graph is identified by the aforementioned identification information, and the entity identification field in the retrieval results can be directly mapped to the corresponding node in the knowledge graph.
[0048] Based on the corresponding nodes in the knowledge graph corresponding to the above search results, a two-hop relation expansion query is performed to enhance the semantic richness of the search results through the knowledge graph expansion mechanism. That is, the search results are used as the starting point for semantic association expansion. The nodes corresponding to the search results are used as the "starting node" for the graph query. Then, according to the predefined relation structure in the graph, relevant information is queried in a one-hop or multi-hop manner. For example, the query statement can be: MATCH (p:Policy)-[:AFFECTS]->(d:Domain)<-[:RESEARCHES]-(t:Talent) WHERE p.id IN $policy_list RETURN t.name, d.name, p.title The above query means: find the talents in research fields affected by the searched policy, that is, "which talents are engaged in research in research fields affected by a certain policy". Here, (p:Policy) represents a policy node; (d:Domain) represents a research field node; (t:Talent) represents a talent node; [:AFFECTS] represents the relationship of "policy-affected fields"; and [:RESEARCHES] represents the relationship of "talents engaged in research fields". The above query takes an average of 0.9 seconds, revealing an implicit association between policy, field, and talent.
[0049] Step 1532: Calculate the relevance of each related entity according to the preset weights to obtain the target output result. Calculate the relevance of related entities in the extended query of the above knowledge graph, and filter the entities to be output based on the relevance.
[0050] In the above query process, relevance = semantic matching degree × authority value × timeliness factor, where the authority value is the dynamic weight of each data node. Preferably, the authority value is set to 1.0 for institutional entities, 0.8 for academic journal entities, and 0.2 for self-media entities. A timeliness decay mechanism is set: for entities with a data age of more than 1 year, the timeliness factor is 0.5; for entities with a data age of more than 3 years, the timeliness factor is 0.1. Semantic matching degree is the semantic similarity between the initial search results and related nodes in the graph, used to filter semantically similar candidate objects in the expanded nodes and avoid introducing irrelevant relationships. Using the above method, the relevance of related entities can be obtained, and related entities with a relevance greater than the preset value are output.
[0051] The method described above in this invention often only covers the keywords or semantically similar content entered by the user, but does not cover implicitly related content. Knowledge graphs, however, can fill in these gaps, for example: When a user searches for "a certain policy," the knowledge graph will automatically identify the research fields and related talents affected by that policy; when a user searches for "a certain field," it can provide supplementary information such as related policies, experts, and institutions; when a user searches for "a certain institution," the knowledge graph can expand to include the projects undertaken by that institution and the personnel involved. The knowledge graph expands the "point-like information" of the initial search results into "related network information," thus forming the target output results.
[0052] The target output consists of the following parts: Preliminary search results (direct hits): Search results obtained by fusing structured and semantic search. Graph expansion results (semantic related items): Highly relevant entities obtained by mapping search results to graph nodes and performing one-hop or two-hop expansion.
[0053] The specific implementation process of the above embodiments of the present invention may include: Step 1, Data Access Phase: After a user uploads an Excel / JSON file, the system automatically performs encoding, delimiter, and header validation; for abnormal files, it returns error code E1002. For API data access, token verification is performed, with a response time of ≤50ms; The request exceeded the limit and returned status code 429, and was recorded in the blacklist.
[0054] The second step, the processing stage: Real-time path: Kafka (distributed stream processing platform), stream processing engine, entity alignment, graph update; Batch processing tasks are executed daily at 02:00 to remove duplicates, generate statistics and reports; Results are synchronized, and the retrieval service is triggered to update the vector index upon completion of the task.
[0055] Step 3, Service Delivery Phase: When a user initiates a search request, the system parses the input text; Call the model to encode and generate vectors; Retrieve the recall candidate set; Extended results of graph reasoning; The final output result after fusion.
[0056] After multiple rounds of testing, the system had an average retrieval latency of 2.1 seconds and a recall rate of 95.3%.
[0057] The embodiments of the present invention acquire multi-source heterogeneous data; standardize the data format and target terms of the multi-source heterogeneous data to obtain standard data; identify duplicate entities in the standard data and remove duplicate entities to obtain target data; acquire entity relationships in the multi-source heterogeneous data and construct a knowledge graph based on the multi-source heterogeneous data; acquire the user's target question, query it in the target data and knowledge graph, and obtain the target output result. It can achieve efficient batch processing at the bottom layer, dynamic graph fusion at the knowledge layer, and high-precision semantic retrieval at the top layer. It features strong stability, high scalability, and good reproducibility in terms of data fusion accuracy, retrieval recall, and real-time response performance.
[0058] like Figure 2 As shown, embodiments of the present invention also provide a retrieval processing device 20 based on multi-source knowledge base data, comprising: Module 21 is used to acquire multi-source heterogeneous data; Processing module 22 is used to standardize the data format and target terms of the multi-source heterogeneous data to obtain standard data; to identify duplicate entities in the standard data and remove duplicate entities to obtain target data; to obtain entity relationships in the multi-source heterogeneous data and construct a knowledge graph based on the multi-source heterogeneous data; to obtain the user's target question, query it in the target data and knowledge graph, and obtain the target output result.
[0059] Optionally, the data format and target terminology of the multi-source heterogeneous data are standardized to obtain standard data, including: The data in different data formats in the multi-source heterogeneous data are converted into standardized data. When a target term is identified in the standardized data, the target term in the standardized data is mapped to a standard term according to the terminology mapping relationship in the preset domain dictionary, thus obtaining the standard data.
[0060] Optionally, duplicate entity identification is performed on the standard data to remove duplicate entities and obtain target data, including: The standard data is processed using natural language processing to generate multiple semantic vectors; The similarity of the multiple semantic vectors is calculated. When the similarity score is greater than a preset threshold, the two semantic vectors corresponding to the similarity score are determined to be the same entity. According to the preset coordination rules, one of the entities in the same entity is retained to obtain the target data.
[0061] Optionally, based on the multi-source heterogeneous data, entity relationships are obtained from the multi-source heterogeneous data, and a knowledge graph is constructed, including: Based on the multi-source heterogeneous data, extract the target entity from the multi-source heterogeneous data; The target entities are classified and aligned, and multiple triples are obtained based on the entity relationships between the target entities. Based on the triples, a knowledge graph is obtained.
[0062] Optionally, the user's target question is obtained, and the target data and knowledge graph are queried to obtain the target output results, including: The target question for obtaining user input; Based on the target question, a search is performed in the target data to obtain search results; Based on the search results, knowledge enhancement is performed using the knowledge graph to obtain the target output result.
[0063] Optionally, based on the target question, a search is performed in the target data to obtain search results, including: Based on the target question, a structured retrieval is performed in the target data using keyword queries to obtain the first retrieval result; The target question is vectorized, and semantic retrieval is performed on the target data to obtain a second retrieval result; The first and second search results are merged and deduplicated, and the results are fused according to a preset weight formula to obtain the search results.
[0064] Optionally, based on the search results, knowledge enhancement is performed using the knowledge graph to obtain the target output results, including: By performing extended queries on the knowledge graph using preset query statements, relevant entities can be obtained; The relevance of each relevant entity is calculated based on a preset weighting mechanism to obtain the target output result.
[0065] It should be noted that this device is the same as the method described above. All implementations in the above method embodiments are applicable to the embodiments of this device and can achieve the same technical effect.
[0066] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method as described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0067] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0068] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0069] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0070] In the embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0071] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0072] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0073] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0074] Furthermore, it should be noted that in the apparatus and method of the present invention, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered equivalent solutions of the present invention. Moreover, the steps performing the above series of processes can naturally be executed in the order described, but are not necessarily required to be executed in chronological order; some steps can be executed in parallel or independently of each other. Those skilled in the art will understand that all or any step or component of the method and apparatus of the present invention can be implemented in any computing device (including processors, storage media, etc.) or network of computing devices, in hardware, firmware, software, or a combination thereof. This is something that those skilled in the art can achieve by using their basic programming skills after reading the description of the present invention.
[0075] Therefore, the object of the present invention can also be achieved by running a program or a set of programs on any computing device. The computing device can be a known general-purpose device. Therefore, the object of the present invention can also be achieved simply by providing a program product containing program code implementing the method or apparatus. That is, such a program product also constitutes the present invention, and the storage medium storing such a program product also constitutes the present invention. Obviously, the storage medium can be any known storage medium or any storage medium developed in the future. It should also be noted that in the apparatus and method of the present invention, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered equivalent to the present invention. Furthermore, the steps performing the above series of processes can naturally be performed in the order described, but are not necessarily required to be performed in chronological order. Some steps can be performed in parallel or independently of each other.
[0076] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A retrieval and processing method based on multi-source knowledge base data, characterized in that, include: Acquire multi-source heterogeneous data; The data format and target terminology of the multi-source heterogeneous data are standardized to obtain standard data; The standard data is subjected to duplicate entity identification to remove duplicate entities from the standard data, thereby obtaining the target data; Based on the multi-source heterogeneous data, obtain the entity relationships in the multi-source heterogeneous data and construct a knowledge graph; The system retrieves the user's target question, queries the target data and knowledge graph, and obtains the target output result.
2. The retrieval and processing method based on multi-source knowledge base data according to claim 1, characterized in that, The data format and target terminology of the multi-source heterogeneous data are standardized to obtain standard data, including: The data in different data formats in the multi-source heterogeneous data are converted into standardized data. When a target term is identified in the standardized data, the target term in the standardized data is mapped to a standard term according to the terminology mapping relationship in the preset domain dictionary, thus obtaining the standard data.
3. The retrieval and processing method based on multi-source knowledge base data according to claim 1, characterized in that, The standard data is subjected to duplicate entity identification to remove duplicate entities, resulting in target data, including: The standard data is processed using natural language processing to generate multiple semantic vectors; The similarity of the multiple semantic vectors is calculated. When the similarity score is greater than a preset threshold, the two semantic vectors corresponding to the similarity score are determined to be the same entity. According to the preset coordination rules, one of the entities in the same entity is retained to obtain the target data.
4. The retrieval and processing method based on multi-source knowledge base data according to claim 1, characterized in that, Based on the multi-source heterogeneous data, entity relationships are obtained from the multi-source heterogeneous data, and a knowledge graph is constructed, including: Based on the multi-source heterogeneous data, extract the target entity from the multi-source heterogeneous data; The target entities are classified and aligned, and multiple triples are obtained based on the entity relationships between the target entities. Based on the triples, a knowledge graph is obtained.
5. The retrieval and processing method based on multi-source knowledge base data according to claim 1, characterized in that, Obtain the user's target question, query it in the target data and knowledge graph, and obtain the target output results, including: The target question for obtaining user input; Based on the target question, a search is performed in the target data to obtain search results; Based on the search results, knowledge enhancement is performed using the knowledge graph to obtain the target output result.
6. The retrieval and processing method based on multi-source knowledge base data according to claim 5, characterized in that, Based on the target question, a search is performed in the target data to obtain search results, including: Based on the target question, a structured retrieval is performed in the target data using keyword queries to obtain the first retrieval result; The target question is vectorized, and semantic retrieval is performed on the target data to obtain a second retrieval result; The first and second search results are merged and deduplicated, and the results are fused according to a preset weight formula to obtain the search results.
7. The retrieval and processing method based on multi-source knowledge base data according to claim 5, characterized in that, Based on the search results, knowledge enhancement is performed using the knowledge graph to obtain the target output results, including: By performing extended queries on the knowledge graph using preset query statements, relevant entities can be obtained; The relevance of each relevant entity is calculated based on a preset weighting mechanism to obtain the target output result.
8. A retrieval and processing device based on multi-source knowledge base data, characterized in that, include: The acquisition module is used to acquire heterogeneous data from multiple sources. The processing module is used to standardize the data format and target terms of the multi-source heterogeneous data to obtain standard data; The standard data is subjected to duplicate entity identification to remove duplicate entities and obtain target data; based on the multi-source heterogeneous data, the entity relationships in the multi-source heterogeneous data are obtained and a knowledge graph is constructed; the user's target question is obtained and queried in the target data and the knowledge graph to obtain the target output result.
9. A computing device, characterized in that, include: A processor, a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A storage instruction that, when executed on a computer, causes the computer to perform the method as described in any one of claims 1 to 7.