Multi-source knowledge processing and querying method and device, equipment and medium
Through distributed data acquisition, standardized processing, multi-dimensional knowledge graph construction and semantic index establishment, the problems of multi-source heterogeneous data integration and knowledge association mining are solved, and efficient data processing and precise query are achieved.
Patent Information
- Application Number
- CN202510278159.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is difficult to effectively integrate multi-source heterogeneous data and mine complex knowledge associations, resulting in inefficient data processing and insufficient query accuracy.
Multi-source heterogeneous data is obtained through the distributed data acquisition module, and cleaned and standardized processing is performed to generate a standardized database. Define the core concepts and association relationships in the field, extract knowledge elements through the feature extraction module, build a multi-dimensional knowledge graph, and establish a semantic index on it. Get the query intent, extract the query constraints, and perform association reasoning in the multi-dimensional knowledge graph to generate the query results.
By building standardized databases and multi-dimensional knowledge graphs, ensure the unified data format and improve data consistency and relevance. 通过语义索引和关联推理,提高查询的精准度和效率,避免信息冗余问题,实现对复杂知识的高效提取和精准查询。
Smart Images

Figure CN120197681A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and in particular, to a multi-source knowledge processing and query method, apparatus, device, and storage medium. Background Art
[0002] With the continuous increase in the amount of data in multiple fields, how to effectively integrate multi-source heterogeneous data and solve problems such as knowledge association mining, query efficiency, and data credibility in related fields has become an urgent challenge. Currently, fields such as Buddhism, healthcare, and finance are facing problems of scattered data sources, inconsistent formats, low query efficiency, and data credibility, which affect the data management and application efficiency in related fields. In these fields, various data sources are usually distributed in documents, databases, and systems with different formats, and there is a lack of an effective integration mechanism, resulting in the inability to efficiently converge information. At the same time, knowledge association mining technology also has difficulties in the face of complex correlations and multi-dimensional data, and cannot comprehensively reflect the deep relationships between knowledge elements, seriously affecting the progress of research and practical applications.
[0003] In the field of cultural research, such as the field of Buddhist studies, relevant materials are scattered in ancient books, electronic documents, and temple materials, lacking a unified management framework, resulting in low efficiency of knowledge integration and transmission. Buddhist knowledge involves complex sectarian doctrines, interpersonal relationships, and historical inheritances, but existing technical means are difficult to effectively identify and mine these complex association information. Especially when facing Buddhist scriptures and literature across time and regions, query systems often cannot fully capture the user's query intent, affecting the efficiency of Buddhist research and application.
[0004] In the field of healthcare, medical records, clinical data, and research results are distributed in different hospitals, databases, and medical platforms, lacking an effective integration mechanism, resulting in the inability to efficiently share and apply medical information and knowledge. At the same time, knowledge association mining in the field of healthcare involves complex relationship networks among diseases, treatment plans, clinical decisions, and medical literature. Existing technologies often cannot efficiently mine these deep associations, affecting the accuracy of clinical decisions and research efficiency. Especially in the fields of disease prevention, diagnosis, and treatment plan recommendation, how to effectively integrate multi-source data and provide accurate knowledge support is an urgent technical problem to be solved.
[0005] The financial field involves a large amount of market data, transaction records, and customer information. These data sources are distributed across multiple financial institutions and markets, with diverse formats and redundant information. Existing financial query systems mainly rely on keyword matching and are difficult to understand users' complex query intentions. Especially when dealing with queries involving multiple related concepts, the query results are often inaccurate and cannot meet users' needs. In terms of knowledge association mining, the financial field involves complex association networks, including multi-dimensional data relationships such as investment portfolios, market fluctuations, and risk control. However, existing technologies are difficult to comprehensively mine these relationships, resulting in low knowledge transfer efficiency and affecting the accuracy of market decision-making and investment analysis.
[0006] In addition, the issue of data credibility in each field is also a key factor restricting development. In the field of Buddhism, the risk of tampering with different versions of Buddhist scriptures and digital materials affects the reliability of research; in the field of healthcare, medical record data and research results are often threatened by tampering and forgery, affecting the accuracy of clinical decision-making and academic research; in the financial field, the authenticity of market data and transaction records is particularly prominent during financial crises and market fluctuations, affecting investment decisions and risk management. Therefore, how to ensure the authority and credibility of data has become a challenge that needs to be solved urgently in each field. Summary of the Invention
[0007] The main objective of the present invention is to provide a multi-source knowledge processing and query method, device, equipment, and storage medium, aiming to solve the technical problems that existing technologies cannot effectively integrate multi-source heterogeneous data and mine complex knowledge associations, resulting in inefficient data processing and insufficient query accuracy.
[0008] To achieve the above objective, the present invention provides a multi-source knowledge processing and query method, including:
[0009] Obtain multi-source heterogeneous data through a distributed data acquisition module, perform cleaning and standardization processing on the multi-source heterogeneous data, and generate a standardized database based on the processed multi-source heterogeneous data;
[0010] Define domain core concepts and association relationships, and extract knowledge elements from the multi-source heterogeneous data through a feature extraction module;
[0011] Construct a multi-dimensional knowledge graph based on the core concepts, association relationships, and knowledge elements;
[0012] Establish a semantic index based on the multi-dimensional knowledge graph, perform semantic label annotation on the elements in the standardized database according to the semantic index, and establish a two-way mapping relationship between the semantic labels of the elements and the nodes of the multi-dimensional knowledge graph;
[0013] Obtain a query intention, and extract query constraint conditions from the query intention;
[0014] Perform associated inference in the multi-dimensional knowledge graph according to the query constraint conditions and the two-way mapping relationship to generate a query result.
[0015] Furthermore, to achieve the above object, the present invention provides a multi-source knowledge processing and query device, including:
[0016] A distributed data acquisition module, configured to acquire multi-source heterogeneous data through the distributed data acquisition module, perform cleaning processing and standardization processing on the multi-source heterogeneous data, and generate a standardized database based on the processed multi-source heterogeneous data;
[0017] A feature extraction module, configured to define domain core concepts and association relationships, and extract knowledge elements from the multi-source heterogeneous data through the feature extraction module;
[0018] A knowledge graph construction module, configured to construct a multi-dimensional knowledge graph based on the core concepts, association relationships, and knowledge elements;
[0019] A semantic indexing module, configured to establish a semantic index based on the multi-dimensional knowledge graph, perform semantic label annotation on elements in the standardized database according to the semantic index, and establish a two-way mapping relationship between the semantic labels of the elements and nodes of the multi-dimensional knowledge graph;
[0020] A query parsing module, configured to obtain a query intention and extract query constraint conditions from the query intention;
[0021] An inference engine module, configured to perform associated inference in the multi-dimensional knowledge graph according to the query constraint conditions and the two-way mapping relationship to generate a query result.
[0022] Furthermore, to achieve the above object, the present invention further provides a computer device, where the computer device includes a memory, a processor, and a multi-source knowledge processing and query program stored on the memory and executable on the processor. When the multi-source knowledge processing and query program is executed by the processor, the steps of the multi-source knowledge processing and query method as described above are implemented.
[0023] Furthermore, to achieve the above object, the present invention further provides a computer-readable storage medium, where a multi-source knowledge processing and query program is stored on the storage medium. When the multi-source knowledge processing and query program is executed by a processor, the steps of the multi-source knowledge processing and query method as described above are implemented.
[0024] Beneficial effects: The present invention relates to the technical field of data analysis and can be applied to business scenarios such as medical health, fintech, and cultural research. It discloses a multi-source knowledge processing and query method, including: acquiring multi-source heterogeneous data and performing cleaning and standardization processing to generate a standardized database; defining core concepts and association relationships, and extracting knowledge elements through feature extraction; constructing a multi-dimensional knowledge graph based on the core concepts, association relationships, and knowledge elements; establishing a semantic index on the multi-dimensional knowledge graph, and performing semantic label annotation on the elements in the standardized database to form a bidirectional mapping relationship; acquiring a query intention, and extracting query constraint conditions from the query intention; based on the query constraint conditions and the bidirectional mapping relationship, performing association reasoning in the multi-dimensional knowledge graph to generate a query result. By constructing a standardized database for multi-source heterogeneous data, the present invention ensures the unity of data format and improves data consistency; through the construction of a multi-dimensional knowledge graph and the establishment of a semantic index, it enhances the relevance and structured expression of data, making the relationships between knowledge elements clearly traceable; through association reasoning based on query constraint conditions, it improves the accuracy and efficiency of queries, avoids the information redundancy problem caused by traditional keyword matching, and realizes the efficient extraction and accurate query of complex knowledge. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0026] Figure 1 is a schematic diagram of an application environment of the multi-source knowledge processing and query method in an embodiment of the present invention;
[0027] Figure 2 is a schematic flowchart of an embodiment of the multi-source knowledge processing and query method of the present invention;
[0028] Figure 3 is a schematic diagram of functional modules of a preferred embodiment of the multi-source knowledge processing and query device of the present invention;
[0029] Figure 4 is a schematic diagram of the structure of a computer device in an embodiment of the present invention;
[0030] Figure 5 is another schematic diagram of the structure of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0032] The multi-source knowledge processing and query method provided by the embodiments of the present invention can be applied in such as Figure 1In the application environment, the client communicates with the server through the network. The server can obtain multi-source heterogeneous data through the client, clean and standardize it to generate a standardized database; define core concepts and their associated relationships, and extract knowledge elements through feature extraction; construct a multi-dimensional knowledge graph based on the core concepts, associated relationships, and knowledge elements; establish a semantic index on the multi-dimensional knowledge graph, and perform semantic label annotation on the elements in the standardized database to form a two-way mapping relationship; obtain the query intent, and extract query constraint conditions from the query intent; based on the query constraint conditions and the two-way mapping relationship, perform association reasoning in the multi-dimensional knowledge graph to generate query results. By constructing a standardized database for multi-source heterogeneous data, the present invention ensures the uniformity of data formats and improves data consistency; through the construction of a multi-dimensional knowledge graph and the establishment of a semantic index, it enhances the relevance and structured expression of data, making the relationships between knowledge elements clearly traceable; through association reasoning based on query constraint conditions, it improves the accuracy and efficiency of queries, avoids the problem of information redundancy caused by traditional keyword matching, and realizes the efficient extraction and accurate query of complex knowledge. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0033] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the multi-source knowledge processing and query method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0034] As Figure 2 shown, the multi-source knowledge processing and query method proposed by the present invention includes the following steps:
[0035] S10, obtain multi-source heterogeneous data through a distributed data acquisition module, perform cleaning processing and standardization processing on the multi-source heterogeneous data, and generate a standardized database based on the processed multi-source heterogeneous data;
[0036] In this embodiment, the core of the distributed data acquisition module lies in constructing an efficient and stable heterogeneous data acquisition mechanism, enabling content from different data sources to be uniformly incorporated into the system for subsequent processing. Data acquisition mainly involves the identification of data sources, the formulation of acquisition strategies, the selection of data acquisition methods, and the establishment of data synchronization mechanisms. In the data source identification stage, the data categories to be acquired are first determined, including data in different formats such as text, images, audio, and video. For different categories of data, appropriate acquisition methods are adopted. For example, for structured data, it can be directly obtained through API interfaces; for unstructured data such as electronic documents and web content, web crawler technology is used for parsing and extraction; for physical data such as paper documents and inscription images, high-resolution scanning devices are used for digitization, and optical character recognition (OCR) technology is combined for text transcription.
[0037] The formulation of the data acquisition strategy is based on the characteristics and update frequency of the data source, and a dynamic task scheduling mechanism is adopted to improve the acquisition efficiency and resource utilization rate. For data sources with high-frequency updates, an incremental acquisition mode can be adopted to regularly capture newly added data; for data sources with low-frequency updates, a full-scale synchronization mode is adopted to ensure data integrity. The selection of data acquisition methods needs to consider the network environment, data source access restrictions, and storage structure, supporting various methods such as HTTP requests, database connections, and file system access.
[0038] The establishment of the data synchronization mechanism needs to ensure the timeliness and consistency of the data. A change detection technology based on timestamps is adopted to perform version control on the newly acquired data to avoid duplicate storage and conflicts. Before data storage, the data from different sources is subjected to format conversion to make it conform to a unified storage standard for subsequent cleaning and standardization processing.
[0039] The cleaning process involves optimizing the quality of the acquired data, removing noise information, and improving the usability of the data. It mainly includes operations such as data deduplication, missing value filling, format standardization, and anomaly detection. For text data, deduplication uses a hash comparison algorithm to identify duplicate content and merge storage; for missing data, an interpolation method based on a statistical model is used for supplementation to keep the data complete; for data with non-standard formats, it is adjusted according to predefined standardization rules, such as unifying date formats and converting numerical units.
[0040] Standardization processing ensures data consistency, enabling it to be efficiently parsed and utilized in subsequent processing steps. The main operations include field mapping, data structure reorganization, semantic tag annotation, etc. Field mapping unifies the conversion of similar fields from different data sources by defining a unified data dictionary. For example, fields such as "Name" and "User Name" in different databases can be mapped to the unified "Name" field. Data structure reorganization involves optimizing the data storage format, such as converting from a nested JSON structure to a relational database structure to improve query performance. Semantic tag annotation assigns clear meanings to data through automated rule matching or manual intervention, enhancing data comprehensibility and retrieval efficiency.
[0041] The construction of the standardized database is carried out after the above data processing is completed. The core objectives are to achieve efficient storage, fast retrieval, and secure management. The database architecture adopts a distributed storage system to ensure data scalability and high-concurrency access capabilities. The database index mechanism is optimized using various methods such as B+ tree indexes and full-text indexes to improve query speed. Data access permission management realizes secure access control for different users through identity authentication and permission grading.
[0042] In the field of healthcare, it can be used for the unified integration of patient medical record data. The data formats and storage methods of different hospitals vary. After adopting the distributed data collection module, patient medical record data can be obtained from multiple medical institutions, and through standardization processing, a unified electronic health record can be generated to improve the usability of medical record data. In the data cleaning process, artificial intelligence technology can be used to automatically identify typos, data missing, and outliers in medical record entries to ensure data quality. In the standardized database, disease codes from different hospitals can be mapped, enabling cross-hospital query of patients' historical medical records, improving doctors' comprehensive understanding of patients' medical histories, and thus enhancing the accuracy of medical decision-making.
[0043] In the financial field, it can be used for the integration and processing of transaction data. The transaction data formats of different financial institutions are different. For example, bank transaction records, securities market data, fund transaction information, etc. all adopt different data storage formats. After adopting this solution, data can be collected from different institutions through the distributed data collection module, and through standardization processing, the transaction data can be converted into a unified format and stored in the standardized database. Through data cleaning, incorrect transaction records can be eliminated and missing data can be filled in to ensure the integrity and accuracy of financial data. With the support of the standardized database, the financial analysis system can retrieve and process market data more quickly, improving the real-time performance and accuracy of financial risk control models.
[0044] In the field of knowledge research, it can be used for the integration and analysis of Buddhist literature. Buddhist scriptures are distributed in different sources such as paper ancient books, electronic documents, and academic platforms, with different data formats and structures. After adopting the distributed data collection module, Buddhist literature from different sources can be integrated, and paper documents can be digitized through optical character recognition technology and stored in a standardized database. Through data cleaning technology, duplicate text fragments can be removed and missing words and sentences can be repaired, making the literature content more complete. The establishment of a standardized database enables researchers to efficiently query different versions of Buddhist scriptures, track the evolution process of Buddhist thoughts, and improve research efficiency.
[0045] By establishing a distributed data collection mechanism, the efficient acquisition of multi-source heterogeneous data is achieved, avoiding the problems of limited data acquisition channels and incompatible data formats in traditional methods. Data cleaning and standardization processing ensure the quality and consistency of data, enabling the data to be efficiently utilized and improving the readability and reliability of the data.
[0046] S20, define the core concepts and their relationships in the domain, and extract knowledge elements from the multi-source heterogeneous data through a feature extraction module;
[0047] In this embodiment, in a multi-source heterogeneous data environment, defining the core concepts and their relationships in the domain is the basis for constructing a high-quality knowledge system. The feature extraction module is responsible for extracting key knowledge elements from the data, making subsequent knowledge modeling and reasoning more accurate and reliable. The definition of core concepts first requires clarifying the data characteristics of the target domain. For example, in the medical field, core concepts may include diseases, symptoms, drugs, treatment methods, etc.; in the financial field, core concepts may involve market indicators, investment strategies, risk assessments, etc.; in the Buddhist field, core concepts involve doctrines, sects, scriptures, figures, etc.
[0048] When defining core concepts, a domain concept set can be constructed based on the knowledge of domain experts and existing literature materials and organized according to hierarchical relationships. The relationships between concepts include inheritance relationships, causal relationships, association strengths, etc. For example, in the medical field, as a disease concept, "diabetes" has a causal relationship with "insulin", and at the same time, there is an inheritance relationship between "diabetes" and "type 2 diabetes"; in the financial field, there is a causal relationship between "market volatility" and "interest rate adjustment", and there is a hierarchical attribution relationship between "stock market" and "securities exchange".
[0049] The core task of the feature extraction module is to extract structured knowledge elements from multi-source heterogeneous data. For text data, natural language processing techniques are adopted, combined with named entity recognition (NER) and relation extraction models, to extract entities, attributes, and relations from unstructured text. For example, in Buddhist literature, concepts such as "Chan Buddhism", "Huineng", and "sudden enlightenment" are identified, and relations such as "Chan Buddhism - Huineng (sect inheritance)" and "Huineng - sudden enlightenment (ideological proposition)" are established. For image data, image semantic segmentation technology is used to extract structured features from images. For example, lesion regions are identified in medical images, and the postures and clothing features of Buddha statues are extracted from Buddhist art works, and they are associated with text concepts through cross-modal alignment technology. For example, when analyzing Buddha statue images, "mudras" can be corresponded to "symbols of wisdom" to form cross-modal knowledge fusion of images and texts.
[0050] By establishing a core concept system and combining feature extraction techniques, knowledge elements can be efficiently extracted from multi-source heterogeneous data, and an accurate concept association network can be constructed. It has significant advantages in terms of the automation degree, accuracy, and multi-modal data fusion ability of knowledge extraction. The introduction of the feature extraction module enables different data types to complement each other, improves the integrity and structuring degree of knowledge, and lays a foundation for subsequent knowledge graph construction, semantic query, and intelligent reasoning.
[0051] S30, construct a multi-dimensional knowledge graph based on the core concepts, association relations, and knowledge elements;
[0052] In this embodiment, the core of constructing a multi-dimensional knowledge graph lies in organizing the core concepts, association relations, and knowledge elements into a structured knowledge network, so that complex knowledge systems can be expressed and stored in an intuitive and computable way. The construction of the knowledge graph involves four key links: concept organization, relation mapping, graph structure optimization, and multi-dimensional knowledge expression.
[0053] First, in the concept organization stage, a concept hierarchy is established based on the previously extracted core concepts. Concepts can be organized according to inheritance relationships and classification hierarchies. For example, in the medical field, diseases can be classified by category. For example, "diabetes" belongs to "endocrine diseases", and "Alzheimer's disease" belongs to "neurological diseases"; in the financial field, "stock market" belongs to "capital market", and "commercial bank" belongs to "financial institutions"; in the Buddhist field, "Chan Buddhism" belongs to "Buddhist sects", and "Huayan Sutra" belongs to "classic categories". The establishment of the concept hierarchy makes the knowledge system clearer and provides a basis for subsequent relation reasoning.
[0054] Secondly, in the relationship mapping stage, establish association relationships for each node in the knowledge graph, including causal relationships, logical relationships, spatial relationships, and temporal relationships, etc. For example, in a medical knowledge graph, establish a causal relationship between "high blood sugar" and "diabetes", and a treatment relationship between "diabetes" and "insulin"; in the financial field, establish an influencing relationship between "interest rate hike" and "stock market volatility"; in the field of Buddhism, establish a inheritance relationship between "Bodhidharma" and "Chan Buddhism", and a concept explanation relationship between "Prajna Paramita" and "Sunyata theory". The establishment of association relationships can be automatically extracted based on natural language processing technology, or defined through expert annotation or knowledge rules.
[0055] The key to graph structure optimization lies in ensuring the reasonable topological structure of the knowledge graph and avoiding the impact of over-connection or redundant relationships on query efficiency. Optimization methods include deduplication, node fusion, relationship weight calculation, etc. For example, if the same concept has multiple expressions in different data sources (such as "COVID-19", "novel coronavirus", "2019-nCoV"), it is necessary to uniformly map them to a standardized concept and merge the corresponding association relationships. In addition, in order to improve the accuracy of reasoning, relationship weight calculation can be determined through data statistics and machine learning methods. For example, in the medical field, the association strength between different treatment plans and a certain disease can be learned based on large-scale clinical data. In the financial field, the correlation degree between different economic policies and market fluctuations can be calculated through historical data modeling.
[0056] In terms of multi-dimensional knowledge representation, the knowledge graph needs to support multi-layer representation in the time dimension, space dimension, and logical dimension. The time dimension can be used to track the evolution of concepts or relationships, such as "the timeline of Buddhism's introduction into China" and "the impact of interest rate policies on the financial market", etc. The space dimension can reflect geographical distribution, such as "the incidence of a specific disease in different regions" and "the differences in the responses of different countries to the same economic policy", etc. The logical dimension is used to express the logical reasoning relationships between abstract concepts, such as "causal relationships" and "subordination relationships", etc. This multi-dimensional expression method enables the knowledge graph to not only store static knowledge but also dynamically depict the development and evolution process of knowledge.
[0057] In practical applications, the methods for constructing a multi-dimensional knowledge graph can be adjusted according to different requirements and data characteristics:
[0058] One of the implementation methods is the rule- and expert system-based approach. For fields with a large amount of structured data, such as medicine and finance, domain experts can predefined concept classification criteria, relationship types, and inference rules. For example, in the medical field, define the classification hierarchy of diseases, treatment relationships, drug interaction rules, etc., and store and query them through a knowledge base management system; in the financial field, define the association rules between macroeconomic indicators and market fluctuations, such as "GDP growth affects the rise of the stock market", and optimize the weights based on statistical data. This method is applicable to fields with relatively mature knowledge systems and can ensure the accuracy and authority of knowledge.
[0059] Another implementation method is the machine learning- and big data analytics-based approach. For fields with a large amount of unstructured data, such as news, social media, and scientific research literature, deep learning techniques can be used to automatically construct knowledge graphs. For example, use pre-trained models such as BERT for text relationship extraction and combine graph neural networks (GNNs) to optimize the structure of the knowledge graph. In the financial field, event entities (such as "the Federal Reserve", "interest rate cut") can be extracted from news texts and a financial event graph can be automatically constructed; in the medical field, the association relationships between different diseases can be learned from a large number of electronic medical records to form a data-driven medical knowledge graph. This method is applicable to scenarios with dynamically updated knowledge systems and can improve the automation degree of knowledge acquisition.
[0060] In addition, a multi-modal fusion method can be adopted to incorporate multi-source data such as text, images, and audio into the knowledge graph. For example, in Buddhist studies, the text content of Buddhist scriptures can be fused with images of Buddha statues, so that the knowledge graph not only contains scriptural explanations but also can display the corresponding visual content. In the medical field, the medical record text of patients can be associated with imaging data (such as MRI, CT) to form a cross-modal medical knowledge graph. This method can make full use of various data types and improve the comprehensiveness of knowledge expression.
[0061] By constructing a multi-dimensional knowledge graph, the ability to extract, organize, and express complex knowledge relationships from multi-source heterogeneous data is achieved. Compared with traditional database storage methods, the knowledge graph has stronger flexibility and scalability and can support complex queries and inferences. Through multi-dimensional expression methods, the knowledge graph can not only store static knowledge but also dynamically depict the evolution process of knowledge, providing more accurate knowledge support for intelligent queries, knowledge discovery, and automatic inferences. In addition, the construction method of the knowledge graph can adapt to the data characteristics of different fields, realize cross-field knowledge fusion, and further improve the relevance and utilization efficiency of information.
[0062] S40. Establish a semantic index based on the multi-dimensional knowledge graph, perform semantic label annotation on the elements in the standardized database according to the semantic index, and establish a two-way mapping relationship between the semantic labels of the elements and the nodes of the multi-dimensional knowledge graph;
[0063] In this embodiment, a semantic index is established based on the multi-dimensional knowledge graph, and the index is used to perform semantic label annotation on the elements in the standardized database, and finally a two-way mapping relationship between the database elements and the knowledge graph nodes is realized to support more accurate query, reasoning, and knowledge retrieval.
[0064] First of all, the establishment of the semantic index requires extracting semantic features from the multi-dimensional knowledge graph. Semantic features include information such as concept hierarchy, temporal relationship, spatial location, and causal association. For example, in the medical field, "diabetes" as a disease concept, its semantic features include attributes such as "chronic disease", "metabolic disorder", and "insulin resistance"; in the financial field, "market volatility" may have "macroeconomic impact" and "policy regulation" as semantic features; in the Buddhist field, "Chan Buddhism" may have semantic associations with "sudden enlightenment" and "division of the Northern and Southern Chan Buddhism". These features are modeled through specific vector representations, so that similar concepts are closer in the semantic space.
[0065] The construction of the semantic index can adopt a hierarchical structure, including a concept hierarchy index, a spatio-temporal index, and a causal association index:
[0066] The concept hierarchy index organizes knowledge according to domain classification. For example, under "cardiovascular diseases", there are subclasses such as "hypertension" and "coronary heart disease"; under "securities market", there are categories such as "stocks" and "futures".
[0067] The spatio-temporal index is applicable to data with time or geographical attributes. For example, "epidemic spread" is associated with a specific time range and geographical location; "development of Buddhist sects" can be associated with the evolution process in different periods.
[0068] The causal association index records the logical connections between concepts, such as "high blood sugar" leading to "diabetes", and "interest rate hike" affecting "market capital liquidity".
[0069] Secondly, perform semantic label annotation based on the semantic index, that is, assign clear semantic identifiers to the data elements in the standardized database according to the established index structure. For example, in a medical database, "medical records of diabetes patients" are labeled as "related to endocrine system diseases"; in a financial database, "news about interest rate hikes" are labeled as "information affected by monetary policies"; in a Buddhist database, "The Platform Sutra of the Sixth Patriarch" is labeled as "a classic of Chan Buddhism". Semantic labels are not only used for concept matching, but also for query optimization and knowledge reasoning.
[0070] Finally, a bidirectional mapping relationship is established between database elements and knowledge graph nodes, so that entity data in the database can be associated with the knowledge graph and query capabilities can be enhanced through graph reasoning. The bidirectional mapping relationship includes:
[0071] Forward mapping (data to graph), such as automatically mapping the "diabetes" diagnosis in the medical record to the "diabetes" concept node in the knowledge graph, and then linking it to the relevant treatment plan.
[0072] Reverse mapping (graph to data), for example, when searching for "stock market fluctuations" in the financial knowledge graph, it can be mapped to historical transaction data to obtain real market change information.
[0073] Example description:
[0074] In the field of healthcare, establishing a semantic index based on a multidimensional knowledge graph and using this index to semantically label medical data in a standardized database can significantly improve the query accuracy of the clinical decision support system. For example, when a doctor queries for "complications of diabetes", the system automatically identifies the relevant nodes of "diabetes" in the knowledge graph based on the semantic index, and associates them with highly relevant concepts such as "diabetic retinopathy" and "diabetic nephropathy", and further maps them to the corresponding case data, research papers, and treatment plans in the standardized database. This not only provides accurate matching query results, but also enables deep reasoning at the knowledge graph level through a bidirectional mapping relationship, mining potential medical association information, and improving the doctor's diagnostic efficiency and the reliability of treatment plan recommendations.
[0075] In the financial field, by establishing a semantic index of a multidimensional knowledge graph in the market analysis system and annotating historical financial data, market news, policy announcements and other data with semantic tags, the ability to analyze and predict financial events can be improved. For example, if an investor wants to query "the impact of the Fed's interest rate hike on the stock market", the system will locate the concept of "interest rate hike" based on the semantic index, and track its association with concepts such as "interest rate changes", "market fluctuations" and "asset prices" in the knowledge graph, and map it to historical market data, expert analysis reports, etc. contained in the standardized database. The establishment of a two-way mapping relationship enables the system to not only return direct historical data, but also predict possible future market reactions based on knowledge reasoning, providing investors with more comprehensive decision support.
[0076] In the field of Buddhist studies, the construction of semantic indexes based on multi-dimensional knowledge graphs can be used to organize and analyze Buddhist scriptures, the inheritance of figures, and the development context of sects. For example, when researchers query the "evolution of Zen thought", the system will, based on the semantic index, identify the core nodes of "Zen" in the knowledge graph and associate them with relevant ideological concepts, such as "sudden enlightenment", "gradual cultivation", and "gong'an", and at the same time map them to the corresponding Buddhist scriptures, inscriptions, and academic papers in the standardized database. The two-way mapping relationship ensures that researchers can not only obtain directly matching classical texts but also explore the understanding and development context of Zen thought by different sects in different periods from the inference structure of the knowledge graph, improving the depth and breadth of academic research.
[0077] By establishing semantic indexes on the basis of multi-dimensional knowledge graphs, semantic tagging of database elements and knowledge association mapping are realized, making data queries more accurate and knowledge reasoning more efficient. Compared with traditional keyword searches, semantic indexes can identify deep semantic relationships and improve the relevance of retrievals. Through the two-way mapping relationship, data queries can not only obtain directly matching content but also infer implicit associated information based on the knowledge graph reasoning, thus enhancing the intelligent application ability of the data.
[0078] S50, obtain the query intention and extract query constraint conditions from the query intention;
[0079] In this embodiment, in the knowledge query system, obtaining the query intention and extracting query constraint conditions are key steps to ensure that the query can accurately match the nodes of the knowledge graph and database elements. This process involves natural language parsing of the user input content, extracting core entities and relationships, and constructing structured constraint conditions according to the query requirements.
[0080] First, the parsing of the query intention depends on natural language processing technology. By performing word segmentation, part-of-speech tagging, syntactic parsing, and semantic analysis on the query statement input by the user, the query target is determined. For example, in the field of medical health, the user may input "the latest treatment plan for diabetes", and the system needs to identify "diabetes" as the core entity of the query and "the latest treatment plan" as the query requirement. In the financial field, when the user queries "the impact of the Fed's interest rate hike on the stock market", the system needs to identify "the Fed's interest rate hike" as the core event and "the impact on the stock market" as the query direction. In the field of Buddhist studies, when the user inputs "the main scriptures of Zen", the system needs to parse "Zen" as the query object and "the main scriptures" as the query scope.
[0081] Secondly, the core components of the query intention usually include:
[0082] Set of core entities: The key concepts involved in the user query, such as "diabetes", "the Fed", and "Zen".
[0083] Set of relational predicates: Predicates that express the relationship between the query target and the core entity, such as "treatment plan", "influence", "classical".
[0084] Implicit semantic information: Information that may not be explicitly mentioned in the query but can be supplemented through reasoning, such as time, space, or logical relationships. For example, "the latest treatment plan" implies a time constraint, and "stock market influence" may imply an association with economic events.
[0085] After completing the query intention parsing, the system needs to extract query constraint conditions to ensure that the query can match the most relevant data. Query constraint conditions are usually divided into:
[0086] Explicit constraint conditions: Structured constraint conditions directly extracted from the query intention, such as "diabetes AND the latest treatment plan".
[0087] Implicit constraint conditions: Constraint conditions supplemented based on knowledge reasoning. For example, "the latest treatment plan for diabetes" implies a time constraint, and the system needs to infer "relevant literature in the past five years" as the screening condition.
[0088] Query scope constraints: Specify the data source, domain, or data type, such as "query the medical database" or "limit the search to financial market data".
[0089] Finally, after the query intention parsing and constraint condition extraction are completed, the system converts them into a standardized query format for subsequent knowledge graph reasoning and database retrieval. For example, it is converted into SPARQL query language, SQL query statement, or knowledge graph reasoning instructions.
[0090] Through the combination of natural language parsing and knowledge reasoning, the accurate understanding of the query intention is achieved, and structured query constraint conditions are automatically generated, improving the accuracy and efficiency of the query. Compared with the traditional keyword matching retrieval method, it can handle more complex query semantics, avoid information omission caused by expression differences, and can provide more intelligent query recommendations through knowledge reasoning, improving the intelligence level of the system.
[0091] S60, according to the query constraint conditions and the bidirectional mapping relationship, perform association reasoning in the multi-dimensional knowledge graph to generate a query result.
[0092] In this embodiment, during the query process of the knowledge graph, association reasoning is performed according to the query constraint conditions and the bidirectional mapping relationship to generate accurate query results. The core of this step is to utilize the structure of the multi-dimensional knowledge graph and the bidirectional mapping relationship, mine the deep knowledge related to the query intention based on the reasoning mechanism, and provide accurate and interpretable retrieval results.
[0093] In the process of performing association reasoning based on query constraints and bidirectional mapping relationships, the core role of association reasoning is to utilize entities, relationships, and semantic rules in the knowledge graph. Based on the existing data, through logical reasoning and semantic expansion, generate results that meet the query objectives. Since the knowledge graph does not directly store query results but stores entities and their relationships, it is necessary to perform association reasoning to calculate and derive the final data required by the user's query.
[0094] The basic process of association reasoning can include path reasoning, semantic matching reasoning, causal reasoning, and association strength calculation. Path reasoning is to find entities or relationships that meet the query constraints by traversing the association edges in the knowledge graph. For example, when the user queries "the latest treatment plan for diabetes", the system will not directly return a database record, but find relevant information through path reasoning such as "diabetes" → "type 2 diabetes" → "insulin treatment" → "the latest research in 2023". Semantic matching reasoning is used to expand the scope of the query objective. For example, when querying "diabetes treatment", the system can recognize the relationships between "diabetes" and "type 2 diabetes", "type 1 diabetes", so as to include relevant treatment plans. Causal reasoning is used to identify causal chains. For example, when querying "the impact of interest rate hikes on the market", the system may reason "interest rate hikes" → "decrease in money supply" → "increase in corporate financing costs" → "market volatility", rather than just returning the historical records of interest rate hikes. Association strength calculation is used to screen the most valuable information in the reasoning path. For example, the system can assign weights to different reasoning results according to factors such as the credibility of the data source, time priority, and matching degree with the query objective, and return the optimal query result.
[0095] The key to generating query results through association reasoning is that the query objective is usually not a single piece of data directly stored in the database, but the final result needs to be derived from multiple data sources and multiple entity relationships. For example, when a user queries "the impact of the latest economic policies on the stock market" in the financial field, the system needs to combine various data sources such as policy information, historical stock market data, and market analysis reports to reason how the policy affects the market, rather than simply returning a certain news or report. Therefore, through association reasoning, the query objective can be matched with the complex relationships in the knowledge graph, and query results based on logical derivation and meeting semantic requirements can be generated, enabling users to obtain more accurate and interpretable answers.
[0096] First, the parsing of query constraints is matched with the initial nodes. Based on the query constraints extracted in the previous step, entities, relationships, or concepts involved in the query are mapped to the initial nodes in the multi-dimensional knowledge graph through a two-way mapping relationship. For example, in the field of healthcare, when a user queries "the latest treatment plan for diabetes", the system maps "diabetes" to the disease node in the knowledge graph and associates it with relevant treatment plan nodes; in the financial field, when querying "the impact of the Fed's interest rate hike on the stock market", the system maps "interest rate hike" to the economic policy node and associates it with nodes such as market fluctuations and stock market indices; in the field of Buddhist studies, when querying "the development context of Chan Buddhism", the system maps "Chan Buddhism" to the sect node and associates it with historical events, representative figures, classic literature, etc.
[0097] Secondly, multi-hop association reasoning is performed based on the knowledge graph. After determining the initial nodes, the system uses knowledge graph reasoning techniques based on the query constraints to perform multi-hop path traversal and filter nodes that meet the constraints. For example:
[0098] In the medical field, when querying "the latest drugs related to diabetes", the system will traverse in the knowledge graph from the "diabetes" node along the "treatment drugs" relationship to obtain directly associated drug nodes, and combine time constraints to filter new drugs in the past five years.
[0099] In the financial field, when querying "the indirect impact of the Fed's interest rate hike on the stock market", the system will traverse from the "interest rate hike" node along the "impact" relationship to the "interest rate change" node, then jump to the "fund liquidity" node, and finally reach the "stock market volatility" node, and perform analysis in combination with market data.
[0100] In the field of Buddhist studies, when querying "the development process of Chan Buddhism", the system will traverse along the "sect evolution" relationship, jump from "Chan Buddhism" to "Bodhidharma", then find "Huineng" along the "inheritance" path, and then reason backward about the development of the "Southern School" and "Northern School" and subsequent sects.
[0101] Then, the screening and optimization of query results. During the reasoning process, the system combines the query constraints to optimize the query path, including:
[0102] Relationship weight screening: If the query involves multiple association paths, the system selects the most relevant path according to the weight calculation in the knowledge graph. For example, in the financial field, if there are multiple different association paths between "market fluctuations" and "policy adjustments", the system can select the most representative path based on the statistical weights of historical data.
[0103] Temporal and Spatial Constraints: For queries involving temporal and spatial dimensions, the system filters relevant nodes according to the timeline or geographical location. For example, in a medical query, if a user wishes to obtain "diabetes treatment plans after 2020", the system will automatically filter nodes whose time attributes meet the constraint conditions.
[0104] Logical Constraint Verification: During the reasoning process, if there are logical conflicts in some association paths (such as A causing B, and B causing A), the system will perform a logical consistency check, delete the conflicting paths, and ensure the rationality of the reasoning results.
[0105] Finally, the formatting and return of the query results. After the reasoning is completed, the system converts the query results into a standard format and presents them in different ways according to the type of query intent. For example:
[0106] In the medical field, when querying "the latest diabetes treatment plan", structured treatment plans are returned, including drug names, indications, clinical trial data, and links to the original literature are provided.
[0107] In the financial field, when querying "the impact of the Fed's interest rate hikes on the stock market", data such as interest rate hike historical events, market reaction analysis reports, and relevant policy interpretations are returned.
[0108] In the field of Buddhist studies, when querying "the development process of Chan Buddhism", information such as the timeline of sect evolution, the relationship map of key figures, and a list of classic literature are returned.
[0109] By combining query constraint conditions and bidirectional mapping relationships, association reasoning is performed in the knowledge graph, realizing intelligent and efficient knowledge retrieval capabilities. Compared with traditional keyword-matching-based search methods, it can deeply understand query intent and provide highly relevant retrieval results based on knowledge reasoning. At the same time, the system supports multi-hop reasoning, can discover indirect relationships between query entities, improve the intelligence level of queries, and can optimize the accuracy and credibility of query results based on technologies such as weight calculation and logical consistency verification.
[0110] The present invention relates to the technical field of data analysis and can be applied to business scenarios such as medical health, fintech, and cultural research. It discloses a multi-source knowledge processing and query method, including: obtaining multi-source heterogeneous data and performing cleaning and standardization processing to generate a standardized database; defining core concepts and association relationships, and extracting knowledge elements through feature extraction; constructing a multi-dimensional knowledge graph based on the core concepts, association relationships, and knowledge elements; establishing a semantic index on the multi-dimensional knowledge graph, and performing semantic label annotation on the elements in the standardized database to form a bidirectional mapping relationship; obtaining a query intention, and extracting query constraint conditions from the query intention; based on the query constraint conditions and the bidirectional mapping relationship, performing association reasoning in the multi-dimensional knowledge graph to generate a query result. By constructing a standardized database for multi-source heterogeneous data, the present invention ensures the uniformity of data formats and improves data consistency; through the construction of a multi-dimensional knowledge graph and the establishment of a semantic index, it enhances the relevance and structured expression of data, making the relationships between knowledge elements clearly traceable; through association reasoning based on query constraint conditions, it improves the accuracy and efficiency of queries, avoids the information redundancy problem caused by traditional keyword matching, and realizes the efficient extraction and accurate query of complex knowledge.
[0111] In one embodiment, the above S10 includes:
[0112] S101, deploying multiple distributed data collection nodes, and each distributed data collection node selects a corresponding collection strategy according to the type of the data source to be collected;
[0113] S102, through the multiple distributed data collection nodes, performing a data collection task on multiple data sources to be collected based on the corresponding collection strategies to obtain the multi-source heterogeneous data;
[0114] S103, performing cleaning processing on the multi-source heterogeneous data to remove duplicate, missing, and invalid data, and generating cleaned data;
[0115] S104, performing standardization processing on the cleaned data to convert data in different formats into standard data in a unified format;
[0116] S105, based on the standard data in the unified format, constructing the standardized database.
[0117] In this embodiment, multi-source heterogeneous data is acquired through a distributed data acquisition module, and the acquired data is cleaned and standardized to build a standardized database, ensuring the high quality, scalability, and unified management of the data. First, multiple distributed data acquisition nodes are deployed so that data acquisition can adapt to the distributed storage characteristics of multi-source heterogeneous data. Different data sources have different structures. For example, structured data comes from electronic medical records, financial statements, etc., and can usually be obtained through APIs or database connections; semi-structured data includes log data in JSON and XML formats, and keyword fields can be extracted through parsing tools; unstructured data includes scientific research papers, image data, audio and video materials, and methods such as web crawlers, OCR technology, and speech transcription are required for acquisition.
[0118] Secondly, based on the acquisition strategy, data acquisition tasks are executed to ensure efficient and accurate acquisition of target data. For data that needs to be continuously updated, a regular acquisition strategy can be adopted. For example, for financial market data and real-time medical monitoring data, the system regularly executes data acquisition tasks to keep the data up-to-date. For data acquisition requirements driven by specific events, such as new case reports in the medical field and market fluctuations in the financial field, event-triggered acquisition can be adopted. When a new event is detected, the data acquisition task is immediately started. For the integration of a large amount of historical data, an incremental acquisition strategy can be adopted, only extracting the data that has changed since the last acquisition, thereby reducing storage redundancy and improving data processing efficiency.
[0119] After the data acquisition is completed, the multi-source heterogeneous data is cleaned to remove invalid or incomplete data and improve data quality. Duplicate data deduplication detects and deletes redundant data through methods such as hash mapping and text similarity matching, thereby ensuring data uniqueness. Missing value filling uses techniques such as statistical analysis, interpolation methods, and machine learning models to complete the missing information based on historical data and improve data integrity. Invalid data filtering targets data entries with format errors, missing fields, or invalid content, and filters them through rule matching and anomaly detection models to remove data that does not meet the standards.
[0120] After the data cleaning is completed, data standardization is performed to ensure the consistency of data formats, enabling data from different sources to be compatible for storage and analysis. Data format conversion unifies data fields from different sources. For example, formats such as time, date, and currency are converted to standard formats to ensure data consistency. Field mapping addresses differences in field naming in different systems. For example, "patient ID" and "patient number" are mapped to the same field for cross-system data compatibility. Unit conversion is used to standardize data with different measurement units. For example, in medical data, "mg / dL" is converted to "mmol / L", and in financial data, "USD / JPY" is converted to "US dollar against Japanese yen" to ensure data comparability.
[0121] Finally, a standardized database is constructed based on the standardized data to support efficient data storage, indexing, and querying. The standardized database adopts a hierarchical storage architecture, storing the original data, cleaned data, and standardized data separately to meet the query requirements at different application levels. Index optimization improves the access speed of frequently queried fields by establishing an efficient index structure, making data retrieval more efficient. The data version management mechanism ensures data traceability, enabling the restoration of different versions of data when retrospective analysis is needed, providing support for medical research, financial analysis, and historical data analysis.
[0122] In this embodiment, through distributed data collection, data cleaning, data standardization, and the construction of a standardized database, the efficient processing and unified storage of multi-source heterogeneous data are realized. It can support the efficient integration of multi-source data during the data collection process, improve the accuracy and integrity of data quality during the data cleaning process, ensure seamless compatibility of data from different sources during the data standardization process, and finally, the constructed standardized database has better query efficiency and data availability.
[0123] In one embodiment, the above S20 includes:
[0124] S201, defining a set of core concepts in the domain according to the sources and characteristics of the multi-source heterogeneous data;
[0125] S202, determining the inheritance relationship, causal relationship, and belonging relationship among the core concepts in the set of core concepts to generate a domain ontology structure containing association constraints;
[0126] S203, extracting entity, attribute, and relationship triples from the text of the multi-source heterogeneous data through a named entity recognition module, and generating structured text knowledge elements based on the entity, attribute, and relationship triples;
[0127] S204, performing region segmentation and feature extraction on the images of the multi-source heterogeneous data through an image semantic segmentation module to generate image semantic feature elements;
[0128] S205, spatially mapping and associating the structured text knowledge elements with the image semantic feature elements through a cross-modal alignment module to generate fused knowledge elements;
[0129] S206, verifying the logical consistency of the fused knowledge elements based on the association constraints of the domain ontology structure to filter out the knowledge elements with conflicting relationships in the fused knowledge elements and generate standardized knowledge elements.
[0130] In this embodiment, the core concepts of the domain and their association relationships are defined, and knowledge elements are extracted from multi-source heterogeneous data through a feature extraction module to construct a structured knowledge system, providing semantic support for subsequent knowledge storage, reasoning, and query. First, according to the sources and characteristics of multi-source heterogeneous data, a set of core concepts within the domain is defined to ensure the basic construction of the knowledge graph. Different data sources may contain knowledge in different domains, such as diseases, symptoms, and treatment plans in the medical and health field, market indices, economic events, and investment tools in the financial field, sects, scriptures, and historical figures in the field of Buddhist studies, etc. Therefore, when defining core concepts, it is necessary to ensure that the core knowledge points of the target domain are covered while avoiding redundant or unnecessary information.
[0131] Secondly, the inheritance relationships, causal relationships, and belonging relationships among the concepts in the set of core concepts are determined to construct a domain ontology structure containing association constraints. Inheritance relationships are used to describe the hierarchical structure between concepts. For example, in the medical field, "diabetes" inherits from "metabolic diseases", and "the Fed's interest rate hike" inherits from "monetary policy adjustments"; causal relationships are used to represent the logical dependencies between concepts. For example, "viral infection" can cause "abnormal immune system", and "interest rate hikes" may lead to "a decline in market liquidity"; belonging relationships are used to identify the belonging of concepts. For example, "Heart Sutra" belongs to "Buddhist scriptures", and "A-shares" belong to "securities markets".
[0132] After determining the concept relationships, entity, attribute, and relationship triples are extracted from the text data through a named entity recognition module to construct structured text knowledge elements. Named entity recognition can extract meaningful information from unstructured text data. For example, in medical literature, it can identify disease names, drug names, and symptom descriptions; in financial reports, it can extract company names, policy announcements, and market indices; in Buddhist scriptures, it can identify people, dharma names, sects, etc., and use attribute analysis to expand the features of each entity, such as the incidence rate of diseases, the indications of drugs, and the market volatility of investment tools.
[0133] Subsequently, the images of multi-source heterogeneous data are subjected to region segmentation and feature extraction through an image semantic segmentation module to enhance the multi-modal expression ability of knowledge. Image semantic segmentation technology can analyze data such as medical images, financial market trend charts, and religious murals to identify key features. For example, in the medical field, information on lesion areas can be extracted from medical images and associated with disease entities; in the financial field, abnormal signals in market trend charts can be analyzed to extract market volatility patterns; in Buddhist studies, information such as Buddha statue features, gestures, and costumes can be extracted from religious art works and associated with text knowledge.
[0134] Through the cross-modal alignment module, the structured text knowledge elements and the image semantic feature elements are spatially mapped and associated to achieve the fusion of different modal data. This process can ensure the semantic matching between text and image data and improve the completeness of the knowledge system. For example, in the medical field, the disease descriptions in the text can be matched with the diagnostic results in the imaging data; in the financial field, the text of policy announcements can be aligned with market trend charts in the time dimension; in Buddhist studies, Buddhist scriptures can be associated with religious murals to construct a cross-modal knowledge graph.
[0135] Finally, based on the association constraints of the domain ontology structure, the logical consistency of the fused knowledge elements is verified, potential conflict information is filtered, and standardized knowledge elements are generated. The verification of logical consistency can ensure the reliability of the knowledge system. For example, in the medical field, if a research literature claims that "a certain drug can cure a certain disease", but historical clinical trial data shows that the drug has no significant effect, this information needs to be marked as low confidence or excluded. In the financial field, if the market prediction model contradicts the trend obtained from historical data analysis, the system needs to adjust the weights or clean the data. In Buddhist studies, if there are contradictory descriptions of the same person in different versions of the scriptures, text traceability analysis can be used to determine a more credible version.
[0136] This embodiment defines domain core concepts, establishes association relationships, and fuses multi-modal data through an automated method, improving the integrity and reasoning ability of the knowledge system. It can efficiently process large-scale heterogeneous data, improve the accuracy and consistency of knowledge extraction. Through knowledge graph construction, the system can support complex query reasoning and enhance the accuracy of information matching through cross-modal data fusion, ultimately constructing a more intelligent knowledge management system.
[0137] In one embodiment, the above S30 includes:
[0138] S301, defining multi-dimensional coordinate axes of the multi-dimensional knowledge graph, where the multi-dimensional coordinate axes include a time dimension axis, a space dimension axis, and a logical dimension axis;
[0139] S302, mapping the time attribute of the knowledge element to the time dimension axis, the space attribute to the space dimension axis, and the logical attribute to the logical dimension axis to generate knowledge nodes containing three-dimensional coordinate labels;
[0140] S303, establishing weighted causal association edges between the knowledge nodes according to the causal relationship in the association relationship;
[0141] S304, constructing an inheritance tree structure with the core concept as the root node according to the inheritance relationship in the association relationship;
[0142] S305. Establish an ownership association edge between the knowledge nodes according to the ownership relationship in the association relationship, and mark the time range and space range of ownership;
[0143] S306. Perform conflict detection on the causal association edge, inheritance tree structure, and ownership association edge. If a conflict relationship is detected between the same set of nodes, determine and delete the conflicting association edge according to the data source priority of the knowledge elements, and generate a verified multi-dimensional knowledge graph.
[0144] In this embodiment, a multi-dimensional knowledge graph is constructed based on core concepts, association relationships, and knowledge elements to support multi-dimensional data organization, query, and reasoning. First, define the multi-dimensional coordinate axes of the multi-dimensional knowledge graph so that the knowledge graph can structurally represent knowledge from multiple dimensions. The multi-dimensional coordinate axes include a time dimension axis, a space dimension axis, and a logical dimension axis, which are used to represent the time attribute, geographical attribute, and logical relationship of knowledge respectively. For example, in a medical knowledge graph, the time dimension can be used to represent the incidence period and treatment process of a disease, the space dimension can be used to mark the prevalence of the disease in different regions, and the logical dimension is used to represent the relationship between diseases, symptoms, and treatment plans. In a financial knowledge graph, the time dimension is used to track market changes, the space dimension is used to represent the region where financial events occur, and the logical dimension is used to represent the causal relationship between economic indicators and policy decisions. In the Buddhist knowledge system, the time dimension can be used to track the historical evolution of scriptures, the space dimension can be used to mark the geographical distribution of the spread of Buddhist culture, and the logical dimension is used to organize the associations between sects, doctrines, and historical figures.
[0145] Secondly, based on the definition of the multi-dimensional coordinate axes, map the time attribute of the knowledge elements to the time dimension axis, the space attribute to the space dimension axis, and the logical attribute to the logical dimension axis, so that the knowledge nodes can carry complete multi-dimensional information. Each knowledge node is defined by a three-dimensional coordinate label to ensure that the semantic level, time evolution, and spatial distribution of knowledge can be completely characterized. For example, in the medical field, a knowledge node can represent a certain disease, the time dimension corresponds to the prevalence trend of the disease, the space dimension corresponds to the spread area of the disease, and the logical dimension corresponds to the causative factors and treatment methods of the disease. In the financial field, a market event node can mark the time when it occurs, the market scope it affects, and its logical relationship with other events. In the field of Buddhist studies, a historical figure node can mark the time of his birth and death, the activity area, and the master-disciple relationship with other figures.
[0146] After the knowledge nodes are established, based on the causal relationships in the association relationships, causal association edges with weight tags are established between the knowledge nodes to represent the causal influences between events or concepts. The weights of the causal association edges reflect the strength and confidence of the relationships. For example, in a medical knowledge graph, the weight of the causal relationship edge between "hypertension" and "heart disease" is relatively high because hypertension is an important factor inducing heart disease. In a financial knowledge graph, the weight value of the causal relationship between policy changes and market fluctuations can be obtained through training with historical data, and the cascade effect of market events can be predicted through causal reasoning methods. In the Buddhist knowledge system, the formation of different sects can be described by causal associations. For example, the degree to which "Chan Buddhism" is influenced by "Taoist thought" can be weighted through knowledge reasoning.
[0147] Then, based on the inheritance relationships in the association relationships, an inheritance tree-like structure with the core concept as the root node is constructed to reflect the hierarchical organization of knowledge. For example, in the medical field, the disease classification system can be represented as an inheritance relationship. For instance, "cancer" is a subclass of "malignant diseases", and "breast cancer" is a subclass of "cancer". In the financial field, the inheritance relationship of asset categories can be used to identify the hierarchy between different market instruments. For example, "stocks" inherit from "securities assets", and "A-shares" inherit from "stocks". In Buddhist studies, sectarian inheritance can be represented by inheritance relationships. For example, "Linji School" inherits from "Chan Buddhism", and the inheritance relationships between different patriarchs can be arranged in sequence.
[0148] Furthermore, based on the belonging relationships in the association relationships, belonging association edges are established between the knowledge nodes, and the time range and space range of belonging are marked to ensure the clear attribution of knowledge. For example, in the medical field, "aspirin" belongs to the category of "anti-inflammatory drugs", and its application range is marked as "prevention and treatment of cardiovascular diseases". In the financial field, "the Fed's interest rate hike" belongs to the category of "monetary policy adjustment", and its applicable range is limited to the "US market". In the field of Buddhist studies, "Heart Sutra" belongs to "Mahayana sutras", and its origin time and geographical information are marked to ensure the accuracy of research.
[0149] Finally, conflict detection is performed on the causal association edges, the inheritance tree-like structure, and the belonging association edges to ensure the logical consistency of the knowledge graph. If a conflict relationship is detected between the same set of nodes, the conflicting association edges are determined and deleted according to the priority of the data sources of the knowledge elements. For example, in the medical field, if two research reports conflict on the efficacy of a certain drug, the system can make adjustments based on the authority of the data sources. In the financial field, if two economic models conflict in their predictions of market trends, the system can make decisions based on the currency and historical credibility of the data. In Buddhist studies, if different versions of Buddhist scriptures differ in their descriptions of a certain historical event, a more credible version can be determined through text traceability analysis. Ultimately, the system generates a verified multi-dimensional knowledge graph to support subsequent semantic reasoning and query analysis.
[0150] In this embodiment, by establishing a multi-dimensional knowledge graph, the unified representation and multi-dimensional analysis of heterogeneous knowledge are realized. It can comprehensively consider time, space, and logical relationships to more accurately describe the evolution and association of knowledge. Through causal reasoning and inheritance structure construction, the system can automatically discover the deep connections between knowledge and use a conflict detection mechanism to ensure the logical consistency of the data.
[0151] In one embodiment, the above S40 includes:
[0152] S401, extracting a semantic feature vector from the nodes of the multi-dimensional knowledge graph, where the semantic feature vector includes a node name, a set of node attributes, and an associated edge type;
[0153] S402, constructing a hierarchical semantic index based on the semantic feature vector;
[0154] S403, performing a similarity match between the text paragraphs and image regions in the standardized database and the hierarchical semantic index to generate fine-grained semantic labels;
[0155] S404, storing a reverse pointer in the nodes of the multi-dimensional knowledge graph, where the reverse pointer points to the corresponding text paragraph or image region in the standardized database;
[0156] S405, establishing a two-way hash mapping relationship between the fine-grained semantic labels and the nodes of the multi-dimensional knowledge graph.
[0157] In this embodiment, a semantic index is established based on the multi-dimensional knowledge graph to improve the accuracy and query efficiency of data retrieval. First, a semantic feature vector is extracted from the nodes of the multi-dimensional knowledge graph. The semantic feature vector of each node consists of a node name, a set of node attributes, and an associated edge type. The node name is used to uniquely identify a knowledge concept, such as "hypertension", "stock market fluctuations", "Zen thought"; the set of node attributes contains the core information of the concept, such as symptom descriptions and drug indications in the medical field, market indicators and policy influencing factors in the financial field, and classic sources and historical backgrounds in Buddhist studies; the associated edge type is used to describe the relationship between this node and other nodes, such as causal relationships, inheritance relationships, and ownership relationships, etc., to ensure the integrity of the knowledge system.
[0158] After extracting semantic feature vectors, a hierarchical semantic index is constructed based on these vectors, enabling efficient retrieval and correlation analysis of semantic information at different levels. The hierarchical semantic index includes a concept hierarchical index, a spatio-temporal index, and a causal association strength index. The concept hierarchical index organizes concepts according to the structure of the knowledge system. For example, in the medical field, diseases can be organized hierarchically as "system - disease category - specific disease"; in the financial field, they can be classified as "industry - market - asset class"; and in the field of Buddhist studies, they can be graded as "sect - scripture - ideological system". The spatio-temporal index indexes knowledge points based on the time dimension and space dimension, such as the epidemic time of diseases, the applicable regions of policy changes, and the dissemination paths of classical thoughts. The causal association strength index is used to identify the degree of causal influence between concepts, such as the impact strength of a certain treatment method on disease remission, the influence weight of a certain policy on market fluctuations, and the influence degree of a certain religious thought on historical and cultural development.
[0159] After constructing the semantic index, the index structure is used to perform semantic matching on text paragraphs and image regions in the standardized database to generate fine-grained semantic labels. For the semantic matching of text paragraphs, deep semantic similarity calculation models, such as natural language processing technologies like BERT, are used to calculate the matching degree between text paragraphs and knowledge graph concepts, so as to return the most relevant results when users query. For the matching of image regions, computer vision technologies are used to extract image features through convolutional neural networks (CNNs) and calculate their matching degree with knowledge graph concepts based on the semantic index. For example, in medical image analysis, lesions in images are automatically identified and matched with corresponding disease concepts; in financial data visualization analysis, the patterns of market trend charts are matched with historical market events; and in Buddhist studies, the themes of Buddhist murals are identified and their associations with classical contents are matched.
[0160] To ensure two-way retrievability of data, reverse pointers are stored in the nodes of the multi-dimensional knowledge graph, and the reverse pointers point to the text paragraphs or image regions in the standardized database related to that node. This can ensure that not only can semantic information be obtained by querying the database, but also the original data sources can be traced back from the knowledge graph. For example, in the medical field, doctors can search for treatment plans for a certain disease through the knowledge graph and directly locate relevant clinical case documents or medical images. In the financial field, analysts can search for historical data on a certain market trend through the knowledge graph and trace back to relevant policy documents or market trend charts. In Buddhist studies, researchers can search for the inheritance context of specific doctrines through the knowledge graph and directly access relevant classical literature or historical relics.
[0161] Finally, a bidirectional hash mapping relationship is established between the fine-grained semantic tags and the multi-dimensional knowledge graph nodes to ensure the efficiency and consistency of data query. The bidirectional hash mapping relationship enables each database element to be quickly mapped to the corresponding knowledge graph node and supports the rapid backtracking of concepts in the knowledge graph to the original database content. For example, in medical research, users can query the indications of a certain drug. The system quickly locates the relevant drug information in the knowledge graph based on the hash mapping and maps back to the standardized database to obtain complete medical literature and experimental data. In financial analysis, investors can input "investment strategies during market turmoil" as a query. The system will match historical market trend data through the hash mapping and return relevant investment strategy reports. In Buddhist studies, researchers can input the name of a historical figure. The system retrieves their life information based on the hash mapping and maps it to relevant classical literature and research papers.
[0162] In this embodiment, a semantic index is established through a multi-dimensional knowledge graph, and a bidirectional mapping relationship is constructed, improving the accuracy of data retrieval and the traceability of knowledge reasoning. It can capture the semantic information of data more accurately and achieve efficient query through a multi-level index structure. Through the reverse pointer and hash mapping mechanism, this solution supports backtracking from the knowledge graph to the original database, improving the transparency and integrity of data query.
[0163] In one embodiment, the above S50 includes:
[0164] S501, obtain a query statement, and perform intent recognition on the query statement through a natural language parsing module to generate a query intent object including a core entity set, a relationship predicate set, and implicit semantics;
[0165] S502, extract explicit constraint conditions from the query intent object, where the explicit constraint conditions include the entity names in the core entity set, the relationship types in the relationship predicate set, and the node attribute conditions defined in the multi-dimensional knowledge graph;
[0166] S503, determine implicit constraint conditions according to the implicit semantics of the query intent object;
[0167] S504, logically fuse the explicit constraint conditions and the implicit constraint conditions to generate a set of structured query constraint conditions.
[0168] In this embodiment, the query intention is obtained and the query constraint conditions are extracted from it to ensure that subsequent queries can accurately match the target data. First, the query statement is obtained, and intention recognition is performed through the natural language parsing module to generate a query intention object containing a core entity set, a relationship predicate set, and implicit semantics. The natural language parsing module performs lexical, syntactic, and semantic level parsing on the query statement based on a pre-trained language model to identify the key entities and relationship structures in the query. The core entity set refers to the main concepts involved in the query, such as disease names, financial terms, religious scriptures, etc.; the relationship predicate set represents the semantic relationships in the query, such as "affect", "develop", "apply to", etc.; the implicit semantics is used to identify the potential requirements in the query, such as limiting time, region, data source, etc.
[0169] After parsing out the query intention object, the explicit constraint conditions are extracted from it. The explicit constraint conditions include the entity names in the core entity set, the relationship types in the relationship predicate set, and the node attribute conditions defined in the knowledge graph. The entity names are used to determine the specific concepts involved in the query, the relationship types are used to constrain the logical relationship of the query content, and the node attribute conditions are used to limit the retrieval scope. For example, in some queries, the user may want to limit the specific attributes of the query object. For example, when querying "the side effects of the COVID-19 vaccine", the system needs to add the node attribute "side effects" to the query conditions to ensure that the retrieved results accurately match the user's needs.
[0170] In addition to the explicit constraint conditions, the implicit constraint conditions also need to be derived from the implicit semantics of the query intention object to supplement the query scope. The implicit constraint conditions mainly include time range derivation, space range derivation, and logical association strength derivation. Time range derivation is used to identify whether the query has timeliness. For example, when querying "the latest treatment plan", the system should automatically derive the latest research results. Space range derivation is used to judge whether the query has geographical restrictions. For example, when querying "financial policies of different countries", the system needs to match the policy data related to different countries. Logical association strength derivation is used to judge the degree of relevance between the concepts involved in the query to exclude low-related or irrelevant information and improve the query accuracy.
[0171] Finally, the explicit constraint conditions and the implicit constraint conditions are logically fused to generate a structured query constraint condition set to ensure that the query request can be effectively executed. The logical fusion involves condition priority adjustment, weight assignment, and conflict resolution mechanisms to ensure the integrity and reasonableness of the query conditions. For example, if there is a conflict between the explicit constraint conditions and the implicit constraint conditions, such as the explicit constraint requires querying data in 2020, while the implicitly derived time range is the last five years, the system can adjust the time range based on the query context to ensure the logical consistency of the query. Finally, the structured query constraint condition set is provided as input to the subsequent knowledge retrieval and reasoning modules to improve the query accuracy and response speed.
[0172] In this embodiment, through the methods of natural language parsing and semantic derivation, the accuracy of query intention parsing is improved, and the retrieval precision is optimized based on the integration of explicit and implicit constraint conditions. It can more accurately understand the user's intention and provide structured query conditions, thereby improving the relevance and integrity of the retrieval.
[0173] In one embodiment, the above S60 includes:
[0174] S601, parsing the target type identifier, explicit constraint condition list, implicit constraint conversion parameter, and priority weight in the structured query constraint condition set;
[0175] S602, according to the forward hash table in the bidirectional mapping relationship, mapping the entity names in the explicit constraint condition list to the initial node set in the multi-dimensional knowledge graph;
[0176] S603, based on the target type identifier, performing an inference operation in the multi-dimensional knowledge graph starting from the initial node set to generate an intermediate result set;
[0177] S604, according to the relationship type in the explicit constraint condition list and the logical association strength threshold, time range, and space range in the implicit constraint conversion parameter, performing node filtering and edge pruning on the intermediate result set to generate a set of valid association paths;
[0178] S605, based on the target node or target relationship information in the set of valid association paths, querying and obtaining the corresponding standardized database elements through the reverse hash table in the bidirectional mapping relationship, and sorting the standardized database elements according to the priority weight to generate the query result.
[0179] In this embodiment, according to the query constraint conditions and the bidirectional mapping relationship, an association inference is performed in the multi-dimensional knowledge graph to generate a result that meets the query requirements. First, the structured query constraint condition set is parsed to extract the target type identifier, explicit constraint condition list, implicit constraint conversion parameter, and priority weight. The target type identifier is used to identify the category of the query target, such as whether the query is about an entity, a relationship, a path, or statistical information. The explicit constraint condition list contains the conditions explicitly specified in the query. For example, when querying "common complications of diabetes", the explicit constraints include "diabetes" as the core entity and "complications" as the relationship predicate. The implicit constraint conversion parameters include the constraints that are not explicitly expressed in the query but are obtained through semantic derivation, such as the time range, geographical range, and logical association strength threshold. The priority weight is used to adjust the influence of different query conditions. For example, if the user is more concerned about the latest data, the weight of the time constraint is higher.
[0180] Next, according to the forward hash table in the bidirectional mapping relationship, map the entity names in the explicit constraint condition list to the initial node set in the multi-dimensional knowledge graph. The forward hash table stores the mapping relationship between the standardized database elements and the knowledge graph nodes, so as to quickly retrieve the target entity during query. For example, when the user queries "the impact of the Fed's interest rate hike on the market", the system can locate the knowledge graph node corresponding to "the Fed's interest rate hike" through the forward hash table and mark it as the initial query node.
[0181] Subsequently, based on the target type identifier, starting from the initial node set, perform inference operations in the multi-dimensional knowledge graph to generate an intermediate result set. The inference operations are based on graph traversal, rule derivation, or machine learning inference models to mine the implicit relationships in the knowledge graph. For example, in a medical query, the system can infer whether the side effects of a certain drug are related to other diseases; in a financial query, the system can infer whether a certain policy change will affect market fluctuations; in a Buddhist studies query, the system can infer how the academic thoughts of a certain eminent monk are influenced by previous figures.
[0182] Then, according to the relationship type in the explicit constraint condition list and the logical association strength threshold, time range, and space range in the implicit constraint conversion parameter, perform node filtering and edge pruning on the intermediate result set to optimize the query result. Node filtering is used to eliminate nodes with low relevance to the query target, such as eliminating irrelevant industry data in financial market analysis; edge pruning is used to reduce unnecessary relationship calculations and improve the inference efficiency. For example, in disease transmission analysis, only retain the causal paths related to epidemiology.
[0183] Finally, based on the reverse hash table in the bidirectional mapping relationship, extract the corresponding node or relationship information from the set of valid association paths, and sort the results according to the priority weight to generate the final query result.
[0184] In the process of querying and obtaining the corresponding standardized database elements based on the target nodes or target relationship information in the set of valid association paths through the reverse hash table in the bidirectional mapping relationship, it is first necessary to extract key information from the inference results of the knowledge graph. The set of valid association paths is generated by query intention parsing, query constraint condition matching, and knowledge graph inference, which contains the target nodes or target relationship information that meets the query requirements. These target nodes or relationships are not only elements of the knowledge graph but also key data points for establishing a bidirectional mapping relationship with the standardized database. Therefore, before executing the query, the system will parse the valid association paths, extract the core nodes and association relationships, and confirm whether these target data have corresponding standardized database elements based on the bidirectional mapping relationship to ensure the accuracy and traceability of data retrieval.
[0185] During the query process, the system retrieves standardized database elements using an inverse hash table. The inverse hash table stores the mapping relationships between the nodes or relationships of the knowledge graph and the specific data in the standardized database. When querying, the system first performs a hash index match for the target node or target relationship to quickly locate its corresponding database entry. This process can adopt a multi-level index structure to enable efficient query completion. For example, in the financial field, if the target node involves the impact of a certain policy on the market, the system can query the specific economic data, policy interpretation documents, or market trend analyses of this policy in the standardized database through the inverse hash table. In the medical field, if the target node involves the latest treatment methods for a certain disease, the system can retrieve relevant medical research papers, experimental data, or guideline documents to support the scientificity and credibility of the query results.
[0186] After obtaining the standardized database elements, the system sorts the data according to the priority weights. The priority weights are determined by multiple factors, including the time priority of the data, the credibility of the source, and the semantic matching degree, etc. The time priority ensures that the latest data is ranked higher. For example, the latest clinical trial data in the medical field is more valuable for reference than outdated treatment plans. The credibility of the source is used to distinguish the authority of the data. For example, the guidelines issued by international medical organizations are more credible than general research papers. The semantic matching degree is calculated based on the similarity between the query intention and the database elements to ensure that the most relevant data is returned first. For example, in the financial field, when a user queries "the impact of the latest interest rate hike policy on the market", the system may obtain multiple research reports, but will give priority to showing the latest analysis issued by the official financial regulatory agency rather than general market prediction data.
[0187] Finally, the system generates query results based on the sorted data. The query results not only include a list of standardized database elements but may also contain the knowledge paths derived during the query process, enabling users to understand the reasoning process of the query results and trace the data sources. For example, in the medical field, when querying "the latest treatment plan for diabetes", the system may return a path such as "diabetes → insulin treatment → latest research in 2023" and provide the corresponding medical guidelines, research papers, and experimental data for this path for users to further consult and verify. This query method not only ensures the accuracy of the results but also enhances the transparency and interpretability of the query data, enabling the query system to meet the professional needs of different fields.
[0188] This embodiment improves the accuracy, logical relevance, and traceability of the query through multi-dimensional knowledge graph reasoning based on bidirectional mapping relationships. It can combine explicit constraint conditions and implicit constraint conversion parameters to intelligently identify the query target and dynamically optimize the query conditions, ensuring that the query results not only meet the user's needs but also deeply mine potential associated information.
[0189] In one embodiment, a multi-source knowledge processing and querying device is provided, and the multi-source knowledge processing and querying device corresponds one-to-one with the multi-source knowledge processing and querying method in the above embodiment. Referring to Figure 3 , Figure 3 FIG. Figure 3 is a schematic diagram of functional modules of a preferred embodiment of the multi-source knowledge processing and querying device of the present invention. Distributed data acquisition module 10, feature extraction module 20, knowledge graph construction module 30, semantic indexing module 40, query parsing module 50, and inference engine module 60. The detailed description of each functional module is as follows:
[0190] The distributed data acquisition module 10 is configured to obtain multi-source heterogeneous data through the distributed data acquisition module, perform cleaning processing and standardization processing on the multi-source heterogeneous data, and generate a standardized database based on the processed multi-source heterogeneous data;
[0191] The feature extraction module 20 is configured to define domain core concepts and association relationships, and extract knowledge elements from the multi-source heterogeneous data through the feature extraction module;
[0192] The knowledge graph construction module 30 is configured to construct a multi-dimensional knowledge graph based on the core concepts, association relationships, and knowledge elements;
[0193] The semantic indexing module 40 is configured to establish a semantic index based on the multi-dimensional knowledge graph, perform semantic label annotation on the elements in the standardized database according to the semantic index, and establish a bidirectional mapping relationship between the semantic labels of the elements and the nodes of the multi-dimensional knowledge graph;
[0194] The query parsing module 50 is configured to obtain a query intention and extract query constraint conditions from the query intention;
[0195] The inference engine module 60 is configured to perform association inference in the multi-dimensional knowledge graph according to the query constraint conditions and the bidirectional mapping relationship to generate a query result.
[0196] In one embodiment, the distributed data acquisition module 10 is specifically configured to:
[0197] Deploy a plurality of distributed data acquisition nodes, and each distributed data acquisition node selects a corresponding acquisition strategy according to the type of the data source to be acquired;
[0198] Execute a data acquisition task on a plurality of data sources to be acquired through the plurality of distributed data acquisition nodes based on the corresponding acquisition strategies to obtain the multi-source heterogeneous data;
[0199] Perform cleaning processing on the multi-source heterogeneous data to remove duplicate, missing, and invalid data, and generate cleaned data;
[0200] Perform standardization processing on the cleaned data to convert data in different formats into standard data in a unified format;
[0201] Construct the standardized database based on the standard data in the unified format.
[0202] In one embodiment, the feature extraction module 20 is specifically configured to:
[0203] Define a set of core concepts in the domain according to the sources and characteristics of the multi-source heterogeneous data;
[0204] Determine the inheritance relationship, causal relationship, and belonging relationship between the core concepts in the set of core concepts to generate a domain ontology structure including association constraints;
[0205] Extract entity, attribute, and relationship triples from the text of the multi-source heterogeneous data through the named entity recognition module, and generate structured text knowledge elements based on the entity, attribute, and relationship triples;
[0206] Perform region segmentation and feature extraction on the images of the multi-source heterogeneous data through the image semantic segmentation module to generate image semantic feature elements;
[0207] Perform spatial mapping association between the structured text knowledge elements and the image semantic feature elements through the cross-modal alignment module to generate fused knowledge elements;
[0208] Verify the logical consistency of the fused knowledge elements based on the association constraints of the domain ontology structure to filter out the knowledge elements with conflicting relationships in the fused knowledge elements and generate standardized knowledge elements.
[0209] In one embodiment, the knowledge graph construction module 30 is specifically configured to:
[0210] Define the multi-dimensional coordinate axes of the multi-dimensional knowledge graph, where the multi-dimensional coordinate axes include a time dimension axis, a space dimension axis, and a logical dimension axis;
[0211] Map the time attribute of the knowledge elements to the time dimension axis, the space attribute to the space dimension axis, and the logical attribute to the logical dimension axis to generate knowledge nodes including three-dimensional coordinate labels;
[0212] Establish a weighted causal association edge between the knowledge nodes according to the causal relationship in the association relationship;
[0213] Construct an inheritance tree structure with the core concept as the root node according to the inheritance relationship in the association relationship;
[0214] Establish a belonging association edge between the knowledge nodes according to the belonging relationship in the association relationship, and mark the belonging time range and space range;
[0215] Perform conflict detection on the causal association edges, inheritance tree structure, and associated edges. If a conflict relationship is detected among the same set of nodes, determine and delete the conflicting association edges according to the data source priority of the knowledge elements to generate a verified multi-dimensional knowledge graph.
[0216] In one embodiment, the semantic indexing module 40 is specifically configured to:
[0217] Extract semantic feature vectors from the nodes of the multi-dimensional knowledge graph, where the semantic feature vectors include node names, node attribute sets, and associated edge types;
[0218] Construct a hierarchical semantic index based on the semantic feature vectors;
[0219] Perform similarity matching between the text paragraphs and image regions in the standardized database and the hierarchical semantic index to generate fine-grained semantic tags;
[0220] Store reverse pointers in the nodes of the multi-dimensional knowledge graph, where the reverse pointers point to the corresponding text paragraphs or image regions in the standardized database;
[0221] Establish a two-way hash mapping relationship between the fine-grained semantic tags and the nodes of the multi-dimensional knowledge graph.
[0222] In one embodiment, the query parsing module 50 is specifically configured to:
[0223] Obtain a query statement, and perform intent recognition on the query statement through the natural language parsing module to generate a query intent object including a core entity set, a relationship predicate set, and implicit semantics;
[0224] Extract explicit constraint conditions from the query intent object, where the explicit constraint conditions include entity names in the core entity set, relationship types in the relationship predicate set, and node attribute conditions defined in the multi-dimensional knowledge graph;
[0225] Determine implicit constraint conditions according to the implicit semantics of the query intent object;
[0226] Logically fuse the explicit constraint conditions and the implicit constraint conditions to generate a structured query constraint condition set.
[0227] In one embodiment, the inference engine module 60 is specifically configured to:
[0228] Parse the target type identifier, explicit constraint condition list, implicit constraint conversion parameter, and priority weight in the structured query constraint condition set;
[0229] Map the entity names in the explicit constraint condition list to the initial node set in the multi-dimensional knowledge graph according to the forward hash table in the two-way mapping relationship;
[0230] Based on the target type identifier, perform an inference operation in the multi-dimensional knowledge graph starting from the initial node set to generate an intermediate result set;
[0231] According to the relationship type in the explicit constraint condition list and the logical association strength threshold, time range, and space range in the implicit constraint conversion parameter, perform node filtering and edge pruning on the intermediate result set to generate a set of valid association paths;
[0232] Based on the target node or target relationship information in the set of valid association paths, query and obtain the corresponding standardized database elements through the reverse hash table in the two-way mapping relationship, and sort the standardized database elements according to the priority weight to generate the query result.
[0233] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a multi-source knowledge processing and query method.
[0234] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a multi-source knowledge processing and query method
[0235] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0236] Collect audio data of the target domain and extract timbre features from the audio data of the target domain;
[0237] Based on the timbre features, train and generate a domain feature timbre generation model;
[0238] Obtain target text data and perform semantic parsing on the target text data to identify the emotional information of the target text data;
[0239] Adjust the basic speech synthesis parameters according to the emotional information to generate an emotion adaptation parameter set;
[0240] Obtain personalized information and construct a personalized parameter mapping table based on the personalized information;
[0241] Fuse the emotion adaptation parameter set with the personalized parameter mapping table to generate a synthesis control parameter sequence;
[0242] Align the synthesis control parameter sequence with the associated text annotation, target domain visual element features, and background music rhythm data on the time axis to establish a parameter modality binding relationship table;
[0243] Drive the domain feature timbre generation model based on the parameter modality binding relationship table to generate multi-modal synthesis speech data.
[0244] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0245] Collect audio data of the target domain and extract timbre features from the audio data of the target domain;
[0246] Based on the timbre features, train and generate a domain feature timbre generation model;
[0247] Obtain target text data and perform semantic parsing on the target text data to identify the emotional information of the target text data;
[0248] Adjust the basic speech synthesis parameters according to the emotional information to generate an emotion adaptation parameter set;
[0249] Obtain personalized information and construct a personalized parameter mapping table based on the personalized information;
[0250] Fuse the emotion adaptation parameter set with the personalized parameter mapping table to generate a synthesis control parameter sequence;
[0251] Time-align the synthetic control parameter sequence with the associated text annotation, the visual element features of the target domain, and the background music rhythm data to establish a parameter-modal binding relationship table;
[0252] Drive the domain feature timbre generation model based on the parameter-modal binding relationship table to generate multi-modal synthetic speech data.
[0253] It should be noted that for the functions or steps that can be achieved by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.
[0254] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to the memory, storage, database, or other media used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0255] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0256] It should be noted that if there are software tools or components of other companies in the embodiments of this application, they are only used for example introduction and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A multi-source knowledge processing and query method, characterized in that: The following steps are involved: Acquire multi-source heterogeneous data through a distributed data acquisition module, clean and standardize the multi-source heterogeneous data, and generate a standardized database based on the processed multi-source heterogeneous data; Defining core concepts and associations in the domain, and extracting knowledge elements from the multi-source heterogeneous data through a feature extraction module; Construct a multidimensional knowledge graph based on the core concepts, associations and knowledge elements; Establishing a semantic index based on the multidimensional knowledge graph, annotating the elements in the standardized database with semantic labels according to the semantic index, and establishing a bidirectional mapping relationship between the semantic labels of the elements and the nodes of the multidimensional knowledge graph; Obtaining query intent, and extracting query constraints from the query intent; According to the query constraint conditions and the bidirectional mapping relationship, associative reasoning is performed in the multidimensional knowledge graph to generate query results.
2. The multi-source knowledge processing and query method according to claim 1, characterized in that: Acquire multi-source heterogeneous data through a distributed data acquisition module, clean and standardize the multi-source heterogeneous data, and generate a standardized database based on the processed multi-source heterogeneous data, including: Deploy multiple distributed data collection nodes, each of which selects a corresponding collection strategy based on the type of data source to be collected; The plurality of distributed data collection nodes execute data collection tasks on the plurality of data sources to be collected based on corresponding collection strategies to obtain the multi-source heterogeneous data; Cleaning the multi-source heterogeneous data to remove duplicate, missing and invalid data and generate cleaned data; Performing standardization processing on the cleaned data to convert data in different formats into standard data in a unified format; The standardized database is constructed based on the standard data in the unified format.
3. The multi-source knowledge processing and query method according to claim 1, characterized in that: Define the core concepts and relationships in the domain, and extract knowledge elements from the multi-source heterogeneous data through the feature extraction module, including: According to the sources and characteristics of the multi-source heterogeneous data, a set of core concepts in the field is defined; Determine the inheritance relationship, causal relationship and belonging relationship between the core concepts in the core concept set to generate a domain ontology structure including association constraints; Extracting entity, attribute and relationship triples from the text of the multi-source heterogeneous data through a named entity recognition module, and generating structured text knowledge elements based on the entity, attribute and relationship triples; Performing region segmentation and feature extraction on the image of the multi-source heterogeneous data through an image semantic segmentation module to generate image semantic feature elements; The structured text knowledge elements are spatially mapped and associated with the image semantic feature elements through a cross-modal alignment module to generate fused knowledge elements; The logical consistency of the fused knowledge elements is verified based on the associated constraints of the domain ontology structure, so as to filter out the knowledge elements with conflicting relationships in the fused knowledge elements and generate standardized knowledge elements.
4. The multi-source knowledge processing and query method according to claim 1, characterized in that: A multi-dimensional knowledge graph is constructed based on the core concepts, associations and knowledge elements, including: Define the multidimensional coordinate axes of the multidimensional knowledge graph, wherein the multidimensional coordinate axes include a time dimension axis, a space dimension axis, and a logic dimension axis; Mapping the time attribute of the knowledge element to the time dimension axis, mapping the space attribute to the space dimension axis, and mapping the logic attribute to the logic dimension axis, to generate a knowledge node containing a three-dimensional coordinate label; According to the causal relationship in the association relationship, a causal association edge with a weight mark is established between the knowledge nodes; According to the inheritance relationship in the association relationship, an inheritance tree structure with the core concept as the root node is constructed; According to the belonging relationship in the association relationship, a belonging association edge is established between the knowledge nodes, and the belonging time range and space range are marked; Conflict detection is performed on the causal association edges, inheritance tree structure and the associated edges. If a conflict relationship is detected between the same group of nodes, the conflicting association edges are determined and deleted based on the data source priority of the knowledge elements to generate a verified multidimensional knowledge graph.
5. The multi-source knowledge processing and query method according to claim 1, characterized in that: Establishing a semantic index based on the multidimensional knowledge graph, annotating the elements in the standardized database with semantic labels according to the semantic index, and establishing a bidirectional mapping relationship between the semantic labels of the elements and the nodes of the multidimensional knowledge graph, including: Extracting semantic feature vectors from the nodes of the multidimensional knowledge graph, wherein the semantic feature vectors include node names, node attribute sets, and associated edge types; Building a hierarchical semantic index based on the semantic feature vector; Performing similarity matching between the text paragraphs and image regions in the standardized database and the hierarchical semantic index to generate fine-grained semantic labels; Storing a reverse pointer in a node of the multidimensional knowledge graph, the reverse pointer pointing to a corresponding text paragraph or image area in the standardized database; A bidirectional hash mapping relationship is established between the fine-grained semantic tags and the nodes of the multidimensional knowledge graph.
6. The multi-source knowledge processing and query method according to claim 1, characterized in that: Obtaining query intent and extracting query constraints from the query intent, including: Obtaining a query statement, and performing intent recognition on the query statement through a natural language parsing module to generate a query intent object including a core entity set, a relationship predicate set, and implicit semantics; Extracting explicit constraints from the query intent object, wherein the explicit constraints include entity names in the core entity set, relationship types in the relationship predicate set, and node attribute conditions defined in the multidimensional knowledge graph; Determining implicit constraints according to the implicit semantics of the query intent object; The explicit constraint conditions are logically integrated with the implicit constraint conditions to generate a structured query constraint condition set.
7. The multi-source knowledge processing and query method according to claim 1, characterized in that: According to the query constraint condition and the bidirectional mapping relationship, associative reasoning is performed in the multidimensional knowledge graph to generate query results, including: Parsing the target type identifier, the explicit constraint list, the implicit constraint conversion parameter and the priority weight in the structured query constraint set; According to the forward hash table in the bidirectional mapping relationship, mapping the entity names in the explicit constraint list to the initial node set in the multidimensional knowledge graph; Based on the target type identifier, perform reasoning operations on the multidimensional knowledge graph with the initial node set as the starting point to generate an intermediate result set; According to the relationship type in the explicit constraint condition list and the logical association strength threshold, time range and space range in the implicit constraint conversion parameter, the intermediate result set is subjected to node filtering and edge pruning to generate a valid association path set; Based on the target node or target relationship information in the valid association path set, the corresponding standardized database elements are queried and acquired through the reverse hash table in the bidirectional mapping relationship, and the standardized database elements are sorted according to the priority weights to generate the query result.
8. A multi-source knowledge processing and query device, characterized in that: The multi-source knowledge processing and query device comprises: A distributed data acquisition module is used to acquire multi-source heterogeneous data through the distributed data acquisition module, clean and standardize the multi-source heterogeneous data, and generate a standardized database based on the processed multi-source heterogeneous data; A feature extraction module is used to define core concepts and associations in a domain, and to extract knowledge elements from the multi-source heterogeneous data through the feature extraction module; A knowledge graph construction module, used to construct a multi-dimensional knowledge graph based on the core concepts, associations and knowledge elements; A semantic indexing module, used to establish a semantic index based on the multidimensional knowledge graph, annotate the elements in the standardized database with semantic labels according to the semantic index, and establish a bidirectional mapping relationship between the semantic labels of the elements and the nodes of the multidimensional knowledge graph; A query parsing module, used to obtain query intent and extract query constraints from the query intent; The inference engine module is used to perform associative reasoning in the multidimensional knowledge graph according to the query constraints and the bidirectional mapping relationship to generate query results.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and a multi-source knowledge processing and query program stored in the memory and executable on the processor. When the multi-source knowledge processing and query program is executed by the processor, the steps of the multi-source knowledge processing and query method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The storage medium stores a multi-source knowledge processing and query program, which, when executed by a processor, implements the steps of the multi-source knowledge processing and query method as described in any one of claims 1-7.
Citation Information
Cited By
Safety control method based on knowledge graph
CN120671683A
Multi-source data federated governance method and system for vocational education
CN120705235A
Question answering method, device and equipment based on geological knowledge map, medium and product
CN120723881A
Question answering method and device based on geological knowledge graph, equipment, medium and product
CN120723881B
Knowledge graph storage and query method and device based on graph database
CN120723947A