Knowledge retrieval method and device based on multi-source data fusion, electronic equipment and medium

CN120670606BActive Publication Date: 2026-09-08PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510772092.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2026-09-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

但是,这种方法依赖的知识图谱仅包含文本数据,使得知识图谱缺乏对不同语言版本数据,以及与文本相关的图像、音频等数据的利用,并且传统的知识图谱也难以准确理解不同数据来源知识的语义差异,导致知识检索的准确性较差

Benefits of technology

[0057] The knowledge retrieval method, apparatus, electronic device, and medium proposed in this application, which integrates multi-source data, firstly acquires and extracts features of target multi-source data containing target text, target images, target audio, and at least two languages. This captures semantic information from different modalities and languages, providing rich data support for subsequent knowledge graph construction. Secondly, it links target image entities, target audio entities, and language entities with target text entities to obtain target knowledge entities. This enables the association of multi-modal and multi-language entities, achieving cross-modal and cross-language knowledge integration, and connecting target knowledge entities with target text entities. This method integrates knowledge with predefined ontology knowledge data, preserving the structured nature of knowledge while dynamically expanding it with new knowledge. This helps address the shortcomings of traditional knowledge graph methods in utilizing data from different language versions and text-related images and audio, as well as the difficulty of accurately understanding the semantic differences between knowledge from different data sources. This approach helps improve the accuracy of knowledge retrieval. Finally, knowledge retrieval of retrieval indication information from the target knowledge graph constructed based on the integrated knowledge data enables knowledge retrieval through a knowledge graph that integrates multimodal and at least two language data, thus improving the accuracy of knowledge retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670606B_ABST
    Figure CN120670606B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of multi-source data fusion knowledge retrieval method and device, electronic equipment and medium, belong to artificial intelligence technical field.The method comprises: the target image entity, target audio entity and language entity corresponding to multi-source data are carried out entity link with target text entity, obtain target knowledge entity, and the target knowledge entity and target text entity relationship are carried out knowledge fusion with predefined ontology knowledge data, and knowledge retrieval is carried out to search instruction information in the target knowledge graph constructed based on fusion knowledge data.The embodiment of the application can carry out knowledge retrieval from the knowledge graph that is fused with multi-modal and at least two kinds of language data by constructing the knowledge graph containing multi-source data, improve the accuracy of knowledge retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a knowledge retrieval method and apparatus based on multi-source data fusion, an electronic device and a medium. Background Art

[0002] Traditional knowledge retrieval methods usually perform node matching between keywords queried by a user and entity nodes in a pre-constructed knowledge graph, and return triple information associated with the nodes. In a Buddhist application scenario, when a user queries Prajna, a pre-constructed knowledge graph is retrieved to obtain entity nodes of Prajna and Buddhist concepts, and the triple information that Prajna belongs to a Buddhist concept is returned. However, the knowledge graph relied on by this method only contains text data, so that the knowledge graph lacks the utilization of data in different language versions and data such as images and audios related to the text. In addition, traditional knowledge graphs also have difficulty in accurately understanding the semantic differences of knowledge from different data sources, resulting in poor accuracy of knowledge retrieval. Therefore, how to improve the accuracy of knowledge retrieval has become an urgent problem to be solved. Summary of the Invention

[0003] The main purpose of the embodiments of the present application is to provide a knowledge retrieval method and apparatus based on multi-source data fusion, an electronic device and a medium, which aims to improve the accuracy of knowledge retrieval.

[0004] To achieve the foregoing objective, a first aspect of the embodiments of the present application provides a knowledge retrieval method based on multi-source data fusion, the method includes:

[0005] in response to a knowledge graph construction request, acquiring target multi-source data; wherein the target multi-source data includes target text, target images, target audio and at least two kinds of language data;

[0006] performing feature extraction on the target multi-source data to obtain multi-source data features;

[0007] performing knowledge extraction on the multi-source data features to obtain multi-source knowledge data; wherein the multi-source knowledge data includes target text entities, target image entities, target audio entities, language entities and relationships between target text entities;

[0008] performing entity linking on the target image entities, the target audio entities and the language entities with the target text entities to obtain target knowledge entities;

[0009] performing knowledge fusion on the target knowledge entities and the relationships between target text entities with predefined ontology knowledge data to obtain fused knowledge data;

[0010] constructing a target knowledge graph based on the fused knowledge data;

[0011] The retrieval instruction information is obtained, and knowledge retrieval is performed from the target knowledge graph to obtain the target retrieval knowledge information.

[0012] In some embodiments, linking the target image entity, the target audio entity, and the language entity with the target text entity to obtain a target knowledge entity includes:

[0013] The similarity between the target image entity and the target text entity is calculated to obtain the image-text entity similarity. Based on the image-text entity similarity, the target image entity and the target text entity are linked to obtain the image-text linked entity.

[0014] The similarity between the target audio entity and the target text entity is calculated to obtain the audio-text entity similarity. Based on the audio-text entity similarity, the target audio entity and the target text entity are linked to obtain the audio-text linked entity.

[0015] The language entity and the target text entity are aligned across languages ​​to obtain a language text link entity.

[0016] The target knowledge entity is determined based on the image text link entity, the audio text link entity, and the language text link entity.

[0017] In some embodiments, the multi-source knowledge data further includes multi-source data temporal entities and multi-source data spatial location entities;

[0018] The step of determining the target knowledge entity based on the image text link entity, the audio text link entity, and the language text link entity includes:

[0019] The time entities of the multi-source data are subjected to time entity embedding processing to obtain time entity features, the spatial location entities of the multi-source data are subjected to spatial location entity embedding processing to obtain spatial location entity features, and the target text entity is subjected to entity embedding processing to obtain text entity features.

[0020] The time entity features, spatial location entity features, and text entity features are weighted and fused to obtain fused spatiotemporal entity features;

[0021] Based on the fused spatiotemporal entity features, the multi-source data time entity, the multi-source data spatial location entity, and the target text entity are fused to obtain a spatiotemporal text entity. The target knowledge entity is then determined based on the image text link entity, the audio text link entity, the language text link entity, and the spatiotemporal text entity.

[0022] In some embodiments, the multi-source knowledge data further includes target text attributes; the ontology knowledge data includes ontology entities, ontology attributes, and ontology entity relationships.

[0023] The step of fusing the target knowledge entity and the target text entity relationship with predefined ontology knowledge data to obtain fused knowledge data includes:

[0024] The target knowledge entity and the ontology entity are aligned to obtain an aligned entity.

[0025] The target text attribute and the ontology attribute are aligned to obtain the alignment attribute.

[0026] The target text entity relationship is aligned with the ontology entity relationship to obtain the aligned entity relationship;

[0027] The fused knowledge data is determined based on the alignment entity, the alignment attribute, and the alignment entity relationship.

[0028] In some embodiments, constructing the target knowledge graph based on the fused knowledge data includes:

[0029] An initial knowledge graph is constructed based on the fused knowledge data;

[0030] Data update detection is performed on the target multi-source data to obtain updated multi-source data;

[0031] Based on the updated multi-source data, updated multi-source knowledge data is determined, and the updated multi-source knowledge data is scored for knowledge value to obtain an updated knowledge value score.

[0032] Based on the updated knowledge value score, the target knowledge entity is updated to obtain the updated knowledge entity;

[0033] Based on the updated knowledge entity, the text entity relationship is updated to obtain the updated entity relationship;

[0034] The fused knowledge data is updated based on the updated knowledge entities and the relationships between the updated entities to obtain updated fused knowledge data;

[0035] The initial knowledge graph is updated based on the updated and fused knowledge data to obtain the target knowledge graph.

[0036] In some embodiments, the multi-source data features include target text features, target image features, target audio features, and language features;

[0037] The step of extracting features from the target multi-source data to obtain multi-source data features includes:

[0038] The target text is subjected to text semantic feature extraction to obtain the target text features;

[0039] The target image is subjected to image object detection to obtain image detection objects, and the image detection objects are subjected to object semantic feature extraction to obtain the target image features;

[0040] The target audio is converted into audio text, and audio text features are extracted from the audio text to obtain the target audio features;

[0041] The language features are obtained by extracting linguistic semantic features from the at least two types of language data.

[0042] In some embodiments, after fusing the target knowledge entity and the target text entity relationship with predefined ontology knowledge data to obtain fused knowledge data, the method further includes:

[0043] Anomaly verification is performed on the fused knowledge data to obtain anomaly verification data;

[0044] The fused knowledge data is subjected to integrity verification to obtain integrity verification data;

[0045] The fused knowledge data is subjected to consistency verification to obtain consistency verification data;

[0046] The fused knowledge data is optimized based on the anomaly verification data, the integrity verification data, and the consistency verification data to obtain optimized fused knowledge data, and the optimized fused knowledge data is used as the fused knowledge data.

[0047] To achieve the above objectives, a second aspect of this application proposes a knowledge retrieval device based on multi-source data fusion, the device comprising:

[0048] A multi-source data acquisition module is used to acquire target multi-source data in response to a knowledge graph construction request; wherein, the target multi-source data includes target text, target image, target audio, and at least two language data;

[0049] The feature extraction module is used to extract features from the target multi-source data to obtain multi-source data features;

[0050] The knowledge extraction module is used to extract knowledge from the features of the multi-source data to obtain multi-source knowledge data; wherein, the multi-source knowledge data includes target text entities, target image entities, target audio entities, language entities, and relationships between target text entities;

[0051] The entity linking module is used to link the target image entity, the target audio entity, and the language entity with the target text entity to obtain the target knowledge entity;

[0052] The knowledge fusion module is used to fuse the target knowledge entity and the target text entity relationship with predefined ontology knowledge data to obtain fused knowledge data.

[0053] The knowledge graph construction module is used to construct a target knowledge graph based on the fused knowledge data;

[0054] The knowledge retrieval module is used to obtain retrieval indication information and perform knowledge retrieval from the target knowledge graph to obtain target retrieval knowledge information.

[0055] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0056] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method of the first aspect described above.

[0057] The knowledge retrieval method, apparatus, electronic device, and medium proposed in this application, which integrates multi-source data, firstly acquires and extracts features of target multi-source data containing target text, target images, target audio, and at least two languages. This captures semantic information from different modalities and languages, providing rich data support for subsequent knowledge graph construction. Secondly, it links target image entities, target audio entities, and language entities with target text entities to obtain target knowledge entities. This enables the association of multi-modal and multi-language entities, achieving cross-modal and cross-language knowledge integration, and connecting target knowledge entities with target text entities. This method integrates knowledge with predefined ontology knowledge data, preserving the structured nature of knowledge while dynamically expanding it with new knowledge. This helps address the shortcomings of traditional knowledge graph methods in utilizing data from different language versions and text-related images and audio, as well as the difficulty of accurately understanding the semantic differences between knowledge from different data sources. This approach helps improve the accuracy of knowledge retrieval. Finally, knowledge retrieval of retrieval indication information from the target knowledge graph constructed based on the integrated knowledge data enables knowledge retrieval through a knowledge graph that integrates multimodal and at least two language data, thus improving the accuracy of knowledge retrieval. Attached Figure Description

[0058] Figure 1 This is a flowchart of the knowledge retrieval method based on multi-source data fusion provided in the embodiments of this application;

[0059] Figure 2 yes Figure 1 The flowchart of step S102 in the document;

[0060] Figure 3 yes Figure 1 The flowchart of step S104 in the process;

[0061] Figure 4 yes Figure 3 The flowchart of step S304 in the process;

[0062] Figure 5 yes Figure 1 The flowchart of step S105 in the process;

[0063] Figure 6 This is another flowchart of the knowledge retrieval method based on multi-source data fusion provided in the embodiments of this application;

[0064] Figure 7 yes Figure 1 The flowchart of step S106 in the process;

[0065] Figure 8 This is a schematic diagram of the structure of the knowledge retrieval device with multi-source data fusion provided in the embodiments of this application;

[0066] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0068] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0069] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0070] First, let's analyze some of the terms used in this application:

[0071] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0072] This application provides a knowledge retrieval method, apparatus, electronic device, and medium based on multi-source data fusion, aiming to improve the accuracy of knowledge retrieval.

[0073] The knowledge retrieval method, apparatus, electronic device, and medium based on multi-source data fusion provided in this application are specifically described through the following embodiments. First, the knowledge retrieval method based on multi-source data fusion in this application is described.

[0074] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0075] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0076] The knowledge retrieval method based on multi-source data fusion provided in this application relates to the field of artificial intelligence technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the knowledge retrieval method based on multi-source data fusion, but is not limited to the above forms.

[0077] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0078] Figure 1 This is an optional flowchart of the knowledge retrieval method based on multi-source data fusion provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S107.

[0079] Step S101: In response to the knowledge graph construction request, obtain target multi-source data; wherein, target multi-source data includes target text, target image, target audio and at least two language data.

[0080] Step S102: Extract features from the target multi-source data to obtain multi-source data features.

[0081] Step S103: Extract knowledge from the features of the multi-source data to obtain multi-source knowledge data; wherein, the multi-source knowledge data includes target text entities, target image entities, target audio entities, language entities and target text entity relationships.

[0082] Step S104: Link the target image entity, target audio entity, and language entity with the target text entity to obtain the target knowledge entity.

[0083] Step S105: The target knowledge entity and the target text entity relationship are fused with the predefined ontology knowledge data to obtain fused knowledge data.

[0084] Step S106: Construct the target knowledge graph based on the fused knowledge data.

[0085] Step S107: Obtain retrieval instruction information and perform knowledge retrieval on the retrieval instruction information from the target knowledge graph to obtain target retrieval knowledge information.

[0086] Steps S101 to S107, as illustrated in this embodiment, firstly, by acquiring and extracting features from target multi-source data containing target text, target image, target audio, and at least two languages, semantic information from different modalities and languages ​​can be captured, providing rich data support for subsequent knowledge graph construction. Secondly, target image entities, target audio entities, and language entities are linked with target text entities to obtain target knowledge entities. This enables the association of multimodal and multilingual entities, achieving cross-modal and cross-lingual knowledge integration, and the relationship between target knowledge entities and target text entities is linked with a predetermined... Knowledge fusion of ontological knowledge data preserves the structured nature of knowledge while dynamically expanding new knowledge. This helps address the shortcomings of traditional knowledge graph methods in utilizing data from different language versions and text-related images and audio, as well as the difficulty of accurately understanding the semantic differences between knowledge from different data sources. This approach helps improve the accuracy of knowledge retrieval. Finally, knowledge retrieval of retrieval indications from a target knowledge graph constructed based on fused knowledge data improves the accuracy of knowledge retrieval by integrating multimodal and at least two language data.

[0087] In step S101 of some embodiments, specifically, the target multi-source data can be a comprehensive data set including target text, target image, target audio and at least two language data, and the multi-source data can be a comprehensive data set in the field of Buddhism.

[0088] Specifically, the target text can be various types of Buddhist texts, including but not limited to Buddhist classics, commentaries on classics, and academic papers by Buddhist scholars.

[0089] Specifically, the target image can be visual materials in the field of Buddhism, including but not limited to images extracted from murals, sculptures, paintings, and printed Buddhist texts.

[0090] Specifically, the target audio can be audio data in the field of Buddhism, including but not limited to audio of Buddhist lectures, recitations, and Buddhist music.

[0091] Specifically, language data refers to Buddhist data in various languages, including but not limited to Buddhist data in Sanskrit, Pali, Tibetan, and Chinese. This language data can be different language data obtained from at least one modality of data, such as target text, target image, and target audio.

[0092] Specifically, in the digitization process of paper-based Buddhist data, a high-resolution, high-speed document scanner is first used to quickly scan the paper-based Buddhist data to obtain complete text and image information. Image quality is improved through preprocessing techniques such as image enhancement, noise reduction, and tilt correction. Then, optical character recognition (OCR) technology, combined with a professional vocabulary database of Buddhist data and contextual semantic analysis, accurately identifies the text in the paper-based Buddhist data, especially special fonts and ancient characters, and converts them into editable text data. Subsequently, the text recognized by OCR is structured according to the ontological structure of Buddhist knowledge, extracting chapters, paragraphs, sentences, and key knowledge elements from the paper-based Buddhist data, and importing them into an electronic Buddhist database to achieve digital storage and management of the paper-based Buddhist data.

[0093] Specifically, target multi-source data can be collected from various websites, digital libraries, and social media platforms using web crawling technology and digital acquisition equipment (such as cameras, drones, and microphones).

[0094] In this embodiment, target multi-source data is obtained in response to the knowledge graph construction request. The target multi-source data includes target text, target image, target audio and data in at least two languages. This avoids the situation where the knowledge system of the constructed knowledge graph is incomplete due to insufficient collection of data in different language versions, art paintings, sculptures, music and other non-text data. It provides rich and multi-dimensional data support for the subsequent construction of the knowledge graph.

[0095] Please see Figure 2 In some embodiments, the multi-source data features include target text features, target image features, target audio features, and language features. Step S102 includes, but is not limited to, steps S201 to S204:

[0096] Step S201: Extract text semantic features from the target text to obtain target text features.

[0097] Step S202: Perform image object detection on the target image to obtain the image detection object, and extract the semantic features of the image detection object to obtain the target image features.

[0098] Step S203: Convert the target audio into audio text, and extract audio text features from the audio text to obtain the target audio features.

[0099] Step S204: Extract language semantic features from at least two types of language data to obtain language features.

[0100] In step S201 of some embodiments, specifically, the multi-source data features can be a set of semantic features extracted from target text, target image, target audio, and at least two language data, specifically including target text features, target image features, target audio features, and language features.

[0101] Specifically, target text features refer to the text semantic feature vectors extracted from the target text.

[0102] Specifically, the BERT model can be used to perform word embedding on the target text to convert it into text vectors and extract the semantic features of the text vectors.

[0103] For example, in Buddhist applications, word embedding can be performed on the Buddhist concept of "Prajnaparamita" to extract its semantic vector.

[0104] In step S202 of some embodiments, specifically, the target image feature refers to the visual semantic feature vector extracted from the target image.

[0105] Specifically, target detection calculations can be performed on target images using faster region-based convolutional neural networks (such as Faster R-CNN, FasterRegion-based Convolutional Neural Network) to extract key objects from the target images.

[0106] For example, in a mural depicting a scene of the Buddha, Faster R-CNN can be used to identify image detection objects such as portraits, thrones, and disciples.

[0107] Furthermore, deep residual networks (such as ResNet) can be used to encode the features of the image detection objects and extract the semantic feature vectors of the image detection objects.

[0108] For example, for the identified human figures, semantic features of the abhaya mudra (gesture of fearlessness) of the subject's hand gestures, as well as features such as the style and color of the subject's clothing, can be extracted.

[0109] In step S203 of some embodiments, specifically, the target audio feature is an audio semantic feature vector extracted from the audio text.

[0110] Specifically, the target audio can be converted into audio text using Automatic Speech Recognition (ASR) technology.

[0111] For example, an audio recording of Buddhist chanting in Sanskrit can be converted into the corresponding Sanskrit text using ASR technology.

[0112] Specifically, the BERT model can be used to perform word embedding on audio text in order to extract semantic feature vectors from the audio text.

[0113] For example, extracting semantic features of "Prajnaparamita" from Sanskrit texts.

[0114] In step S204 of some embodiments, specifically, semantic features refer to semantic feature vectors extracted from different language data that can characterize language content.

[0115] Specifically, cross-lingual models (such as XLM-RoBERTa, Cross-lingual LanguageModel-RoBERTa) can be used to compare and learn words from different languages ​​to reduce the cross-lingual distance of words with the same concept, so as to map words from different languages ​​to the same vector space and calculate the semantic similarity of words from different languages. Combined with attention alignment technology, words with the same concept in different languages ​​are aligned to generate an alignment matrix and its corresponding attention weights. Based on the alignment matrix and attention weights, the expression form of the same concept word in different languages ​​is determined.

[0116] In this embodiment, by extracting linguistic semantic features from at least two language data, the same semantic information in different languages ​​can be transformed into unified semantic features, which facilitates the subsequent construction of a knowledge graph that can understand and process multilingual data.

[0117] Through steps S201 to S204, semantic features are extracted from multi-source data, which can comprehensively capture key information in different modalities and languages. This helps to comprehensively consider textual semantic information, visual information, audio information and multilingual information when constructing the knowledge graph, and improves the ability of the constructed knowledge graph to understand and process different data sources.

[0118] In step S103 of some embodiments, specifically, multi-source knowledge data refers to structured knowledge data extracted from multi-source data features. This multi-source knowledge data may include, but is not limited to, target text entities, target image entities, target audio entities, language entities, multi-source data time entities, multi-source data spatial location entities, target text entity relationships, target image entity relationships, target audio entity relationships, language entity relationships, and target text attributes.

[0119] Specifically, target text entities, target image entities, target audio entities, linguistic entities and the like in multi-source data features can be extracted through named entity recognition technology.

[0120] For example, in the field of Buddhism, when the target text feature is "Prajnaparamita", text entities such as "Prajna" and "Paramita" can be identified; for target image features, target image entities such as "lotus position" and "dhyana mudra" can be identified; for target audio features, rhythmic entities such as "serene" and "solemn" can be identified; for linguistic features, Sanskrit entities such as "Prajna" and "meditation" can be identified.

[0121] Further, dependency syntactic analysis can be performed on multimodal entities and multilingual entities to extract entity relationships between entities, and the entity relationships can be synonymous relationships, hypernym-hyponym relationships, causal relationships, guiding relationships, inclusion relationships and the like between entities.

[0122] For example, the entity relationship between text entities such as "Prajna" and "Paramita" is the guiding relationship where "Prajna" guides "Paramita"; the entity relationship between the image entities "lotus position" and "dhyana mudra" is a concomitant relationship; the entity relationship between rhythmic entities such as "serene" and "solemn" is a synergistic relationship; the entity relationship between Sanskrit entities such as "Prajna" and "meditation" is an inclusion relationship.

[0123] In this embodiment, knowledge extraction on multi-source data features can extract entities and entity relationships from the multi-source data features, further providing comprehensive and abundant data support for the subsequent construction of structured knowledge data including multiple modalities and multiple languages.

[0124] Please refer to Figure 3 , in some embodiments, step S104 includes but is not limited to steps S301 to S304:

[0125] Step S301: calculating the similarity between the target image entities and the target text entities to obtain image-text entity similarity, and performing image-text entity linking on the target image entities and the target text entities based on the image-text entity similarity to obtain image-text linked entities.

[0126] Step S302: calculating the similarity between the target audio entities and the target text entities to obtain audio-text entity similarity, and performing audio-text entity linking on the target audio entities and the target text entities based on the audio-text entity similarity to obtain audio-text linked entities.

[0127] Step S303: performing cross-language alignment processing on the linguistic entities and the target text entities to obtain linguistic-text linked entities.

[0128] Step S304, determining a target knowledge entity based on image-text linked entities, audio-text linked entities and language-text linked entities.

[0129] In step S301 of some embodiments, specifically, performing cosine similarity calculation on the target image entity and the target text entity can determine the image-text entity similarity.

[0130] For example, cosine similarity calculation is performed on the image entity of the "Abhaya Mudra" gesture extracted from a Buddha statue mural and the text entity of "Abhaya-giver", and an image-text entity similarity of 0.9 is obtained.

[0131] Specifically, an image-text linked entity refers to a unified entity representation formed after associating a target image entity with a target text entity.

[0132] For example, if the image-text entity similarity of 0.9 is higher than a set threshold (e.g., 0.85), then the "Abhaya Mudra" gesture is linked with "Abhaya-giver", and the formed image-text linked entity is "Abhaya Mudra-Abhaya-giver".

[0133] In step S302 of some embodiments, specifically, performing cosine similarity calculation on the target audio entity and the target text entity can determine the audio-text entity similarity.

[0134] For example, cosine similarity calculation is performed on the "serene" prosodic entity extracted from Buddhist chanting audio and the "compassionate" text entity, and an audio-text entity similarity of 0.86 is obtained.

[0135] Specifically, an audio-text linked entity refers to a unified entity representation formed after associating a target audio entity with a target text entity.

[0136] For example, if the audio-text entity similarity of 0.86 is higher than a set threshold (e.g., 0.85), then "serene" is linked with "compassionate", and the formed audio-text linked entity is "serene-compassionate".

[0137] In step S303 of some embodiments, specifically, a language-text linked entity refers to a unified entity representation formed after associating a language entity with a target text entity.

[0138] Specifically, different language entities and the target text entity can be mapped to the same vector space by using a cross-lingual model (e.g., XLM-RoBERTa), so as to calculate the cosine similarity between the language entities and the text entities, and perform entity linking on the language entities and the target text entities according to the cosine similarity.

[0139] For example, cosine similarity calculation is performed on the Sanskrit entity "一切皆空" (everything is void) and the Chinese entity "诸法空相" (all dharmas are empty in nature), a linguistic text entity similarity of 0.95 is obtained, and if the linguistic text entity similarity of 0.95 is higher than a set threshold (e.g., 0.85), the Sanskrit entity "一切皆空" and the Chinese entity "诸法空相" are linked, forming a linguistic text linked entity as "一切皆空 (Sanskrit) - 诸法空相 (Chinese)".

[0140] See Figure 4 , in some embodiments, the multi-source knowledge data further includes multi-source data time entities and multi-source data spatial location entities, step S304 includes but is not limited to steps S401 to S403:

[0141] Step S401: performing time entity embedding processing on the multi-source data time entities to obtain time entity features, performing spatial location entity embedding processing on the multi-source data spatial location entities to obtain spatial location entity features, and performing entity embedding processing on target text entities to obtain text entity features.

[0142] Step S402: performing weighted fusion on the time entity features, the spatial location entity features and the text entity features to obtain fused space-time entity features.

[0143] Step S403: performing entity fusion on the multi-source data time entities, the multi-source data spatial location entities and the target text entities based on the fused space-time entity features to obtain space-time text entities, and determining target knowledge entities based on image-text linked entities, audio-text linked entities, linguistic text linked entities and the space-time text entities.

[0144] In step S401 of some embodiments, specifically, the multi-source data time entities can be creation years of Buddhist works, construction time of Buddhist statues, etc.

[0145] Specifically, the multi-source data spatial location entities can be excavation sites of Buddhist statues, performance venues of Buddhist music, etc.

[0146] Specifically, entity embedding processing can be performed on the multi-source data time entities and the target text entities through a BERT model, so as to convert time information and text information into feature vectors.

[0147] For example, for a historical event in Buddhist documents, such as "the Buddhist figure attained enlightenment at Bodh Gaya in 528 BC", the time entity "528 BC" is converted into a time entity feature vector through word embedding processing.

[0148] For example, "Buddha" in the description of the target text entity is converted into a text entity feature vector through word embedding processing.

[0149] Furthermore, geographic information systems (GIS) can be used to encode spatial location entities from multi-source data to transform geographic location information into feature vectors.

[0150] For example, to determine the geographical location of Buddhist murals, geographic coordinate entities (such as latitude and longitude entities) can be obtained. These entities can then be converted into low-dimensional vectors using GIS, thereby transforming geographic location information into spatial location entity features.

[0151] In step S402 of some embodiments, specifically, fusing spatiotemporal entity features refers to an entity feature vector that integrates time entity features, spatial location entity features, and text entity features.

[0152] Specifically, the fused spatiotemporal entity features can be obtained by acquiring the temporal weight of temporal entity features, the spatial weight of spatial entity features, and the text weight of text entity features, and by adding the product of temporal entity features and temporal weights, the product of spatial entity features and spatial weights, and the product of the entity feature and text weights.

[0153] For example, by weighting and fusing the characteristics of "528 BC", "geographical coordinates of Buddhist murals" and "Buddha", a comprehensive spatiotemporal entity feature is obtained.

[0154] In step S403 of some embodiments, specifically, multi-source data temporal entities, multi-source data spatial location entities, and target text entities can be fused using graph convolutional networks (GCNs).

[0155] For example, by integrating "528 BC", "geographical coordinates of Buddhist murals", and the entity of "Buddha", a unified spatiotemporal text entity is obtained: "Buddha attained enlightenment in Bodh Gaya in 528 BC".

[0156] Furthermore, by combining image text link entities (such as "Abhaya Mudra - the one who bestows fearlessness"), audio text link entities (such as "peace and harmony - compassion"), and language text link entities (such as "all is emptiness Sanskrit - the emptiness of all phenomena Chinese"), it is possible to identify target knowledge entities that contain the above-mentioned link entities.

[0157] It provides comprehensive multimodal and multilingual knowledge for the construction of knowledge graphs.

[0158] Through steps S401 to S403, by embedding multi-source data time entities and multi-source data spatial location entities, and fusing them with text entities, a comprehensive knowledge representation containing time, space, and semantic information can be constructed. This enriches the content of the knowledge graph and helps the subsequently constructed knowledge graph to more accurately represent the time and place of historical events, thus supporting more complex historical and cultural research.

[0159] Through steps S301 to S304, entity linking of multimodal and multilingual data can solve the problem of the lack of utilization of different language versions of data and text-related image, audio and other data in the subsequent knowledge graph construction. It also improves the ability of the subsequent knowledge graph construction to understand and process different data sources, so as to better support knowledge retrieval in a multilingual environment and thus improve the accuracy of subsequent knowledge retrieval.

[0160] Please see Figure 5 In some embodiments, the multi-source knowledge data also includes target text attributes; the ontology knowledge data includes ontology entities, ontology attributes, and ontology entity relationships. Step S105 includes, but is not limited to, steps S501 to S504:

[0161] Step S501: Perform entity alignment processing on the target knowledge entity and the ontology entity to obtain the aligned entity.

[0162] Step S502: Align the target text attributes with the ontology attributes to obtain the aligned attributes.

[0163] Step S503: Align the target text entity relationship with the ontology entity relationship to obtain the aligned entity relationship.

[0164] Step S504: Determine the fused knowledge data based on the aligned entities, aligned attributes, and aligned entity relationships.

[0165] In step S501 of some embodiments, specifically, ontology knowledge data refers to structured knowledge data in the target domain, which may include ontology entities, which refer to basic concepts or objects in the target domain knowledge system.

[0166] For example, in the field of Buddhism, the essential entity can be "meditation" or "wisdom".

[0167] Specifically, knowledge fusion between target knowledge entities and ontology entities can be achieved through a pre-constructed Buddhist ontology model. This model comprises a core concept layer, an attribute and relationship layer, and an instance data layer. The core concept layer defines ontology entities (such as figures, concepts, and texts within the Buddhist field); the attribute and relationship layer defines ontology attributes (such as the Buddhist school to which a figure belongs) and entity relationships (such as the spatiotemporal relationship between Buddhist sculptures and texts); and the instance data layer instantiates and concretizes specific multi-source data with ontology knowledge within the ontology model, mapping entities, attributes, and relationships from multi-source data onto ontology knowledge to form concrete knowledge fusion instances.

[0168] Specifically, ontology attributes can be described using the Web Ontology Language (OWL) to semantically annotate ontology entities, thereby defining entity relationships or attributes between ontology entities.

[0169] Specifically, the semantic similarity between the target knowledge entity and the ontology entity can be calculated through the instance data layer, and the target knowledge entity and the ontology entity can be aligned based on the semantic similarity.

[0170] For example, the target knowledge entity "wisdom" is a Buddhist concept entity after multi-source data fusion. There is also an entity "prajna" in the ontology knowledge data. Although they have different names, semantic analysis reveals that "prajna" is usually interpreted as "wisdom" in Buddhism. It can be determined that the two entities are different manifestations of the same concept. Through entity alignment processing, "prajna" and "wisdom" are aligned, that is, mapped to the same concept, thus obtaining the aligned entity "prajna-wisdom".

[0171] In this embodiment, by performing entity alignment between the target knowledge entity and the ontology entity, entities from different data sources can be integrated into a unified knowledge system through ontology mapping, which can effectively eliminate semantic differences and conceptual conflicts between different data sources.

[0172] In step S502 of some embodiments, specifically, the target text attribute refers to the structured features used to describe the specific meaning of the target text entity.

[0173] Specifically, ontology knowledge data also includes ontology attributes, which are structured features used to describe the specific meaning of ontology entities within a predefined ontology knowledge system.

[0174] Specifically, alignment attributes refer to structured features that combine target text attributes and ontology attributes.

[0175] Specifically, the cosine similarity between the target text attribute and the ontology attribute can be calculated through the instance data layer, and the target text attribute and the ontology attribute can be aligned based on the cosine similarity.

[0176] For example, the target text entity "Alaya-vijnana" has the target text attribute "possessing the three meanings of being able to store, being stored, and holding", and the ontological attributes are "being able to store", "being stored", and "holding". Through similarity calculation, it is ensured that being able to store, being stored, and holding are accurately mapped to the ontological model to obtain the aligned attributes "being able to store", "being stored", and "holding".

[0177] In step S503 of some embodiments, specifically, ontology entity relationship refers to the structured relationship that defines and describes the logical or semantic connection between ontology entities in the ontology knowledge system.

[0178] Specifically, the semantic similarity between the target text entity relationship and the ontology entity relationship can be calculated through the instance data layer, and the target text entity relationship and the ontology entity relationship can be aligned based on the semantic similarity of the relationship.

[0179] For example, if the target text entity relationship is the entity relationship between "Alaya-vijnana" and "ignorance" extracted from Buddhist literature, namely "Alaya-vijnana is the root of ignorance", and the ontological entity relationship has a causal relationship of "ignorance leads to attachment of Alaya-vijnana", by aligning the entity relationships, it can be confirmed that the two are semantically related and logically consistent. Therefore, aligning the target text entity relationship "Alaya-vijnana is the root of ignorance" with the ontological entity relationship "ignorance leads to attachment of Alaya-vijnana" yields the aligned entity relationship "causal relationship".

[0180] In step S504 of some embodiments, specifically, the fused knowledge data refers to knowledge system data that includes aligned entities, aligned attributes, and relationships between aligned entities.

[0181] For example, integrated knowledge data can be represented as: entity "Alaya consciousness", attributes "storehouse, stored, and holding", and inclusion relationship "Alaya consciousness has the three meanings of storehouse, stored, and holding".

[0182] In this embodiment, the fusion knowledge data is determined based on the alignment entity, alignment attribute, and alignment entity relationship. This enables the effective integration of knowledge elements from multi-source data with elements in a predefined ontology knowledge system, generating a structured and systematic knowledge system.

[0183] Through steps S501 to S504, by aligning multi-source data with ontology knowledge data through entity alignment, attribute alignment, and relation alignment, knowledge elements in multi-source data can be effectively associated with elements in the predefined ontology knowledge system. This makes the data after knowledge fusion more accurate and complete, which helps to integrate the aligned elements into a unified knowledge system, forming a structured knowledge graph that resolves semantic differences and conceptual conflicts between different sources of data, thus improving the accuracy of subsequent knowledge retrieval.

[0184] In one optional embodiment of this application, when defining ontology knowledge data, the quality of ontology knowledge data is ensured by using annotation guidelines, and the quality of ontology knowledge data is ensured by combining cross-review within the group to verify the format compliance, entity integrity and basic logic of the annotated data.

[0185] For example, in the field of Buddhist studies, the data that needs to be labeled includes Buddhist entities (such as Buddhist concepts, figures, and scriptures) and relationships within Buddhist entities (such as teacher-student relationships and interpretations). Labeling must adhere to detailed guidelines, including tagging systems and boundary rules. For instance, the "Four Noble Truths" must be fully labeled as suffering, its origin, its cessation, and the path to its cessation. This is combined with cross-review within the group to verify the format compliance, entity completeness, and basic logic of the labeled data. For example, checking whether the "Twelve Links of Dependent Origination" are fully labeled and whether the dates when Buddhist figures discovered Buddhist scriptures are correct. Furthermore, language-level review focuses on the accuracy of labeled terminology, grammatical correctness, and disambiguation handling, such as for Sanskrit. It is uniformly translated as "emptiness," etc.

[0186] Please see Figure 6 In some embodiments, the knowledge retrieval method based on multi-source data fusion also includes, but is not limited to, steps S601 to S604:

[0187] Step S601: Perform anomaly verification on the fused knowledge data to obtain anomaly verification data.

[0188] Step S602: Perform integrity verification on the fused knowledge data to obtain integrity verification data.

[0189] Step S603: Perform consistency verification on the fused knowledge data to obtain consistency verification data.

[0190] Step S604: Optimize the fused knowledge data based on the anomaly verification data, integrity verification data, and consistency verification data to obtain optimized fused knowledge data, and use the optimized fused knowledge data as the fused knowledge data.

[0191] In step S601 of some embodiments, specifically, a preset knowledge verification rule can be learned through an LSTM model to perform anomaly verification on the fused knowledge data.

[0192] For example, in the field of Buddhism, an LSTM model can be used to perform anomaly checks on the "Four Noble Truths" in fused knowledge data to determine whether the "Four Noble Truths" has been incorrectly written as "Four Holy Emperors," or to perform anomaly checks on the birth and death years of Buddhist figures, such as 702 AD, to determine that the birth and death years of the Buddhist figures were from 602 AD to 664 AD. In addition, cross-validation can be used to verify the source of the cited Buddhist data, such as if there are differences between the content recorded in the cited classics in the Buddhist data and the Buddhist data itself.

[0193] In this embodiment, by performing anomaly verification on the fused knowledge data, errors in the fused knowledge data can be detected and corrected in a timely manner, thereby improving the accuracy and reliability of the data.

[0194] In step S602 of some embodiments, the integrity of the fused knowledge data can be verified by methods such as field fill rate monitoring, association integrity check, and cross-source completion.

[0195] For example, in the integration of Buddhist knowledge, the fill rate of core Buddhist fields (such as concepts, figures, and classics) can be monitored by the field fill rate to ensure that the missing rate does not exceed 2%. The completeness of the relationship between entities can be checked by querying the knowledge system path, such as whether the classic school is correctly associated with Buddhist figures. Furthermore, when missing integrated knowledge data is found, the missing data can be supplemented by cross-source completion methods, such as calling the Buddhist classic API to supplement Pali literature.

[0196] In step S603 of some embodiments, specifically, consistency verification of fused knowledge data can be performed through methods such as spatiotemporal logic verification, cross-version comparison, and relationship conflict detection.

[0197] Specifically, we can identify the consistency of time logic in the fused knowledge data by constructing an event timeline, and analyze whether the fused knowledge data in different language versions is consistent by cross-version comparison. Furthermore, we can detect whether there are relationship conflicts in the fused knowledge data by combining relationship conflict detection.

[0198] For example, an event timeline can be constructed to ensure that the year of composition of Buddhist scriptures is no earlier than the year of death of the corresponding Buddhist figures. By comparing the same Buddhist scriptures in Sanskrit, Chinese and Tibetan versions, the similarity of the expressions in the scriptures can be ensured to exceed 0.75, and the core concepts of different schools of Buddhism can be further examined to see if they are incorrectly marked as completely identical.

[0199] In this embodiment, by performing consistency verification on the fused knowledge data, the logical consistency of the fused knowledge data can be ensured, avoiding the chaos of the knowledge system caused by logical errors or relationship conflicts.

[0200] In step S604 of some embodiments, specifically, optimizing the fused knowledge data refers to the fused knowledge data after correcting abnormal data, supplementing missing data, and adjusting inconsistent data.

[0201] For example, incorrectly labeled terms or contradictory years in the fused knowledge data are corrected, missing Pali documents are supplemented based on integrity verification data, and incorrect relational annotations or logical conflicts are adjusted based on consistency verification data.

[0202] Through steps S601 to S604, through anomaly verification, integrity verification, and consistency verification, errors, omissions, and inconsistencies in the fused knowledge data can be comprehensively identified and corrected, ensuring the accuracy and integrity of the data. This helps to build a more accurate knowledge graph in the future, thereby improving the accuracy of knowledge retrieval.

[0203] Please see Figure 7 In some embodiments, step S106 includes, but is not limited to, steps S701 to S707:

[0204] Step S701: Construct an initial knowledge graph based on the fused knowledge data.

[0205] Step S702: Perform data update detection on the target multi-source data to obtain updated multi-source data.

[0206] Step S703: Based on the updated multi-source data, determine the updated multi-source knowledge data, and score the knowledge value of the updated multi-source knowledge data to obtain the updated knowledge value score.

[0207] Step S704: Update the target knowledge entity based on the updated knowledge value score to obtain the updated knowledge entity.

[0208] Step S705: Update the entity relations of the text entity relations based on the updated knowledge entities to obtain the updated entity relations.

[0209] Step S706: Update the fused knowledge data based on the updated knowledge entities and updated entity relationships to obtain updated fused knowledge data.

[0210] Step S707: Update the initial knowledge graph based on the updated and fused knowledge data to obtain the target knowledge graph.

[0211] In step S701 of some embodiments, specifically, entities in the fused knowledge data are used as nodes, attributes are used as labels for nodes, and entity relationships are used as edges connecting two nodes, so as to organize all the fused knowledge data into a graph structure and form an initial knowledge graph.

[0212] In this embodiment, an initial knowledge graph is constructed based on fused knowledge data, which can present complex knowledge data in an intuitive graphical form, providing a data foundation for subsequent knowledge updates and knowledge retrieval.

[0213] In step S702 of some embodiments, specifically, it can be determined whether the multi-source data has been updated by detecting whether the timestamp of the target text has changed, or whether an image has been added to the image database, or whether the timestamp of the audio file has changed. If the text, image, or audio data has changed, the changed data is marked as updated multi-source data.

[0214] In this embodiment, by performing data update detection on the target multi-source data, changes in the multi-source data can be detected in a timely manner, ensuring that the constructed knowledge graph can reflect the latest knowledge information.

[0215] In step S703 of some embodiments, specifically, the knowledge value score is used to assess the importance and contribution of updating multi-source data to the knowledge graph.

[0216] Specifically, the importance and contribution of updated multi-source data to the knowledge graph can be evaluated using a knowledge value scoring model. This model includes multiple evaluation dimensions, such as innovativeness, authority, and completeness. Innovativeness assesses whether the knowledge data provides new perspectives or discoveries; authority assesses the credibility of the knowledge sources; and completeness assesses whether the knowledge data comprehensively covers the relevant topics.

[0217] For example, in research papers on Buddhist scriptures, the updated multi-source knowledge data can be scored by analyzing the citation rate and the proportion of technical terms in the research paper. If the citation rate is 60% and the density of technical terms exceeds a set threshold (such as 70%), the research paper can obtain a knowledge value score of 65 points.

[0218] In this embodiment, updated multi-source knowledge data is determined based on updated multi-source data, and the knowledge value of the updated multi-source knowledge data is scored. This helps to identify and prioritize updated data that makes significant contributions to the knowledge graph, ensuring that the knowledge graph prioritizes high-value updated data and improving the quality and usability of the knowledge graph.

[0219] In step S704 of some embodiments, specifically, high-value update data can be filtered based on the updated knowledge value score to update the target knowledge entity in the knowledge graph. For example, update data with a knowledge value score exceeding a certain threshold (e.g., 0.8) can be selected. For example, if a newly discovered Buddha image displays new gesture features, it may be necessary to update the attributes of the "Buddha image" entity and add a new gesture description.

[0220] In this embodiment, by performing data update detection on the target multi-source data, it can be ensured that the entities in the knowledge graph are always up-to-date and most accurate, which helps to improve the accuracy and reliability of the knowledge graph.

[0221] In step S705 of some embodiments, specifically, updating entity relationships refers to the updated relationships between entities.

[0222] For example, after the knowledge graph reasoning engine detects the update of the "Buddhist figure" entity, it can identify the entity relationship related to "gesture" and adjust the entity relationship according to the updated entity information. If the updated gesture is related to "hand gesture mudra", then the relationship between "Buddhist figure" and "hand gesture mudra" will be used as the updated entity relationship.

[0223] In step S706 of some embodiments, specifically, the updated knowledge entities (such as Buddhist figures) and updated entity relationships (such as the relationship between "Buddhist figures" and "hands together in prayer") can be integrated into the fused knowledge data, replacing the original fused knowledge data.

[0224] In step S707 of some embodiments, specifically, graph embedding technology can be used to convert the updated knowledge nodes and nodes in the initial knowledge graph into low-dimensional vector representations, calculate the cosine similarity between the vectors, identify the semantic relationship between the updated knowledge nodes and existing entity nodes, expand the relationship between the updated knowledge nodes and existing nodes through the path sorting algorithm, generate updated entity relationship edges, and add the updated entity relationship edges to the initial knowledge graph to form the target knowledge graph.

[0225] For example, when updating the knowledge node "Theravada Buddhist Thought", the update entity relationship edge between "Theravada Buddhist Thought" and "Yogacara" is determined and marked as "related relationship". The "difference relationship" edge between "Theravada Buddhist Thought" and "Mahayana Buddhist Thought" is determined by the path sorting algorithm, and the "difference relationship" edge connects the "Theravada Buddhist Thought" node and the "Yogacara" node.

[0226] In one optional embodiment of this application, when updating the knowledge map, OWL technology is also used to perform logical verification on the updated knowledge relationships in the updated and fused knowledge data.

[0227] For example, if there is a logical conflict between the updated knowledge relation and the existing axioms, an updated knowledge relation branch is created using OWL technology to store the updated knowledge relation in the branch, while keeping the existing axioms unchanged to ensure the logical consistency of the constructed knowledge graph; if there is no conflict between the updated knowledge relation and the existing axioms, an RDF triple update transaction is generated to add the updated entity relation to the initial knowledge graph.

[0228] In this embodiment, the initial knowledge graph is updated based on the updated and fused knowledge data. This allows for the dynamic expansion of the knowledge graph according to the updated knowledge entities, ensuring the completeness and timeliness of the final generated knowledge graph. Furthermore, by combining logical verification of the relationships between updated entities, conflicts between updated knowledge entities and existing knowledge can be resolved, ensuring the logical consistency and completeness of the final constructed knowledge graph.

[0229] Through steps S701 to S707, an initial knowledge graph is constructed and the updates of target multi-source data are continuously monitored. Based on the value scores of the updated multi-source knowledge data, the entities and their relationships in the initial knowledge graph are accurately updated to ensure that the knowledge graph is always up-to-date and the most accurate. The knowledge graph is also dynamically expanded according to the updated knowledge entities to ensure the completeness and timeliness of the final generated knowledge graph.

[0230] In step S107 of some embodiments, specifically, the retrieval instruction information refers to the query request entered by the user, which usually appears in the form of keywords, phrases or natural language questions, and is used to specify the target content to be retrieved.

[0231] For example, in a Buddhist knowledge graph, a user might enter "the definition of Alaya-vijnana" as a search term.

[0232] Specifically, target retrieval knowledge information refers to knowledge content related to retrieval indication information.

[0233] Specifically, word embedding processing is performed on the retrieval indication information to extract the key semantic feature vector of the retrieval indication information, and knowledge retrieval is performed in the target knowledge graph to quickly locate the target retrieval node corresponding to the key semantics, and all other nodes associated with the target retrieval node and the relationships between nodes are used as target retrieval knowledge information.

[0234] For example, if a user enters "definition of Alaya-vijnana", by searching in the target knowledge graph, the node related to "Alaya-vijnana" can be located, and the attribute information of the node can be extracted, such as "the three meanings of the storehouse, the stored, and the holding".

[0235] In this embodiment, retrieval instruction information is obtained, and knowledge retrieval of the retrieval instruction information is performed from the target knowledge graph to obtain target retrieval knowledge information. This can quickly and accurately respond to the user's query request, and provide knowledge content that is highly relevant to the retrieval instruction information from the constructed target knowledge graph. Furthermore, since the target knowledge graph is a knowledge system that integrates data from different language versions, as well as text-related image, audio, and other data, the accuracy of knowledge retrieval is significantly improved.

[0236] This application first acquires and extracts features from target multi-source data, including target text, target images, target audio, and data in at least two languages. This captures semantic information from different modalities and languages, providing rich data support for subsequent knowledge graph construction. Second, it links target image entities, target audio entities, and language entities with target text entities to obtain target knowledge entities. This enables the association of multimodal and multilingual entities, achieving cross-modal and cross-lingual knowledge integration. Furthermore, it fuses target knowledge entities and target text entity relationships with predefined ontology knowledge data, maintaining the structured nature of knowledge while dynamically expanding new knowledge. This helps address the shortcomings of traditional knowledge graph methods in utilizing different language versions of data and text-related image and audio data, and the difficulty of accurately understanding the semantic differences between knowledge from different data sources. This improves the accuracy of subsequent knowledge retrieval. Finally, it performs knowledge retrieval of retrieval indication information from the target knowledge graph constructed based on fused knowledge data. This allows for knowledge retrieval through a knowledge graph that integrates multimodal and at least two language data, improving the accuracy of knowledge retrieval.

[0237] Please see Figure 8 This application also provides a knowledge retrieval device based on multi-source data fusion, which can implement the above-mentioned knowledge retrieval method based on multi-source data fusion. The device includes:

[0238] The multi-source data acquisition module is used to acquire target multi-source data in response to knowledge graph construction requests; wherein, the target multi-source data includes target text, target image, target audio and at least two language data;

[0239] The feature extraction module is used to extract features from the target multi-source data to obtain multi-source data features;

[0240] The knowledge extraction module is used to extract knowledge from the features of multi-source data to obtain multi-source knowledge data; among which, multi-source knowledge data includes target text entities, target image entities, target audio entities, language entities, and relationships between target text entities;

[0241] The entity linking module is used to link target image entities, target audio entities, and language entities with target text entities to obtain target knowledge entities;

[0242] The knowledge fusion module is used to fuse target knowledge entities and target text entity relationships with predefined ontology knowledge data to obtain fused knowledge data.

[0243] The knowledge graph construction module is used to construct target knowledge graphs based on fused knowledge data;

[0244] The knowledge retrieval module is used to obtain retrieval indication information and perform knowledge retrieval on the retrieval indication information from the target knowledge graph to obtain the target retrieval knowledge information.

[0245] The specific implementation of this multi-source data fusion knowledge retrieval device is basically the same as the specific implementation of the multi-source data fusion knowledge retrieval method described above, and will not be repeated here.

[0246] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned knowledge retrieval method based on multi-source data fusion. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0247] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0248] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0249] The memory 902 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 902 can store the processing system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 to execute the multi-source data fusion knowledge retrieval method of the embodiments of this application.

[0250] The 903 input / output interface is used to implement information input and output.

[0251] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0252] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);

[0253] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0254] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned knowledge retrieval method based on multi-source data fusion.

[0255] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0256] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0257] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0258] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0259] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0260] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0261] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0262] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. The coupling or direct coupling or communication connection between the shown or discussed units may be through some interfaces, or indirect coupling or communication connection between the apparatus or units, and may be electrical, mechanical, or other forms.

[0263] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0264] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0265] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0266] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A knowledge retrieval method based on multi-source data fusion, characterized in that, The method includes: In response to a knowledge graph construction request, target multi-source data is acquired; wherein, the target multi-source data includes target text, target image, target audio, and at least two language data; Feature extraction is performed on the target multi-source data to obtain multi-source data features; Knowledge extraction is performed on the features of the multi-source data to obtain multi-source knowledge data; wherein, the multi-source knowledge data includes target text entities, target image entities, target audio entities, language entities, and relationships between target text entities; The target image entity, the target audio entity, and the language entity are linked with the target text entity to obtain the target knowledge entity; The target knowledge entity and the target text entity relationship are fused with predefined ontology knowledge data to obtain fused knowledge data. Construct a target knowledge graph based on the fused knowledge data; Obtain retrieval indication information, and perform knowledge retrieval from the target knowledge graph to obtain target retrieval knowledge information; The step of linking the target image entity, the target audio entity, and the language entity with the target text entity to obtain the target knowledge entity includes: The similarity between the target image entity and the target text entity is calculated to obtain the image-text entity similarity. Based on the image-text entity similarity, the target image entity and the target text entity are linked to obtain the image-text linked entity. The similarity between the target audio entity and the target text entity is calculated to obtain the audio-text entity similarity. Based on the audio-text entity similarity, the target audio entity and the target text entity are linked to obtain the audio-text linked entity. The language entity and the target text entity are aligned across languages ​​to obtain a language text link entity. The target knowledge entity is determined based on the image text link entity, the audio text link entity, and the language text link entity.

2. The method according to claim 1, characterized in that, The multi-source knowledge data also includes multi-source data time entities and multi-source data spatial location entities; The step of determining the target knowledge entity based on the image text link entity, the audio text link entity, and the language text link entity includes: The time entities of the multi-source data are subjected to time entity embedding processing to obtain time entity features, the spatial location entities of the multi-source data are subjected to spatial location entity embedding processing to obtain spatial location entity features, and the target text entity is subjected to entity embedding processing to obtain text entity features. The time entity features, spatial location entity features, and text entity features are weighted and fused to obtain fused spatiotemporal entity features; Based on the fused spatiotemporal entity features, the multi-source data time entity, the multi-source data spatial location entity, and the target text entity are fused to obtain a spatiotemporal text entity. The target knowledge entity is then determined based on the image text link entity, the audio text link entity, the language text link entity, and the spatiotemporal text entity.

3. The method according to claim 1, characterized in that, The multi-source knowledge data also includes target text attributes; the ontology knowledge data includes ontology entities, ontology attributes, and ontology entity relationships. The step of fusing the target knowledge entity and the target text entity relationship with predefined ontology knowledge data to obtain fused knowledge data includes: The target knowledge entity and the ontology entity are aligned to obtain an aligned entity. The target text attribute and the ontology attribute are aligned to obtain the alignment attribute. The target text entity relationship is aligned with the ontology entity relationship to obtain the aligned entity relationship; The fused knowledge data is determined based on the alignment entity, the alignment attribute, and the alignment entity relationship.

4. The method according to claim 1, characterized in that, The construction of the target knowledge graph based on the fused knowledge data includes: An initial knowledge graph is constructed based on the fused knowledge data; Data update detection is performed on the target multi-source data to obtain updated multi-source data; Based on the updated multi-source data, updated multi-source knowledge data is determined, and the updated multi-source knowledge data is scored for knowledge value to obtain an updated knowledge value score. Based on the updated knowledge value score, the target knowledge entity is updated to obtain the updated knowledge entity; Based on the updated knowledge entity, the text entity relationship is updated to obtain the updated entity relationship; The fused knowledge data is updated based on the updated knowledge entities and the relationships between the updated entities to obtain updated fused knowledge data; The initial knowledge graph is updated based on the updated and fused knowledge data to obtain the target knowledge graph.

5. The method according to claim 1, characterized in that, The multi-source data features include target text features, target image features, target audio features, and language features; The step of extracting features from the target multi-source data to obtain multi-source data features includes: The target text is subjected to text semantic feature extraction to obtain the target text features; Image object detection is performed on the target image to obtain the image detection object, and the semantic features of the image detection object are extracted to obtain the target image features; The target audio is converted into audio text, and audio text features are extracted from the audio text to obtain the target audio features; The language features are obtained by extracting semantic features from the at least two language data.

6. The method according to any one of claims 1 to 5, characterized in that, After fusing the target knowledge entity and the target text entity relationship with predefined ontology knowledge data to obtain fused knowledge data, the method further includes: Anomaly verification is performed on the fused knowledge data to obtain anomaly verification data; The fused knowledge data is subjected to integrity verification to obtain integrity verification data; The fused knowledge data is subjected to consistency verification to obtain consistency verification data; The fused knowledge data is optimized based on the anomaly verification data, the integrity verification data, and the consistency verification data to obtain optimized fused knowledge data, and the optimized fused knowledge data is used as the fused knowledge data.

7. A knowledge retrieval device based on multi-source data fusion, characterized in that, The device includes: A multi-source data acquisition module is used to acquire target multi-source data in response to a knowledge graph construction request; wherein, the target multi-source data includes target text, target image, target audio, and at least two language data; The feature extraction module is used to extract features from the target multi-source data to obtain multi-source data features; The knowledge extraction module is used to extract knowledge from the features of the multi-source data to obtain multi-source knowledge data; wherein, the multi-source knowledge data includes target text entities, target image entities, target audio entities, language entities, and relationships between target text entities; The entity linking module is used to link the target image entity, the target audio entity, and the language entity with the target text entity to obtain the target knowledge entity; The knowledge fusion module is used to fuse the target knowledge entity and the target text entity relationship with predefined ontology knowledge data to obtain fused knowledge data. The knowledge graph construction module is used to construct a target knowledge graph based on the fused knowledge data; The knowledge retrieval module is used to obtain retrieval indication information and perform knowledge retrieval from the target knowledge graph to obtain target retrieval knowledge information; The step of linking the target image entity, the target audio entity, and the language entity with the target text entity to obtain the target knowledge entity includes: The similarity between the target image entity and the target text entity is calculated to obtain the image-text entity similarity. Based on the image-text entity similarity, the target image entity and the target text entity are linked to obtain the image-text linked entity. The similarity between the target audio entity and the target text entity is calculated to obtain the audio-text entity similarity. Based on the audio-text entity similarity, the target audio entity and the target text entity are linked to obtain the audio-text linked entity. The language entity and the target text entity are aligned across languages ​​to obtain a language text link entity. The target knowledge entity is determined based on the image text link entity, the audio text link entity, and the language text link entity.

8. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the knowledge retrieval method of multi-source data fusion as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the knowledge retrieval method of multi-source data fusion as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Information search method and device based on large model, electronic equipment and storage medium

    CN119917672A

  • Knowledge-derived search suggestion

    US20220253477A1