Multi-language knowledge graph construction method, system, equipment and medium
By standardizing the multi-source heterogeneous data and semantic annotation, aligning multi-language embedded representations with comparative learning technology, and using incremental update algorithm to update the knowledge graph in real time, the problems of complex semantic analysis, multi-language support and dynamic update in the existing technology are solved, and efficient and accurate knowledge graph construction and update are achieved.
Patent Information
- Application Number
- CN202510028705.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-09
AI Technical Summary
The existing knowledge graph construction technology has problems such as inaccurate semantic labeling, insufficient multilingual support, and inefficient dynamic updates when dealing with complex semantics, multilingual environments and dynamic knowledge updates.
By acquiring multi-source heterogeneous data, standardized processing is performed and mapped to a unified semantic representation space, the embedded representations of different languages are aligned using contrast learning technology, and the knowledge graph is updated in real time based on the incremental update algorithm.
It realizes accurate analysis of complex semantics, efficient alignment of multilingual knowledge and real-time update of dynamic knowledge graphs, significantly improving the system's response speed and knowledge update efficiency, and reducing the cost of manual intervention.
Smart Images

Figure CN119962655A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a multilingual knowledge graph construction method, system, device and medium. Background Art
[0002] Knowledge graph is a technology that represents entities and their relationships in the form of a graphical structure. It is widely used in search engines, question-answering systems, recommendation engines, intelligent customer service, etc. Existing knowledge graph construction technologies mainly rely on artificial rules or traditional machine learning methods. These methods have the following shortcomings when dealing with complex semantics, multilingual environments, and dynamic knowledge updates:
[0003] 1. Limitations of semantic annotation: Existing technologies mainly rely on predefined rules, dictionaries or shallow machine learning models to complete semantic annotation. Due to the static nature of the rules, the system has difficulty processing complex contexts, implicit semantics or dynamically evolving knowledge, and the annotation results are often inaccurate or incomplete.
[0004] 2. Insufficient support for multilingual knowledge graphs: Many knowledge graph construction technologies focus on a single language (such as English). In a multilingual environment, there is a lack of effective semantic mapping and alignment methods for multilingual data. Especially in scenarios with large differences in semantic complexity and language habits, existing technologies are difficult to meet the needs.
[0005] 3. Inefficiency of dynamic updates: Knowledge graphs need to be updated frequently in practical applications to reflect the generation of new knowledge. However, most existing methods rely on manual intervention or regular batch updates, which makes it difficult to achieve real-time response. This inefficiency makes knowledge graphs lack practicality in rapidly changing scenarios. The core of these problems lies in the fact that existing technologies cannot fully utilize the capabilities of generative artificial intelligence, especially in terms of dynamic semantic understanding, multilingual support, and knowledge extraction and updating. There are significant bottlenecks. Summary of the invention
[0006] The purpose of the present invention is to provide a multilingual knowledge graph construction method, system, device and medium to solve the problems existing in the above-mentioned prior art.
[0007] To achieve the above object, the present invention provides a method for constructing a multilingual knowledge graph, comprising:
[0008] Acquire multi-source heterogeneous data, wherein the multi-source heterogeneous data includes text data, voice data, and image data;
[0009] Standardizing the text data, the voice data, and the image data to obtain multi-source heterogeneous data in a unified format;
[0010] Map the multi-source heterogeneous data with unified format into a unified semantic representation space and perform semantic annotation to generate corresponding semantic annotation results;
[0011] Align the embedding representations of different languages based on contrastive learning technology, map the embeddings of multiple languages to a unified semantic space, and update the semantic annotation results;
[0012] A knowledge graph is constructed based on the updated semantic annotation results, and the knowledge graph is updated in real time based on an incremental update algorithm.
[0013] Optionally, the standardization processing of the text data, the voice data and the image data specifically includes:
[0014] The text data is processed by word segmentation, stop word removal and syntactic analysis to obtain preprocessed text data; the voice data is processed by text conversion to obtain preprocessed voice data, and the image data is processed by text recognition to obtain preprocessed image data; by preprocessing the data of each modality, multi-source heterogeneous data with a unified format is obtained.
[0015] Optionally, mapping the multi-source heterogeneous data in a unified format to a unified semantic representation space and performing semantic annotation specifically includes:
[0016] A semantic annotation model is built based on the Transformer architecture. The semantic annotation model is used to map different modal data into a unified semantic representation space. The cosine similarity formula is used to calculate the semantic relevance of each modal data. The semantic annotation results are generated based on the Beam Search technology / Sampling technology combined with the semantic relevance. The semantic annotation model is dynamically optimized based on the incremental learning mechanism.
[0017] Optionally, aligning the embedding representations of different languages based on contrastive learning technology and mapping the embeddings of multiple languages to a unified semantic space specifically includes:
[0018] Based on the SimCSE model, the embedding spaces of different languages are aligned, and the aligned multilingual embedding representation is input into the mBERT model to map the multilingual embeddings into a unified semantic space.
[0019] Optionally, the constructing of a knowledge graph based on the updated semantic annotation results specifically includes:
[0020] Based on named entity recognition technology, key entities are extracted, the key entities are input into a relationship prediction model for prediction and classification, and the relationship between entities is output; wherein the relationship prediction model is built based on the BERT+CRF model.
[0021] A multilingual knowledge graph construction system, comprising:
[0022] A data acquisition module is used to acquire multi-source heterogeneous data, wherein the multi-source heterogeneous data includes text data, voice data and image data; and perform standardization processing on the text data, voice data and image data to obtain the multi-source heterogeneous data in a unified format;
[0023] The semantic annotation module is used to map the multi-source heterogeneous data with unified format into a unified semantic representation space and perform semantic annotation to generate corresponding semantic annotation results;
[0024] The multilingual semantic mapping module is used to align the embedding representations of different languages based on contrastive learning technology, map the embeddings of multiple languages to a unified semantic space, and update the semantic annotation results;
[0025] The dynamic knowledge graph construction module is used to construct a knowledge graph according to the updated semantic annotation results, and to update the knowledge graph in real time based on an incremental update algorithm.
[0026] An electronic device includes a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the multilingual knowledge graph construction method.
[0027] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the multilingual knowledge graph construction method.
[0028] The technical effects of the present invention are:
[0029] The present invention can achieve accurate analysis of complex semantics and efficient annotation of unseen knowledge. By constructing a multilingual knowledge graph system, the present invention can effectively support semantic alignment and mapping of multilingual data and solve the problem of semantic inconsistency between different languages. The present invention realizes dynamic real-time updating of knowledge graphs, significantly improves the response speed of the system and the efficiency of knowledge updating, and reduces the cost of manual intervention. The present invention provides a modular system design, which is convenient for large-scale deployment and adapts to cross-domain knowledge management and application needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0031] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0032] Figure 1 is a construction flow chart in an embodiment of the present invention;
[0033] Figure 2 It is a schematic diagram of the structure of each module in the embodiment of the present invention. DETAILED DESCRIPTION
[0034] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as limiting the present invention, but should be understood as a more detailed description of certain aspects, features, and embodiments of the present invention.
[0035] It should be understood that the terms described in the present invention are only for describing special embodiments and are not intended to limit the present invention. In addition, for the numerical range in the present invention, it should be understood that each intermediate value between the upper and lower limits of the scope is also specifically disclosed. Each smaller range between the intermediate value in any stated value or stated range and any other stated value or intermediate value in the described range is also included in the present invention. The upper and lower limits of these smaller ranges can be independently included or excluded in the scope.
[0036] It will be apparent to those skilled in the art that various modifications and variations may be made to the specific embodiments of the present invention description without departing from the scope or spirit of the present invention. Other embodiments derived from the present invention description will be apparent to those skilled in the art. The present application description and examples are exemplary only.
[0037] The words “include,” “including,” “have,” “contain,” etc. used in this article are open-ended terms, meaning including but not limited to.
[0038] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0039] Embodiment 1
[0040] like Figure 1-Figure 2As shown, this embodiment provides a method for constructing a multilingual knowledge graph, including: obtaining multi-source heterogeneous data, the multi-source heterogeneous data including text data, voice data and image data; standardizing the text data, voice data and image data to obtain multi-source heterogeneous data with a unified format; mapping the multi-source heterogeneous data with a unified format to a unified semantic representation space and performing semantic annotation to generate corresponding semantic annotation results; aligning embedding representations of different languages based on contrastive learning technology, mapping the embeddings of multiple languages to a unified semantic space, and updating the semantic annotation results; constructing a knowledge graph based on the updated semantic annotation results, and updating the knowledge graph in real time based on an incremental update algorithm.
[0041] This embodiment designs a dynamic semantic annotation and multilingual knowledge graph construction system combined with generative AI, and its technical solution includes the following modules and working principles:
[0042] 1. System module design, including:
[0043] (1) Data input module:
[0044] Function: Receive multi-source heterogeneous data, including text, voice, image and other formats; standardize the input data and convert it into the structured form required for semantic analysis.
[0045] Features: Supports multiple data interfaces, compatible with real-time data streams and static data files.
[0046] (2) Generative AI semantic annotation module:
[0047] Function: Based on the context-aware capabilities of generative AI (such as the GPT model), it performs semantic analysis and annotation on input data to generate structured knowledge fragments.
[0048] Implementation: Use pre-trained large-scale language models combined with domain-specific fine-tuning techniques to achieve dynamic understanding and precise annotation of complex semantics.
[0049] Features: Supports multi-layer semantic parsing, including entity recognition, relationship extraction, event detection, and implicit semantic inference.
[0050] (3) Multilingual semantic mapping module:
[0051] Function: Realize semantic alignment and knowledge fusion between different languages, and build a language-independent unified knowledge representation.
[0052] Implementation: Through the multilingual capabilities of generative AI, combined with alignment algorithms and translation technology, multilingual data can be automatically processed for semantic consistency.
[0053] Features: Supports mainstream languages and low-resource languages, and solves the problem of differences in semantic habits.
[0054] (4) Dynamic knowledge graph construction module:
[0055] Function: Based on the semantic annotation results, construct the node and edge relationships of the knowledge graph; update the existing graph in real time and dynamically.
[0056] Implementation: Use incremental update algorithms and real-time knowledge extraction technology to maintain the freshness and integrity of knowledge graph content.
[0057] Features: Modular architecture, supporting distributed storage and efficient retrieval.
[0058] (5) Feedback optimization module:
[0059] Function: Use reinforcement learning (RLHF) and user feedback to continuously optimize the generative AI model to improve the accuracy of semantic annotation and knowledge construction.
[0060] Implementation: Design a dynamic feedback loop and take user input, annotation results, and the actual application effect of the knowledge graph as the optimization goal.
[0061] Features: Adaptive optimization mechanism to improve long-term system performance.
[0062] 2. The system workflow of this embodiment:
[0063] (1) The data input module receives multi-source data provided by the user and completes standardization processing.
[0064] (2) Standardized data enters the generative AI semantic annotation module to generate preliminary structured knowledge fragments.
[0065] (3) The multilingual semantic mapping module performs language alignment and unified knowledge representation processing on the annotation results.
[0066] (4) The dynamic knowledge graph construction module integrates knowledge fragments, updates or constructs the knowledge graph, and provides real-time query and display services.
[0067] (5) The feedback optimization module optimizes the generative AI model based on user feedback and the application effect of the knowledge graph.
[0068] 3. Technical implementation details
[0069] (1) Data processing algorithm: A generative model based on the Transformer architecture that supports batch and real-time input.
[0070] (2) Knowledge fusion method: semantic similarity calculation and logical reasoning technology are used to solve the problems of knowledge duplication and conflict.
[0071] (3) Update mechanism: Based on timestamps and data change logs, an efficient incremental update algorithm is implemented.
[0072] Compared with the prior art, this embodiment has the following significant advantages:
[0073] 1. Dynamicity: Utilize the contextual understanding ability of generative AI to achieve real-time updating of knowledge graphs and reduce human involvement.
[0074] 2. Multilingual adaptability: The multilingual semantic mapping module significantly improves the accuracy and efficiency of knowledge alignment between different languages.
[0075] 3. High degree of automation: The entire process from data input to knowledge graph construction is automated, reducing development and maintenance costs.
[0076] 4. System performance optimization: Continuously optimize the generative AI model through user feedback and reinforcement learning to improve the long-term reliability and accuracy of the system.
[0077] In the overall system architecture of this embodiment, the core algorithm processing process of each module is as follows, which shows in detail the main calculation process and key technologies from data input to knowledge graph construction:
[0078] 1. Data input module, including:
[0079] Preprocessing algorithm: For multi-source data (text, voice, image), specific preprocessing techniques are used:
[0080] For text: word segmentation, stop word removal, and syntactic analysis;
[0081] For speech: speech-to-text (ASR, Automatic Speech Recognition);
[0082] For images: OCR (Optical Character Recognition) technology extracts text information.
[0083] Calculation process:
[0084] Text segmentation: Use Double-Array Trie to speed up word segmentation.
[0085] OCR process: Convolutional neural network (CNN) combined with recurrent neural network (RNN) recognizes text sequences in image areas.
[0086] 2. Generative AI semantic annotation module, including:
[0087] Dynamic semantic annotation based on generative pre-trained models (such as GPT-4, LLaMA):
[0088] For input processing:
[0089] Multimodal embedding: Through the Transformer architecture, the input text, speech, and image are uniformly mapped to the semantic space.
[0090] Semantic similarity is calculated using the cosine similarity formula:
[0091] Sim(A,B)=∑Ai·Bi∑Ai2·∑Bi2\text{Sim}(A,B)=\frac{\sum A_i\cdot B_i}{\sqrt{\sum A_i^2}\cdot\sqrt{\sum B_i^2}}
[0092] The semantic relevance is determined based on the above cosine similarity formula.
[0093] Generate labels: Use beam search or sampling technology to generate the optimal label sequence, and dynamically adjust the labeling results: Dynamically optimize the semantic labeling model through incremental learning.
[0094] 3. Multilingual semantic mapping module, including:
[0095] Cross-language semantic mapping based on contrastive learning and language alignment techniques:
[0096] Contrastive Learning: Use the SimCSE (Simple Contrastive Learning of Sentence Embeddings) model to align the embedding spaces of different languages.
[0097] The loss function is contrast loss:
[0098]
[0099] Among them, τ\tau is the temperature parameter, xix_i and yiy_i are semantic embeddings.
[0100] Cross-lingual alignment: Use alignment models such as mBERT to map embeddings of multiple languages into a unified semantic space.
[0101] 4. Dynamic knowledge graph construction module, including:
[0102] Entity extraction: Use named entity recognition (NER) technology to extract key entities. The algorithm uses the BER T+CRF model, which has higher accuracy.
[0103] Relationship extraction: Use a relationship prediction algorithm based on a graph neural network (GNN) to infer the relationship between entities.
[0104] The loss function is categorical cross entropy:
[0105]
[0106] Among them, yiy_i is the true relationship label, and y^i\hat{y}_i is the predicted value.
[0107] Dynamic update: Based on the incremental graph update algorithm, newly discovered entities and relationships are merged in real time.
[0108] 5. Feedback optimization module, using the policy gradient method of reinforcement learning to optimize the system, including:
[0109] Reward function design: positive scores of user feedback and coverage and accuracy of entity-relation pairs in the knowledge graph.
[0110] Strategy optimization formula:
[0111] Among them, πθ\pi_\theta is the policy network, and Q(s,a)Q(s,a) is the state action value.
[0112] Interactive relationship operation process:
[0113] 1. After the data is preprocessed by the input module, it enters the generative AI module for semantic analysis and annotation.
[0114] 2. The generated labels flow to the semantic mapping module to complete multi-language alignment.
[0115] 3. Entity and relationship information is fed into the knowledge graph construction module to update the knowledge structure.
[0116] 4. User feedback flows back to the optimization module to iteratively improve system performance.
[0117] Each module is tightly integrated through data flow and feedback loops. This architecture achieves highly intelligent and dynamic updating of semantic annotation and multilingual knowledge graph construction through modular division of labor and overall collaborative optimization.
[0118] This embodiment provides a dynamic semantic annotation method, which uses the context-awareness capability of generative AI to achieve accurate parsing of complex semantics and efficient annotation of unseen knowledge.
[0119] This embodiment builds a multilingual knowledge graph system that can effectively support the semantic alignment and mapping of multilingual data and solve the problem of semantic inconsistency between different languages.
[0120] This embodiment can realize dynamic real-time updating of the knowledge graph, significantly improve the system's response speed and knowledge updating efficiency, and reduce the cost of manual intervention.
[0121] This embodiment provides a modular system design to facilitate large-scale deployment and meet cross-domain knowledge management and application requirements.
[0122] An electronic device includes a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the multilingual knowledge graph construction method.
[0123] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the multilingual knowledge graph construction method.
[0124] The above is only a preferred specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for constructing a multilingual knowledge graph, characterized in that: include: Acquire multi-source heterogeneous data, wherein the multi-source heterogeneous data includes text data, voice data, and image data; Standardizing the text data, the voice data, and the image data to obtain multi-source heterogeneous data in a unified format; Map the multi-source heterogeneous data with unified format into a unified semantic representation space and perform semantic annotation to generate corresponding semantic annotation results; Align the embedding representations of different languages based on contrastive learning technology, map the embeddings of multiple languages to a unified semantic space, and update the semantic annotation results; A knowledge graph is constructed based on the updated semantic annotation results, and the knowledge graph is updated in real time based on an incremental update algorithm.
2. A multilingual knowledge graph construction method according to claim 1, characterized in that: The standardization processing of the text data, the voice data and the image data specifically includes: The text data is processed by word segmentation, stop word removal and syntactic analysis to obtain preprocessed text data; the voice data is processed by text conversion to obtain preprocessed voice data, and the image data is processed by text recognition to obtain preprocessed image data; by preprocessing the data of each modality, multi-source heterogeneous data with a unified format is obtained.
3. A multilingual knowledge graph construction method according to claim 1, characterized in that: The mapping of the multi-source heterogeneous data in a unified format into a unified semantic representation space and performing semantic annotation specifically includes: A semantic annotation model is built based on the Transformer architecture. The semantic annotation model is used to map different modal data into a unified semantic representation space. The cosine similarity formula is used to calculate the semantic relevance of each modal data. The semantic annotation results are generated based on the BeamSearch technology / Sampling technology combined with the semantic relevance. The semantic annotation model is dynamically optimized based on the incremental learning mechanism.
4. A multilingual knowledge graph construction method according to claim 1, characterized in that: The method of aligning the embedding representations of different languages based on contrastive learning technology and mapping the embeddings of multiple languages into a unified semantic space specifically includes: Based on the SimCSE model, the embedding spaces of different languages are aligned, and the aligned multilingual embedding representation is input into the mBERT model to map the multilingual embeddings into a unified semantic space.
5. A multilingual knowledge graph construction method according to claim 1, characterized in that: The step of constructing a knowledge graph based on the updated semantic annotation results specifically includes: Based on named entity recognition technology, key entities are extracted, the key entities are input into a relationship prediction model for prediction and classification, and the relationship between entities is output; wherein the relationship prediction model is built based on the BERT+CRF model.
6. A multilingual knowledge graph construction system, characterized in that: include: A data acquisition module, used to acquire multi-source heterogeneous data, wherein the multi-source heterogeneous data includes text data, voice data and image data; Standardizing the text data, the voice data, and the image data to obtain multi-source heterogeneous data in a unified format; The semantic annotation module is used to map the multi-source heterogeneous data with unified format into a unified semantic representation space and perform semantic annotation to generate corresponding semantic annotation results; The multilingual semantic mapping module is used to align the embedding representations of different languages based on contrastive learning technology, map the embeddings of multiple languages to a unified semantic space, and update the semantic annotation results; The dynamic knowledge graph construction module is used to construct a knowledge graph according to the updated semantic annotation results, and to update the knowledge graph in real time based on an incremental update algorithm.
7. An electronic device, characterized in that: It includes a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute a multilingual knowledge graph construction method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements a multilingual knowledge graph construction method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Chinese cross-language event detection method based on graph neural network
CN118069839A
Knowledge graph driven semantic governance scheme dynamic generation method and system
CN118114758A
Cerebral hemorrhage personalized treatment scheme optimization method and system based on big data analysis
CN119153117A
Knowledge map construction method and apparatus, and storage medium
WO2018153266A1
Knowledge extraction method, apparatus, electronic device, and storage medium
WO2021212682A1
Cited By
Vibrating compaction forming energy transfer efficiency evaluation method
CN120597061A
A method for evaluating energy transmission efficiency of vibration compaction forming
CN120597061B
Key attribute extraction method and system based on multi-modal normalization
CN120705539A
Information extraction and dynamic updating method, system, equipment and medium
CN120804455A
File information association analysis method based on heterogeneous knowledge graph
CN121743747A