Multi-language knowledge graph construction method, system, equipment and medium

By standardizing the multi-source heterogeneous data and semantic annotation, aligning multi-language embedded representations with comparative learning technology, and using incremental update algorithm to update the knowledge graph in real time, the problems of complex semantic analysis, multi-language support and dynamic update in the existing technology are solved, and efficient and accurate knowledge graph construction and update are achieved.

CN119962655APending Publication Date: 2025-05-09HUNAN INSTITUTE OF ENGINEERING

Patent Information

Application Number
CN202510028705.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing knowledge graph construction technology has problems such as inaccurate semantic labeling, insufficient multilingual support, and inefficient dynamic updates when dealing with complex semantics, multilingual environments and dynamic knowledge updates.

Method used

By acquiring multi-source heterogeneous data, standardized processing is performed and mapped to a unified semantic representation space, the embedded representations of different languages ​​are aligned using contrast learning technology, and the knowledge graph is updated in real time based on the incremental update algorithm.

Benefits of technology

It realizes accurate analysis of complex semantics, efficient alignment of multilingual knowledge and real-time update of dynamic knowledge graphs, significantly improving the system's response speed and knowledge update efficiency, and reducing the cost of manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962655A_ABST
    Figure CN119962655A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-language knowledge graph construction method, system and device and a medium. The method comprises the steps of obtaining text data, voice data and image data; performing standardization processing on the text data, the voice data and the image data to obtain multi-source heterogeneous data after format unification; mapping the multi-source heterogeneous data with the unified format into a unified semantic representation space, and performing semantic annotation to generate a corresponding semantic annotation result; the embedding representations of different languages are aligned based on a comparative learning technology, the embedding of multiple languages is mapped to a unified semantic space, and a semantic annotation result is updated; and constructing a knowledge graph based on the updated semantic annotation result, and updating the knowledge graph in real time based on an incremental updating algorithm. According to the technical scheme, precise analysis of complex semantics and efficient annotation of unseen knowledge can be achieved, dynamic real-time updating of the knowledge graph is achieved, and the response speed and knowledge updating efficiency of a system are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a multilingual knowledge graph construction method, system, device and medium. Background Art

[0002] Knowledge graph is a technology that represents entities and their relationships in the form of a graphical structure. It is widely used in search engines, question-answering systems, recommendation engines, intelligent customer service, etc. Existing knowledge graph construction technologies mainly rely on artificial rules or traditional machine learning methods. These methods have the following shortcomings when dealing with complex semantics, multilingual environments, and dynamic knowledge updates:

[0003] 1. Limitations of semantic annotation: Existing technologies mainly rely on predefined rules, dictionaries or shallow machine learning models to complete semantic annotation. Due to the static nature of the rules, the system has difficulty processing complex contexts, implicit semantics or dynamically evolving knowledge, and the annotation results are often inaccurate or incomplete.

[0004] 2. Insufficient support for multilingual knowledge graphs: Many knowledge graph construction technologies focus on a single language (such as English). In a multilingual environment, there is a lack of effective semantic mapping and alignment methods for multilingual data. Especially in scenarios with large differences in semantic complexity and language habits, existing technologies are difficult to meet the needs.

[0005] 3. Inefficiency of dynamic updates: Knowledge graphs need to be updated frequently in practical applications to reflect the generation of new knowledge. However, most existing methods rely on manual intervention or regular batch updates, which makes it difficult to achieve real-time response. This inefficiency makes knowledge graphs lack practicality in rapidly changing scenarios. The core of these problems lies in the fact that existing technologies cannot fully utilize the capabilities of generative artificial intelligence, especially in terms of dynamic semantic understanding, multilingual support, and knowledge extraction and updating. There are significant bottlenecks. Summary of the invention

[0006] The purpose of the present invention is to provide a multilingual knowledge graph construction method, system, device and medium to solve the problems existing in the above-mentioned prior art.

[0007] To achieve the above object, the present invention provides a method for constructing a multilingual knowledge graph, comprising:

[0008] Acquire multi-source heterogeneous data, wherein the multi-source heterogeneous data includes text data, voice data, and image data;

[0009] Standardizing the text data, the voice data, and the image data to obtain multi-source heterogeneous data in a unified format;

[0010] Map the multi-source heterogeneous data with unified format into a unified semantic representation space and perform semantic annotation to generate corresponding semantic annotation results;

[0011] Align the embedding representations of different languages ​​based on contrastive learning technology, map the embeddings of multiple languages ​​to a unified semantic space, and update the semantic annotation results;

[0012] A knowledge graph is constructed based on the updated semantic annotation results, and the knowledge graph is updated in real time based on an incremental update algorithm.

[0013] Optionally, the standardization processing of the text data, the voice data and the image data specifically includes:

[0014] The text data is processed by word segmentation, stop word removal and syntactic analysis to obtain preprocessed text data; the voice data is processed by text conversion to obtain preprocessed voice data, and the image data is processed by text recognition to obtain preprocessed image data; by preprocessing the data of each modality, multi-source heterogeneous data with a unified format is obtained.

[0015] Optionally, mapping the multi-source heterogeneous data in a unified format to a unified semantic representation space and performing semantic annotation specifically includes:

[0016] A semantic annotation model is built based on the Transformer architecture. The semantic annotation model is used to map different modal data into a unified semantic representation space. The cosine similarity formula is used to calculate the semantic relevance of each modal data. The semantic annotation results are generated based on the Beam Search technology / Sampling technology combined with the semantic relevance. The semantic annotation model is dynamically optimized based on the incremental learning mechanism.

[0017] Optionally, aligning the embedding representations of different languages ​​based on contrastive learning technology and mapping the embeddings of multiple languages ​​to a unified semantic space specifically includes:

[0018] Based on the SimCSE model, the embedding spaces of different languages ​​are aligned, and the aligned multilingual embedding representation is input into the mBERT model to map the multilingual embeddings into a unified semantic space.

[0019] Optionally, the constructing of a knowledge graph based on the updated semantic annotation results specifically includes:

[0020] Based on named entity recognition technology, key entities are extracted, the key entities are input into a relationship prediction model for prediction and classification, and the relationship between entities is output; wherein the relationship prediction model is built based on the BERT+CRF model.

[0021] A multilingual knowledge graph construction system, comprising:

[0022] A data acquisition module is used to acquire multi-source heterogeneous data, wherein the multi-source heterogeneous data includes text data, voice data and image data; and perform standardization processing on the text data, voice data and image data to obtain the multi-source heterogeneous data in a unified format;

[0023] The semantic annotation module is used to map the multi-source heterogeneous data with unified format into a unified semantic representation space and perform semantic annotation to generate corresponding semantic annotation results;

[0024] The multilingual semantic mapping module is used to align the embedding representations of different languages ​​based on contrastive learning technology, map the embeddings of multiple languages ​​to a unified semantic space, and update the semantic annotation results;

[0025] The dynamic knowledge graph construction module is used to construct a knowledge graph according to the updated semantic annotation results, and to update the knowledge graph in real time based on an incremental update algorithm.

[0026] An electronic device includes a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the multilingual knowledge graph construction method.

[0027] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the multilingual knowledge graph construction method.

[0028] The technical effects of the present invention are:

[0029] The present invention can achieve accurate analysis of complex semantics and efficient annotation of unseen knowledge. By constructing a multilingual knowledge graph system, the present invention can effectively support semantic alignment and mapping of multilingual data and solve the problem of semantic inconsistency between different languages. The present invention realizes dynamic real-time updating of knowledge graphs, significantly improves the response speed of the system and the efficiency of knowledge updating, and reduces the cost of manual intervention. The present invention provides a modular system design, which is convenient for large-scale deployment and adapts to cross-domain knowledge management and application needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0031] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0032] Figure 1 is a construction flow chart in an embodiment of the present invention;

[0033] Figure 2 It is a schematic diagram of the structure of each module in the embodiment of the present invention. DETAILED DESCRIPTION

[0034] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as limiting the present invention, but should be understood as a more detailed description of certain aspects, features, and embodiments of the present invention.

[0035] It should be understood that the terms described in the present invention are only for describing special embodiments and are not intended to limit the present invention. In addition, for the numerical range in the present invention, it should be understood that each intermediate value between the upper and lower limits of the scope is also specifically disclosed. Each smaller range between the intermediate value in any stated value or stated range and any other stated value or intermediate value in the described range is also included in the present invention. The upper and lower limits of these smaller ranges can be independently included or excluded in the scope.

[0036] It will be apparent to those skilled in the art that various modifications and variations may be made to the specific embodiments of the present invention description without departing from the scope or spirit of the present invention. Other embodiments derived from the present invention description will be apparent to those skilled in the art. The present application description and examples are exemplary only.

[0037] The words “include,” “including,” “have,” “contain,” etc. used in this article are open-ended terms, meaning including but not limited to.

[0038] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0039] Embodiment 1

[0040] like Figure 1-Figure 2As shown, this embodiment provides a method for constructing a multilingual knowledge graph, including: obtaining multi-source heterogeneous data, the multi-source heterogeneous data including text data, voice data and image data; standardizing the text data, voice data and image data to obtain multi-source heterogeneous data with a unified format; mapping the multi-source heterogeneous data with a unified format to a unified semantic representation space and performing semantic annotation to generate corresponding semantic annotation results; aligning embedding representations of different languages ​​based on contrastive learning technology, mapping the embeddings of multiple languages ​​to a unified semantic space, and updating the semantic annotation results; constructing a knowledge graph based on the updated semantic annotation results, and updating the knowledge graph in real time based on an incremental update algorithm.

[0041] This embodiment designs a dynamic semantic annotation and multilingual knowledge graph construction system combined with generative AI, and its technical solution includes the following modules and working principles:

[0042] 1. System module design, including:

[0043] (1) Data input module:

[0044] Function: Receive multi-source heterogeneous data, including text, voice, image and other formats; standardize the input data and convert it into the structured form required for semantic analysis.

[0045] Features: Supports multiple data interfaces, compatible with real-time data streams and static data files.

[0046] (2) Generative AI semantic annotation module:

[0047] Function: Based on the context-aware capabilities of generative AI (such as the GPT model), it performs semantic analysis and annotation on input data to generate structured knowledge fragments.

[0048] Implementation: Use pre-trained large-scale language models combined with domain-specific fine-tuning techniques to achieve dynamic understanding and precise annotation of complex semantics.

[0049] Features: Supports multi-layer semantic parsing, including entity recognition, relationship extraction, event detection, and implicit semantic inference.

[0050] (3) Multilingual semantic mapping module:

[0051] Function: Realize semantic alignment and knowledge fusion between different languages, and build a language-independent unified knowledge representation.

[0052] Implementation: Through the multilingual capabilities of generative AI, combined with alignment algorithms and translation technology, multilingual data can be automatically processed for semantic consistency.

[0053] Features: Supports mainstream languages ​​and low-resource languages, and solves the problem of differences in semantic habits.

[0054] (4) Dynamic knowledge graph construction module:

[0055] Function: Based on the semantic annotation results, construct the node and edge relationships of the knowledge graph; update the existing graph in real time and dynamically.

[0056] Implementation: Use incremental update algorithms and real-time knowledge extraction technology to maintain the freshness and integrity of knowledge graph content.

[0057] Features: Modular architecture, supporting distributed storage and efficient retrieval.

[0058] (5) Feedback optimization module:

[0059] Function: Use reinforcement learning (RLHF) and user feedback to continuously optimize the generative AI model to improve the accuracy of semantic annotation and knowledge construction.

[0060] Implementation: Design a dynamic feedback loop and take user input, annotation results, and the actual application effect of the knowledge graph as the optimization goal.

[0061] Features: Adaptive optimization mechanism to improve long-term system performance.

[0062] 2. The system workflow of this embodiment:

[0063] (1) The data input module receives multi-source data provided by the user and completes standardization processing.

[0064] (2) Standardized data enters the generative AI semantic annotation module to generate preliminary structured knowledge fragments.

[0065] (3) The multilingual semantic mapping module performs language alignment and unified knowledge representation processing on the annotation results.

[0066] (4) The dynamic knowledge graph construction module integrates knowledge fragments, updates or constructs the knowledge graph, and provides real-time query and display services.

[0067] (5) The feedback optimization module optimizes the generative AI model based on user feedback and the application effect of the knowledge graph.

[0068] 3. Technical implementation details

[0069] (1) Data processing algorithm: A generative model based on the Transformer architecture that supports batch and real-time input.

[0070] (2) Knowledge fusion method: semantic similarity calculation and logical reasoning technology are used to solve the problems of knowledge duplication and conflict.

[0071] (3) Update mechanism: Based on timestamps and data change logs, an efficient incremental update algorithm is implemented.

[0072] Compared with the prior art, this embodiment has the following significant advantages:

[0073] 1. Dynamicity: Utilize the contextual understanding ability of generative AI to achieve real-time updating of knowledge graphs and reduce human involvement.

[0074] 2. Multilingual adaptability: The multilingual semantic mapping module significantly improves the accuracy and efficiency of knowledge alignment between different languages.

[0075] 3. High degree of automation: The entire process from data input to knowledge graph construction is automated, reducing development and maintenance costs.

[0076] 4. System performance optimization: Continuously optimize the generative AI model through user feedback and reinforcement learning to improve the long-term reliability and accuracy of the system.

[0077] In the overall system architecture of this embodiment, the core algorithm processing process of each module is as follows, which shows in detail the main calculation process and key technologies from data input to knowledge graph construction:

[0078] 1. Data input module, including:

[0079] Preprocessing algorithm: For multi-source data (text, voice, image), specific preprocessing techniques are used:

[0080] For text: word segmentation, stop word removal, and syntactic analysis;

[0081] For speech: speech-to-text (ASR, Automatic Speech Recognition);

[0082] For images: OCR (Optical Character Recognition) technology extracts text information.

[0083] Calculation process:

[0084] Text segmentation: Use Double-Array Trie to speed up word segmentation.

[0085] OCR process: Convolutional neural network (CNN) combined with recurrent neural network (RNN) recognizes text sequences in image areas.

[0086] 2. Generative AI semantic annotation module, including:

[0087] Dynamic semantic annotation based on generative pre-trained models (such as GPT-4, LLaMA):

[0088] For input processing:

[0089] Multimodal embedding: Through the Transformer architecture, the input text, speech, and image are uniformly mapped to the semantic space.

[0090] Semantic similarity is calculated using the cosine similarity formula:

[0091] Sim(A,B)=∑Ai·Bi∑Ai2·∑Bi2\text{Sim}(A,B)=\frac{\sum A_i\cdot B_i}{\sqrt{\sum A_i^2}\cdot\sqrt{\sum B_i^2}}

[0092] The semantic relevance is determined based on the above cosine similarity formula.

[0093] Generate labels: Use beam search or sampling technology to generate the optimal label sequence, and dynamically adjust the labeling results: Dynamically optimize the semantic labeling model through incremental learning.

[0094] 3. Multilingual semantic mapping module, including:

[0095] Cross-language semantic mapping based on contrastive learning and language alignment techniques:

[0096] Contrastive Learning: Use the SimCSE (Simple Contrastive Learning of Sentence Embeddings) model to align the embedding spaces of different languages.

[0097] The loss function is contrast loss:

[0098]

[0099] Among them, τ\tau is the temperature parameter, xix_i and yiy_i are semantic embeddings.

[0100] Cross-lingual alignment: Use alignment models such as mBERT to map embeddings of multiple languages ​​into a unified semantic space.

[0101] 4. Dynamic knowledge graph construction module, including:

[0102] Entity extraction: Use named entity recognition (NER) technology to extract key entities. The algorithm uses the BER T+CRF model, which has higher accuracy.

[0103] Relationship extraction: Use a relationship prediction algorithm based on a graph neural network (GNN) to infer the relationship between entities.

[0104] The loss function is categorical cross entropy:

[0105]

[0106] Among them, yiy_i is the true relationship label, and y^i\hat{y}_i is the predicted value.

[0107] Dynamic update: Based on the incremental graph update algorithm, newly discovered entities and relationships are merged in real time.

[0108] 5. Feedback optimization module, using the policy gradient method of reinforcement learning to optimize the system, including:

[0109] Reward function design: positive scores of user feedback and coverage and accuracy of entity-relation pairs in the knowledge graph.

[0110] Strategy optimization formula:

[0111] Among them, πθ\pi_\theta is the policy network, and Q(s,a)Q(s,a) is the state action value.

[0112] Interactive relationship operation process:

[0113] 1. After the data is preprocessed by the input module, it enters the generative AI module for semantic analysis and annotation.

[0114] 2. The generated labels flow to the semantic mapping module to complete multi-language alignment.

[0115] 3. Entity and relationship information is fed into the knowledge graph construction module to update the knowledge structure.

[0116] 4. User feedback flows back to the optimization module to iteratively improve system performance.

[0117] Each module is tightly integrated through data flow and feedback loops. This architecture achieves highly intelligent and dynamic updating of semantic annotation and multilingual knowledge graph construction through modular division of labor and overall collaborative optimization.

[0118] This embodiment provides a dynamic semantic annotation method, which uses the context-awareness capability of generative AI to achieve accurate parsing of complex semantics and efficient annotation of unseen knowledge.

[0119] This embodiment builds a multilingual knowledge graph system that can effectively support the semantic alignment and mapping of multilingual data and solve the problem of semantic inconsistency between different languages.

[0120] This embodiment can realize dynamic real-time updating of the knowledge graph, significantly improve the system's response speed and knowledge updating efficiency, and reduce the cost of manual intervention.

[0121] This embodiment provides a modular system design to facilitate large-scale deployment and meet cross-domain knowledge management and application requirements.

[0122] An electronic device includes a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the multilingual knowledge graph construction method.

[0123] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the multilingual knowledge graph construction method.

[0124] The above is only a preferred specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for constructing a multilingual knowledge graph, characterized in that: include: Acquire multi-source heterogeneous data, wherein the multi-source heterogeneous data includes text data, voice data, and image data; Standardizing the text data, the voice data, and the image data to obtain multi-source heterogeneous data in a unified format; Map the multi-source heterogeneous data with unified format into a unified semantic representation space and perform semantic annotation to generate corresponding semantic annotation results; Align the embedding representations of different languages ​​based on contrastive learning technology, map the embeddings of multiple languages ​​to a unified semantic space, and update the semantic annotation results; A knowledge graph is constructed based on the updated semantic annotation results, and the knowledge graph is updated in real time based on an incremental update algorithm.

2. A multilingual knowledge graph construction method according to claim 1, characterized in that: The standardization processing of the text data, the voice data and the image data specifically includes: The text data is processed by word segmentation, stop word removal and syntactic analysis to obtain preprocessed text data; the voice data is processed by text conversion to obtain preprocessed voice data, and the image data is processed by text recognition to obtain preprocessed image data; by preprocessing the data of each modality, multi-source heterogeneous data with a unified format is obtained.

3. A multilingual knowledge graph construction method according to claim 1, characterized in that: The mapping of the multi-source heterogeneous data in a unified format into a unified semantic representation space and performing semantic annotation specifically includes: A semantic annotation model is built based on the Transformer architecture. The semantic annotation model is used to map different modal data into a unified semantic representation space. The cosine similarity formula is used to calculate the semantic relevance of each modal data. The semantic annotation results are generated based on the BeamSearch technology / Sampling technology combined with the semantic relevance. The semantic annotation model is dynamically optimized based on the incremental learning mechanism.

4. A multilingual knowledge graph construction method according to claim 1, characterized in that: The method of aligning the embedding representations of different languages ​​based on contrastive learning technology and mapping the embeddings of multiple languages ​​into a unified semantic space specifically includes: Based on the SimCSE model, the embedding spaces of different languages ​​are aligned, and the aligned multilingual embedding representation is input into the mBERT model to map the multilingual embeddings into a unified semantic space.

5. A multilingual knowledge graph construction method according to claim 1, characterized in that: The step of constructing a knowledge graph based on the updated semantic annotation results specifically includes: Based on named entity recognition technology, key entities are extracted, the key entities are input into a relationship prediction model for prediction and classification, and the relationship between entities is output; wherein the relationship prediction model is built based on the BERT+CRF model.

6. A multilingual knowledge graph construction system, characterized in that: include: A data acquisition module, used to acquire multi-source heterogeneous data, wherein the multi-source heterogeneous data includes text data, voice data and image data; Standardizing the text data, the voice data, and the image data to obtain multi-source heterogeneous data in a unified format; The semantic annotation module is used to map the multi-source heterogeneous data with unified format into a unified semantic representation space and perform semantic annotation to generate corresponding semantic annotation results; The multilingual semantic mapping module is used to align the embedding representations of different languages ​​based on contrastive learning technology, map the embeddings of multiple languages ​​to a unified semantic space, and update the semantic annotation results; The dynamic knowledge graph construction module is used to construct a knowledge graph according to the updated semantic annotation results, and to update the knowledge graph in real time based on an incremental update algorithm.

7. An electronic device, characterized in that: It includes a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute a multilingual knowledge graph construction method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements a multilingual knowledge graph construction method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Chinese cross-language event detection method based on graph neural network

    CN118069839A

  • Knowledge graph driven semantic governance scheme dynamic generation method and system

    CN118114758A

  • Cerebral hemorrhage personalized treatment scheme optimization method and system based on big data analysis

    CN119153117A

  • Knowledge map construction method and apparatus, and storage medium

    WO2018153266A1

  • Knowledge extraction method, apparatus, electronic device, and storage medium

    WO2021212682A1

Cited By

  • Vibrating compaction forming energy transfer efficiency evaluation method

    CN120597061A

  • A method for evaluating energy transmission efficiency of vibration compaction forming

    CN120597061B

  • Key attribute extraction method and system based on multi-modal normalization

    CN120705539A

  • Information extraction and dynamic updating method, system, equipment and medium

    CN120804455A

  • File information association analysis method based on heterogeneous knowledge graph

    CN121743747A