Aerospace application knowledge graph construction method, device, system, equipment and medium

By constructing a knowledge graph for aerospace applications and using a large language model to process multimodal data, the problem of underutilization of multimodal data in the aerospace application field has been solved, and efficient and accurate knowledge question answering and intelligent services have been achieved.

CN120930740APending Publication Date: 2025-11-11AEROSPACE INFORMATION RES INST CAS
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510950602.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

In existing aerospace applications, multimodal data has not been fully explored and integrated, affecting the intuitiveness and accuracy of knowledge services. Furthermore, traditional methods are insufficient to meet the demands for intelligent and efficient knowledge-based question answering.

Method used

By acquiring multimodal data from aerospace knowledge files, using a large language model for data processing and extraction, lightweight markup language text and embedding vectors are constructed. Combined with entity relation triples, an aerospace application knowledge graph is built and stored in a graph, vector, and relational database, achieving deep fusion of multimodal data and intelligent question answering.

Benefits of technology

It enables efficient and accurate presentation and intelligent question answering of knowledge in the field of aerospace applications, thereby improving the intelligence and efficiency of knowledge services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930740A_ABST
    Figure CN120930740A_ABST
Patent Text Reader

Abstract

The invention provides an aerospace application knowledge graph construction method, device, system and equipment and a medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining multi-modal data; based on the data types, data processing is carried out on the modal data, and lightweight markup language texts and embedded vectors are obtained; performing knowledge extraction on the lightweight markup language text on the basis of a preset cue word and a large language model to obtain metadata information, and constructing an entity relationship triple on the basis of a relationship between entities and the metadata information; and based on the aerospace knowledge file number information, determining an association relationship among the aerospace knowledge file, the lightweight markup language text, the embedded vector and the entity relationship triple, and constructing the aerospace application knowledge graph according to the association relationship and the entity relationship triple. According to the space application knowledge graph constructed by the method, the space application field entities and the association relationship can be accurately presented, and efficient and accurate knowledge support is provided for space application knowledge questions and answers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, system, device, and medium for constructing a knowledge graph for aerospace applications. Background Technology

[0002] Driven by the iterative development of information technology, the business architecture of the aerospace application field continues to expand. This is accompanied by the continuous growth of knowledge achievements and diverse heterogeneous data in the field. These data exhibit cross-dimensional and multi-modal fusion characteristics, forming a knowledge system for the aerospace application field with significant complexity and heterogeneity.

[0003] Existing data acquisition and retrieval models based on manual or database searches suffer from low efficiency and poor accuracy. While the use of knowledge graphs to build question-answering systems has been proposed, it faces two major limitations in the aerospace application field. Firstly, multimodal data (a combination of geographic vector data and text / structured data tables) has not been fully explored and integrated. When purely text-based responses lack intuitiveness, the inability to effectively incorporate image references to enhance the concrete presentation of questions and answers hinders the effectiveness of knowledge services. Secondly, the sheer volume of knowledge in the aerospace domain makes it difficult to meet the demands for intelligent and efficient aerospace knowledge services and question-answering.

[0004] Therefore, there is an urgent need for a method, device, system, equipment, and medium for constructing aerospace application knowledge graphs to solve the above problems. Summary of the Invention

[0005] To address the problems existing in the prior art, this invention provides a method, apparatus, system, equipment, and medium for constructing aerospace application knowledge graphs.

[0006] This invention provides a method for constructing a knowledge graph for aerospace applications, including: Acquire multimodal data of target aerospace knowledge from various aerospace knowledge files; Based on the data type of multimodal data, data processing is performed on each modality of the target aerospace knowledge multimodal data to obtain the lightweight markup language text and lightweight markup language text embedding vector corresponding to each modality of data. Based on preset aerospace knowledge extraction prompts and a large language model, aerospace knowledge extraction and knowledge completion processing are performed on the lightweight markup language text to obtain metadata information corresponding to each entity in the aerospace knowledge file. Based on the relationship between each entity in the aerospace knowledge file and the metadata information, a knowledge entity relationship triplet is constructed. Based on the numbering information of the aerospace knowledge file, the association relationship between the aerospace knowledge file, the lightweight markup language text, the lightweight markup language text embedding vector, and the knowledge entity relationship triple is determined, and an aerospace application knowledge graph is constructed according to the association relationship and the knowledge entity relationship triple.

[0007] According to a method for constructing a space application knowledge graph provided by the present invention, after determining the association relationship between the space knowledge file, the lightweight markup language text, the lightweight markup language text embedding vector, and the knowledge entity relation triple based on the numbering information of the space knowledge file, and constructing the space application knowledge graph based on the association relationship and the knowledge entity relation triple, the method further includes: A graph database is constructed based on the aforementioned aerospace application knowledge graph; A vector database is constructed based on the lightweight markup language text embedding vectors corresponding to the lightweight markup language texts. A relational database is constructed based on the aerospace knowledge file and the document access path information corresponding to the aerospace knowledge file; The graph database, the vector database, and the relational database are used to perform cross-database association retrieval through the association relationships during the knowledge question-and-answer process for aerospace applications.

[0008] According to the method for constructing aerospace application knowledge graph provided by the present invention, the target aerospace knowledge multimodal data includes at least text data, tabular data and image data; The data type based on multimodal data involves processing each modality of the target aerospace knowledge multimodal data to obtain lightweight markup language text and lightweight markup language text embedding vectors corresponding to each modality, including: Lightweight language conversion processing is performed on the text data to obtain the lightweight markup language text and the lightweight markup language text embedding vector corresponding to the text data; The table data is preprocessed to obtain the lightweight markup language text and the lightweight markup language text embedding vector corresponding to the table data. The data preprocessing includes at least data deduplication, missing value imputation, outlier removal and format normalization. Obtain the string description content corresponding to the image data, and construct the lightweight markup language text and the lightweight markup language text embedding vector corresponding to the image data based on the string description content.

[0009] According to a method for constructing a knowledge graph for aerospace applications provided by the present invention, the step of obtaining the string description content corresponding to the image data, and constructing the lightweight markup language text and the lightweight markup language text embedding vector corresponding to the image data based on the string description content, includes: Based on a visual language model, the image data is parsed to obtain the semantic information of the image data. The context information of the image data in the aerospace knowledge file is obtained, and the context information is input into the large language model to obtain the image and text supplementary information corresponding to the image data; A hash code value corresponding to the image data is generated, and the lightweight markup language text and the lightweight markup language text embedding vector corresponding to the image data are constructed based on the image semantic information, the image and text supplementary information and the hash code value.

[0010] According to a method for constructing aerospace application knowledge graphs provided by the present invention, the method further includes: Based on the large language model, geographic information is identified in the lightweight markup language text. If there is geospatial semantics in the lightweight markup language text, the corresponding geographic unit information in the lightweight markup language text is obtained. Based on a preset geocoding algorithm, the geographic unit information is converted into corresponding geocoding information; The geocoding information is added to the lightweight markup language text to obtain lightweight markup language text with geocoding information.

[0011] According to the aerospace application knowledge graph construction method provided by the present invention, the preset aerospace knowledge extraction prompt words include aerospace knowledge completion prompt words and aerospace knowledge extraction prompt words; The process involves extracting and completing space-related knowledge from the lightweight markup language text based on preset space-related knowledge extraction prompts and a large language model. This yields metadata information corresponding to each entity in the space-related knowledge file. Furthermore, based on the relationships between entities in the space-related knowledge file and the metadata information, a knowledge entity relationship triplet is constructed, including: Based on the aforementioned aerospace knowledge completion prompts and the aforementioned large language model, knowledge completion information corresponding to the lightweight markup language text is generated within a three-dimensional ontology framework. The three-dimensional ontology framework includes the aerospace information's business system, technical system, and data system. Based on the knowledge completion information, the lightweight markup language text is completed to obtain the completed lightweight markup language text; Based on the aforementioned aerospace knowledge extraction prompts and the large language model, aerospace knowledge extraction processing is performed on the completed lightweight markup language text to obtain metadata information corresponding to each entity in the completed lightweight markup language text. Based on the relationships between each entity in the completed lightweight markup language text and the metadata information, the knowledge entity relationship triplet is constructed.

[0012] The present invention also provides a device for constructing a knowledge graph for aerospace applications, comprising: The multimodal data acquisition module is used to acquire multimodal data of target aerospace knowledge from various aerospace knowledge files; The multimodal data processing module is used to process each modal data in the target aerospace knowledge multimodal data according to the data type of the multimodal data, and obtain the lightweight markup language text and lightweight markup language text embedding vector corresponding to each modal data. The aerospace knowledge extraction module is used to perform aerospace knowledge extraction and knowledge completion processing on the lightweight markup language text based on preset aerospace knowledge extraction prompts and a large language model, to obtain metadata information corresponding to each entity in the aerospace knowledge file, and to construct knowledge entity relationship triples based on the relationship between each entity in the aerospace knowledge file and the metadata information. The aerospace application knowledge graph construction module is used to determine the association relationship between the aerospace knowledge file, the lightweight markup language text, the lightweight markup language text embedding vector and the knowledge entity relationship triple based on the numbering information of the aerospace knowledge file, and to construct the aerospace application knowledge graph based on the association relationship and the knowledge entity relationship triple.

[0013] This invention also provides a space application knowledge question-and-answer system, comprising: The question receiving module is used to acquire target aerospace knowledge questions; The problem parsing module is used to obtain the aerospace application knowledge graph information corresponding to the target aerospace knowledge problem from the graph database, wherein the aerospace application knowledge graph information is generated based on the above-mentioned aerospace application knowledge graph construction method; The question-and-answer module is used to input the aerospace application knowledge graph information into the large language model to obtain the question-and-answer results corresponding to the target aerospace knowledge question.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the aerospace application knowledge graph construction method described above.

[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the aerospace application knowledge graph construction method as described above.

[0016] The aerospace application knowledge graph construction method, apparatus, system, device, and medium provided by this invention obtains multimodal data of target aerospace knowledge from various aerospace knowledge files. Based on data type, different modalities of data are processed and transformed into lightweight markup language text. Then, using preset aerospace knowledge extraction prompts and a large language model, metadata information corresponding to each entity is extracted from the text. Knowledge entity relationship triples are then constructed by combining the relationships between entities. Finally, based on the aerospace knowledge file number information, the association between the file, lightweight markup language text, embedding vector, and knowledge entity relationship triples is clarified. This constructs an aerospace application knowledge graph that accurately presents entities and relationships in the aerospace application domain, providing efficient and accurate knowledge support for aerospace application knowledge question answering. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating the aerospace application knowledge graph construction method provided by this invention; Figure 2 A schematic diagram of the aerospace application knowledge graph provided by this invention; Figure 3 This is a schematic diagram of the multimodal data processing flow provided by the present invention; Figure 4 A schematic diagram illustrating the image semantic parsing process based on a visual language model provided by this invention; Figure 5 A schematic diagram illustrating the knowledge extraction and knowledge completion of the aerospace application knowledge graph provided by this invention. Figure 6 A schematic diagram illustrating the overall process of constructing the aerospace application knowledge graph provided by this invention; Figure 7 A schematic diagram of the aerospace application knowledge graph construction device provided by the present invention; Figure 8 This is a schematic diagram of the structure of the aerospace application knowledge question-and-answer system provided by the present invention; Figure 9This is a schematic diagram illustrating the database retrieval and question-answer generation process for different types of questions provided by the present invention; Figure 10 This invention provides a schematic diagram of a question-and-answer process for spatial regions within aerospace knowledge files. Figure 11 A schematic diagram illustrating the tracing process of original aerospace knowledge documents provided by this invention; Figure 12 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0020] In the current aerospace application knowledge system, data acquisition and retrieval modes based on manual or database searches face the dual limitations of efficiency bottlenecks and insufficient accuracy. Existing methods using knowledge graphs to unify and integrate data to build aerospace knowledge question-and-answer systems have two major limitations: First, in the aerospace application field, geographic knowledge graphs are typically constructed using multimodal data graphs combining geographic vector data, text data, and structured data tables. This "text-image hybrid" multimodal data has not been fully explored and integrated. In scenarios where purely textual responses lack intuitiveness, image reference mechanisms cannot be effectively introduced to enhance the concrete presentation of question-and-answer content, thus affecting the effectiveness of knowledge services. Second, the aerospace application field contains a vast amount of knowledge. Traditional methods relying on manual or machine learning instance extraction for knowledge graph construction lack flexibility, have low intelligence, and significantly limit knowledge modeling efficiency. They also lack standardized and universal knowledge systems and storage architectures, making it difficult to meet the intelligent and efficient needs of aerospace application knowledge services and question-and-answer work.

[0021] To address the problems existing in the prior art, this invention designs an aerospace application knowledge system and knowledge domain ontology based on the aerospace domain. It utilizes the semantic understanding and generation capabilities of a large language model to assist in knowledge extraction and completion, constructing a three-layer storage system for aerospace application knowledge that integrates multimodal data. Based on this, an intelligent question-and-answer system for aerospace knowledge is built. On one hand, based on an image semantic parsing framework using a visual language model, it achieves deep fusion of multimodal data through "image-text hybrid layout + multimodal data embedding." On the other hand, it introduces intelligent graph knowledge extraction and completion technology driven by a large language model, designing a three-layer knowledge data storage architecture consisting of a vector database, a graph database, and a relational database. This architecture decomposes and classifies knowledge information for storage, providing effective support for intelligent question-and-answering of aerospace application knowledge information.

[0022] Figure 1 This is a flowchart illustrating the aerospace application knowledge graph construction method provided by the present invention, as shown below. Figure 1 As shown, this invention provides a method for constructing a knowledge graph for aerospace applications, including: Step 101: Obtain the target aerospace knowledge multimodal data from each aerospace knowledge file.

[0023] In this invention, key information is extracted from multiple aerospace knowledge files to provide foundational data for subsequent knowledge processing and atlas construction. Specifically, in the aerospace field, knowledge is often stored in various forms in different files, such as documents containing text descriptions, image sets recording image data, and tabular files storing structured information. This invention comprehensively collects these aerospace knowledge files through file reading and data interface calls, and then filters out target aerospace knowledge multimodal data relevant to the current application scenario or research objective, covering multiple modalities such as text, images, videos, and structured data. For example, when studying satellite orbit data, it is necessary to obtain text files containing satellite orbit parameters and image monitoring data of satellite operational status.

[0024] Step 102: Based on the data type of the multimodal data, perform data processing on each modal data in the target aerospace knowledge multimodal data to obtain the lightweight markup language text and lightweight markup language text embedding vector corresponding to each modal data.

[0025] In this invention, aerospace knowledge data of different modalities are uniformly converted into lightweight markup language text (such as Markdown) to facilitate subsequent knowledge extraction and processing.

[0026] Since different modalities of data have different formats and structures, they need to be processed separately according to their data types. For text data, preprocessing operations such as text cleaning, word segmentation, part-of-speech tagging, and named entity recognition are performed to remove noise information, extract key semantic components, and convert them into a lightweight markup language format, using specific tags to label different semantic units.

[0027] For image data, image recognition technology is used to identify key objects and scenes in the image, converting this visual information into text descriptions, and then encoding them according to lightweight markup language rules. For example, information such as cloud distribution and terrain features in satellite images can be extracted and converted into text descriptions, which are then labeled with appropriate tags.

[0028] For structured data, it is converted into lightweight markup language text to preserve the structured information of the data, such as the row and column relationships in a table, which can be reflected through the tag hierarchy.

[0029] In this invention, for the above-mentioned different modalities of aerospace knowledge data, when obtaining the corresponding lightweight markup language text, the corresponding lightweight markup language text embedding vector is also obtained.

[0030] Step 103: Based on the preset aerospace knowledge extraction prompts and large language model, perform aerospace knowledge extraction and knowledge completion processing on the lightweight markup language text to obtain the metadata information corresponding to each entity in the aerospace knowledge file, and construct knowledge entity relationship triples based on the relationship between each entity in the aerospace knowledge file and the metadata information.

[0031] In this invention, key knowledge is extracted from lightweight markup language text and represented in the form of structured knowledge entity relation triples, facilitating the storage and reasoning of the knowledge graph. Specifically, specialized prompts for the aerospace domain (i.e., preset aerospace knowledge extraction prompts) are pre-designed to guide the large language model in understanding the semantics of the text. In this invention, the lightweight markup language text is input into the large language model, and the semantic understanding and knowledge reasoning capabilities of the large language model are utilized to extract the knowledge entities (such as algorithm models, thematic data, application cases, etc.) corresponding to the aerospace application knowledge text, as well as the metadata information (such as knowledge categories, knowledge providers and their organizations, etc.) corresponding to the entities.

[0032] Furthermore, the relationships between entities in the aerospace knowledge files are analyzed, such as the "belonging" relationship between algorithm models and knowledge categories, and the "attribution" relationship between thematic data and knowledge providers. Combining the extracted metadata information, entities and relationships are represented as triples, thus obtaining knowledge entity-relationship triples, such as (algorithm model A, belongs to, knowledge category B), (thematic data A, belongs to, knowledge provider C), etc.

[0033] Step 104: Based on the numbering information of the aerospace knowledge file, determine the association relationship between the aerospace knowledge file, the lightweight markup language text, the lightweight markup language text embedding vector, and the knowledge entity relationship triplet, and construct an aerospace application knowledge graph based on the association relationship and the knowledge entity relationship triplet.

[0034] In this invention, all the processed knowledge obtained in the above embodiments is integrated to construct a complete knowledge graph, providing comprehensive knowledge support for applications such as aerospace knowledge question answering. First, using the identification information of aerospace knowledge files (i.e., the document ID corresponding to each aerospace knowledge file), the association between aerospace knowledge files, lightweight markup language text, lightweight markup language text embedding vectors, and knowledge entity relation triples is established. The identification information serves as a unique identifier, linking data within the same file and clarifying the connections between data in different files. For example, the identification information can trace which specific aerospace knowledge file a knowledge entity relation triple originates from, and the corresponding lightweight markup language text content of that file.

[0035] Then, based on the established relationships, all knowledge entity relationship triples are integrated into a unified knowledge graph. The knowledge graph stores knowledge in a graph structure, with entities as nodes and relationships as edges. Through the connections between nodes and edges, the knowledge system of the aerospace domain is fully presented, providing an accurate and comprehensive foundation for knowledge query and reasoning for the aerospace knowledge question-and-answer system.

[0036] The aerospace application knowledge graph construction method provided by this invention obtains multimodal data of target aerospace knowledge from various aerospace knowledge files, processes different modal data according to data type, and transforms them into lightweight markup language text; then, using preset aerospace knowledge extraction prompts and a large language model, it extracts metadata information corresponding to each entity from the text, and constructs knowledge entity relationship triples by combining the relationships between entities; finally, based on the aerospace knowledge file number information, it clarifies the association relationship between the file, lightweight markup language text, embedding vector, and knowledge entity relationship triples, thereby constructing an aerospace application knowledge graph that can accurately present entities and relationships in the aerospace application field, providing efficient and accurate knowledge support for aerospace application knowledge question answering.

[0037] Based on the above embodiments, after determining the association relationship between the aerospace knowledge file, the lightweight markup language text, the lightweight markup language text embedding vector, and the knowledge entity relation triplet based on the numbering information of the aerospace knowledge file, and constructing the aerospace application knowledge graph according to the association relationship and the knowledge entity relation triplet, the method further includes: A graph database is constructed based on the aforementioned aerospace application knowledge graph; A vector database is constructed based on the lightweight markup language text embedding vectors corresponding to the lightweight markup language texts. A relational database is constructed based on the aerospace knowledge file and the document access path information corresponding to the aerospace knowledge file; The graph database, the vector database, and the relational database are used to perform cross-database association retrieval through the association relationships during the knowledge question-and-answer process for aerospace applications.

[0038] In this invention, the knowledge information (i.e., aerospace application knowledge graph, lightweight markup language text, and aerospace knowledge files) obtained through the above embodiments is mainly stored in three types of databases: graph database, vector database, and relational database. The graph database stores metadata information of knowledge entities in the aerospace application knowledge graph, providing basic data support for question-and-answer retrieval of knowledge metadata; the vector database stores the generated lightweight markup language text, providing data support for semantic retrieval; and the relational database stores the original file data of knowledge entities (i.e., aerospace knowledge files), providing underlying data assurance for data tracing and querying. The three types of knowledge information are linked by the identification information of the aerospace knowledge files, jointly constructing a hierarchical knowledge storage and retrieval system.

[0039] Specifically, in this invention, the graph database is mainly used to store knowledge metadata information (such as knowledge title, knowledge author, and knowledge publication time), and its main function is to provide data support for efficient retrieval of knowledge metadata. Based on the knowledge entity relation triples extracted from the large language model in the above embodiments, preliminary disambiguation fusion is performed to form standardized knowledge representation units, which are then stored in the graph database. During this process, the identification information of the aerospace knowledge files is also stored in the knowledge graph as an attribute of the knowledge instance. Figure 2 This is a schematic diagram of the aerospace application knowledge graph provided by the present invention. The constructed aerospace application knowledge graph can be referred to. Figure 2 As shown.

[0040] Long textual knowledge is prevalent in aerospace applications. Existing segmentation methods easily lead to semantic fragmentation of the text, resulting in a lack of correlation between knowledge units and information silos. To address this issue, this invention introduces a vector database storage mechanism, storing the lightweight markup language text processed in the above embodiments into a vector database for subsequent semantic similarity retrieval. The specific process is as follows: First, the lightweight markup language (LTL) text, along with its corresponding identifiers and segment IDs, is extracted. Then, a text vectorization tool is used to convert the LTL text into high-dimensional vectors. Next, the high-dimensional vector data is stored back into a vector database, constructing a vector database system that supports semantic similarity retrieval.

[0041] In this invention, the construction of a vector database can effectively establish a semantic association network across graph nodes, effectively solve the semantic discretization problem in long text processing, and provide strong support for the rapid query of semantically similar information in subsequent question answering.

[0042] In this invention, the relational database constructs an association storage mechanism between document access path information (such as Uniform Resource Locator URL) and number information to achieve structured management of the correspondence between the two, providing users with source tracing support for the original document content. Through the mapping relationship between unique number information and corresponding access path, it ensures that users can efficiently locate and trace the source of the original document content.

[0043] Based on the above embodiments, the target aerospace knowledge multimodal data includes at least text data, tabular data, and image data; The data type based on multimodal data involves processing each modality of the target aerospace knowledge multimodal data to obtain lightweight markup language text and lightweight markup language text embedding vectors corresponding to each modality, including: Lightweight language conversion processing is performed on the text data to obtain the lightweight markup language text and the lightweight markup language text embedding vector corresponding to the text data; The table data is preprocessed to obtain the lightweight markup language text and the lightweight markup language text embedding vector corresponding to the table data. The data preprocessing includes at least data deduplication, missing value imputation, outlier removal and format normalization. Obtain the string description content corresponding to the image data, and construct the lightweight markup language text and the lightweight markup language text embedding vector corresponding to the image data based on the string description content.

[0044] In this invention, the data types of target aerospace knowledge multimodal data mainly include text, images, tables, and nested data with "mixed text and images + multimodal data embedding". Figure 3 This is a schematic diagram of the multimodal data processing flow provided by the present invention, which can be referred to. Figure 3As shown, taking text documents as an example, file formats include docx, ofd, pdf, doc, and wps. In the field of aerospace applications, these text documents often vary in length, covering data descriptions, algorithm explanations, application cases, and knowledge achievements. Furthermore, these documents frequently employ a composite structure of "text and images mixed with multimodal data embedding," containing numerous complex reference tables. Not only are there a large number of tables and images, but the tables often span multiple pages and involve extensive cell merging. Although current large language models possess document parsing capabilities, they still face the dual challenges of ensuring data integrity and accurately reconstructing semantic relationships when processing nested table structures within documents.

[0045] This invention considers the composite modal data contained in aerospace knowledge files. By parsing the nested file content types, it extracts plain text data, tabular data, and image data for step-by-step processing. After targeted optimization based on the data characteristics of each modal data, it is uniformly converted into a lightweight markup language (Markdown) recording format. Each piece of knowledge is assigned a number based on its source, and then processed into chunks using the `chunk` function. Figure 3 In M 1. M 2 and M 3. This facilitates subsequent index creation and knowledge storage.

[0046] Specifically, in this invention, for textual data with low complexity, a lightweight language conversion framework can be directly used for conversion. However, tabular data is typically large in quantity and may contain semantic noise including meaningless strings and blank values; furthermore, multi-page tables, exceeding the segmentation range of a large language model, are often dynamically segmented into discrete, independent semantic units, resulting in impaired logical coherence of the table content. Therefore, this invention addresses structured nested tables by combining pandas tools with a lightweight markup language, with the specific steps as follows: Step S1, Data Deduplication: Remove duplicate rows to ensure that each data entry is unique; Step S2, fill in missing values ​​or meaningless strings: use front and back padding to fill in missing values ​​in text tables, and use linear interpolation to fill in missing data points; Step S3, Remove obviously unreasonable outliers: Set a threshold or use statistical methods to identify and remove values ​​that exceed the expected range; Step S4, Standardization Processing: Analyze each list field to ensure that all fields have a consistent format and units.

[0047] After preprocessing the tabular data, the preprocessed tabular data is converted into Markdown format. Markdown format is concise and lightweight, which can eliminate format differences between multiple documents, improve the recognition and retrieval capabilities of large language models, and can well parse cross-page and cross-table content to clearly display data structure and analysis results, while ensuring the accuracy and reliability of subsequent analysis.

[0048] Furthermore, this invention utilizes image recognition and understanding technology to extract and analyze features from image data, identifying key information such as objects, scenes, human behavior, and color distribution within the image. This visual information is then converted into strings describing the image in natural language. For example, for an image of a satellite flying in space, elements such as the satellite, the space background, and its flight attitude are identified, and corresponding string descriptions are generated. Based on these string descriptions, a lightweight markup language text is constructed corresponding to the image data.

[0049] Based on the above embodiments, the step of obtaining the string description content corresponding to the image data, and constructing the lightweight markup language text and the lightweight markup language text embedding vector corresponding to the image data based on the string description content, includes: Based on a visual language model, the image data is parsed to obtain the semantic information of the image data. The context information of the image data in the aerospace knowledge file is obtained, and the context information is input into the large language model to obtain the image and text supplementary information corresponding to the image data; A hash code value corresponding to the image data is generated, and the lightweight markup language text and the lightweight markup language text embedding vector corresponding to the image data are constructed based on the image semantic information, the image and text supplementary information and the hash code value.

[0050] In this invention, key example images are often embedded in professional textual materials in the field of aerospace applications. These image data are mainly used to intuitively demonstrate logical processes, example images, and other important information, playing an irreplaceable role in understanding and mastering relevant knowledge. However, existing technologies have weak analytical capabilities for the information contained in the image data itself, directly limiting its practical value in complex application scenarios, especially in applications that require visual information to assist in understanding, judgment, and decision-making. Figure 4 The diagram below illustrates the image semantic parsing process based on a visual language model provided by this invention. Figure 4 As shown, this invention utilizes the image understanding capabilities of a visual language model to generate string descriptions for image data that do not exceed a predefined length.

[0051] Specifically, in this invention, for image data that does not contain geographic information, a visual language model is used to identify the core thematic content contained in the image data, including the image description subject, the main behavior of the subject, and its significance; for thematic maps that contain geographic information, it is also necessary to extract the thematic map's metadata information (such as the mapping time, mapping unit, and thematic map name) as well as the region to which the thematic map belongs and the image information, thereby obtaining the image's semantic information.

[0052] Furthermore, the preceding and following paragraphs of the aerospace knowledge file containing the image data are obtained and appropriately supplemented in combination with the original image description. This process can utilize a large language model to generate context-coherent enhanced graphic descriptions, i.e., supplementary graphic information.

[0053] In this invention, considering the possibility that the same description results may be generated, the hash value of the image data is calculated and used as the name for storing the image data.

[0054] Next, the semantic information of the image and the supplementary text information are used as the image description content in the lightweight markup language text Markdown, and combined with the local location (if any) to form the Markdown representation of the image data.

[0055] For example, you can refer to Figure 4 As shown, the lightweight markup language text corresponding to the final generated image data is: ![Introduction to Remote Sensing Satellites]( / md_image / ca190867872ede5323d47c8 96fd8a55ff72e9d999bcf3ab1c86baccfd4d1fcda.jpg); The "Remote Sensing Satellite Introduction" section summarizes the relevant content of the image data (i.e., semantic information) and provides supplementary information generated based on the context (i.e., supplementary text and image information). "ca190867872ede5323d47c896fd8a55ff72e9d999bcf3ab1c86baccfd4d1fcda" is the hash code corresponding to the image data. Finally, the Markdown representation of the image is re-embedded into the original aerospace knowledge file at the location of the image data to ensure the accuracy and contextual relevance of the lightweight markup language text corresponding to the image data.

[0056] Based on the above embodiments, the method further includes: Based on the large language model, geographic information is identified in the lightweight markup language text. If there is geospatial semantics in the lightweight markup language text, the corresponding geographic unit information in the lightweight markup language text is obtained. Based on a preset geocoding algorithm, the geographic unit information is converted into corresponding geocoding information; The geocoding information is added to the lightweight markup language text to obtain lightweight markup language text with geocoding information.

[0057] Existing technologies lack effective methods for parsing and fusing geospatial semantic elements such as administrative division names and landmark references contained in aerospace knowledge files. This invention integrates a large language model and geocoding algorithms to construct a collaborative identification and parsing framework for geographic information. The specific process is as follows: First, a large language model is used to identify lightweight markup language data and extract geographic unit information (such as country, province, city, district, and county). Then, a preset geocoding algorithm is used to convert the extracted geographic units into corresponding geocoded information, which is then written into the corresponding lightweight markup language data to form complete information with geocoding.

[0058] In this invention, after data preprocessing, the organized lightweight markup language text is assigned corresponding numbering information to establish a traceable index system. Then, the lightweight markup language text is segmented into chunks according to specified tokens using the chunking function to generate unique chunk_ids, providing structured data support for subsequent semantic retrieval and knowledge graph construction.

[0059] Based on the above embodiments, the preset aerospace knowledge extraction prompts include aerospace knowledge completion prompts and aerospace knowledge extraction prompts; The process involves extracting and completing space-related knowledge from the lightweight markup language text based on preset space-related knowledge extraction prompts and a large language model. This yields metadata information corresponding to each entity in the space-related knowledge file. Furthermore, based on the relationships between entities in the space-related knowledge file and the metadata information, a knowledge entity relationship triplet is constructed, including: Based on the aforementioned aerospace knowledge completion prompts and the aforementioned large language model, knowledge completion information corresponding to the lightweight markup language text is generated within a three-dimensional ontology framework. The three-dimensional ontology framework includes the aerospace information's business system, technical system, and data system. Based on the knowledge completion information, the lightweight markup language text is completed to obtain the completed lightweight markup language text; Based on the aforementioned aerospace knowledge extraction prompts and the large language model, aerospace knowledge extraction processing is performed on the completed lightweight markup language text to obtain metadata information corresponding to each entity in the completed lightweight markup language text. Based on the relationships between each entity in the completed lightweight markup language text and the metadata information, the knowledge entity relationship triplet is constructed.

[0060] Figure 5 This is an overall schematic diagram of knowledge extraction and knowledge completion for the aerospace application knowledge graph provided by the present invention, which can be referred to. Figure 5 As shown, in this invention, a large language model is introduced. Based on the lightweight markup language text obtained by processing multimodal data in the above embodiments, an aerospace knowledge system and domain ontology are established. The automatic generation of the knowledge system is driven by preset aerospace knowledge extraction prompts, and a three-layer storage is achieved by integrating vector databases, graph databases, and relational databases.

[0061] Specifically, given the relatively complex nature of knowledge in the aerospace application field, encompassing multiple systems, this invention aims to further classify and store aerospace knowledge. Based on existing aerospace knowledge, a three-dimensional ontology framework is constructed, centered on a business system (such as aerospace knowledge application scenarios), a technical system (such as algorithms involved in aerospace knowledge), and a data system (such as related datasets of aerospace knowledge). The business system is divided into four levels: "domain-direction-sub-direction-scenario," involving 17 business directions, 58 sub-directions, and 231 application scenarios across four key areas: natural resources, agriculture and rural areas, disaster reduction and emergency response, and ecological environment. The technical system involves five key technical areas and 34 technical categories: thematic information extraction, remote sensing data preprocessing, remote sensing inversion and assimilation, remote sensing application services, and remote sensing product authenticity verification. The data system is divided into "major data category-minor data category-data product," involving four major categories: remote sensing data, common remote sensing products, thematic application products, and other data, 23 minor data categories, and 106 data products.

[0062] In this invention, the business system adopts a multi-level classification logic, constructing a systematic business application architecture through a progressive division from domain, direction, sub-direction to scenario. Starting with a specific professional field, the business system is refined layer by layer to specific application scenarios, forming a complete business framework from macro-level to micro-level practice. This aims to provide clear path guidance for knowledge transformation and practical application, achieving scientific classification and standardized management of complex business scenarios.

[0063] The technology system focuses on the organic integration of key technology areas and specific technology categories. Based on a reasonable division of technology areas, it subdivides and sorts out systematic technology categories, forming a clear and semantically unified technology framework, which provides support for the construction of a space-air application knowledge graph.

[0064] The data system follows a hierarchical logic of "major data categories - minor data categories - data products". By integrating basic data, common products, special application data and other supplementary data, it forms a well-defined and complementary data ecosystem, which is the core resource hub driving the efficient operation of the aerospace knowledge management system.

[0065] Meanwhile, based on the subordinate instances under the established knowledge classification system, this invention designs the ontology, inter-class relationships and related attributes according to the metadata information covered by common knowledge. For specific design of the aerospace application knowledge ontology, please refer to Table 1. It should be noted that the spatial range is represented by tile coordinates and the specific coordinates are described in string encoding form.

[0066] Table 1. Design of Knowledge Ontology for Aerospace Applications

[0067] Existing knowledge extraction methods for aerospace applications rely heavily on manual identification. This invention introduces a large language model to assist in extraction, and designs preset aerospace knowledge extraction prompts corresponding to knowledge extraction and completion, namely aerospace knowledge completion prompts and aerospace knowledge extraction prompts, to identify, extract and complete aerospace knowledge documents.

[0068] Specifically, this invention utilizes a large language model to assist knowledge providers in intelligently recommending the required knowledge metadata when sharing knowledge. This involves identifying the knowledge content provided by the provider, reasoning to identify and recommend information such as the applicable spatial region and knowledge category (business system, technology system, data system) to which the knowledge belongs, so as to complete the information based on this knowledge metadata and construct knowledge triples. It should be noted that this invention does not limit the specific form of the aerospace knowledge completion prompts. In one embodiment, the aerospace knowledge completion prompts include a role (e.g., defining the large language model as an expert in aerospace knowledge system identification), a goal (e.g., requiring the large language model to automatically identify and input the business system, technology system, and data system covered by aerospace knowledge), constraints (e.g., requiring the large language model to strictly follow the specifications and requirements of the relevant knowledge management system during the identification process), an output format (e.g., requiring the large language model to output the knowledge system classification and spatial region), a workflow (used to guide the large language model to perform data processing and identification according to a specific process), and examples (e.g., input objects, output results, and identified spatial regions).

[0069] Furthermore, after completing the lightweight markup language text, the large language model is used to extract metadata information such as knowledge title, author, publication time, keywords, abstract, research area, and knowledge system from the completed lightweight markup language text, and form knowledge entity relation triples. Similarly, this invention does not limit the specific form of the aerospace knowledge extraction prompts. In one embodiment, the aerospace knowledge extraction prompts include roles (e.g., defining the large language model as a semantic analysis engineer of the aerospace knowledge system), objectives (e.g., requiring the large language model to automatically extract relevant metadata information and form knowledge entity relation triples), constraints (e.g., requiring the large language model to ensure the integrity and usability of aerospace knowledge during the extraction process), output format (e.g., requiring the large language model to output the knowledge entity relation triples as a graphical knowledge graph), workflow (used to guide the large language model to perform data processing and extraction according to a specific process), and examples (e.g., input objects, output results, and three-dimensional ontology frameworks).

[0070] Figure 6 This is a schematic diagram of the overall process of constructing the aerospace application knowledge graph provided by the present invention, which can be referred to. Figure 6 As shown, after obtaining the aerospace files and aerospace table data, multiple knowledge blocks are obtained through knowledge block processing (such as...). F 1. F 2 and F 3, etc.), and perform corresponding processing on the data type of each knowledge segment (such as tabular data, text data, or image data, etc.). For example, for tables... T 1. Perform preprocessing to obtain the preprocessed table. T 2; Regarding the text W 1. Perform geocoding algorithms to obtain text. W Geographic information in 1; image processing using a Visual Language Model (VLM). P 1. Process the image to obtain the corresponding image description. P 2. Alternatively, these knowledge chunks can be input into a Large Language Model (LLM) to extract relevant image descriptions. P 2. Further, the processed data is converted into Markdown and then segmented into multiple chunks. M 1. M 2, ..., M nThese data chunks are then input into a large language model (LLM) to construct an aerospace knowledge graph. Simultaneously, the processed multimodal data is stored in the corresponding database. Finally, in the intelligent question-and-answer process for aerospace applications, upon receiving a user's question, the LLM is used to retrieve keywords. Based on the retrieval requirements, relevant results are retrieved from a graph database storing the aerospace knowledge graph and a vector database storing lightweight markup language text. These results are then input into the LLM to obtain the corresponding answer and document ID. For questions requiring tracing back to the original text, the corresponding original aerospace knowledge file can also be retrieved from a relational database based on the document ID.

[0071] This invention combines multi-source heterogeneous information to propose a method and question-answering system for constructing a multimodal aerospace application knowledge graph through large language model collaboration. It introduces the deep semantic understanding capabilities of a large language model and the image understanding capabilities of a visual language model to construct a professional aerospace application knowledge system and an intelligent knowledge storage architecture generation technology. This overcomes the limitations of traditional single-modal knowledge representation, enabling multi-source information coupling and nested structural semantic recognition in multiple scenarios. It provides a structured knowledge foundation with associative characteristics for downstream applications of knowledge graphs, such as intelligent question-answering systems in the aerospace knowledge domain, effectively solving the problem of collaborative association of multi-dimensional knowledge elements in aerospace application scenarios.

[0072] This invention collects multimodal data on aerospace application knowledge, designs a multi-stage, hierarchical data processing flow for text, images, and mixed text-image data, forms a standardized multimodal aerospace data processing framework, and proposes an image semantic collaborative parsing framework based on a visual language model, providing innovative technical support for the in-depth mining and intelligent application of aerospace application knowledge.

[0073] Furthermore, this invention addresses key issues in the current construction and data storage of aerospace knowledge systems, such as low construction efficiency and insufficient utilization of knowledge data. It proposes a three-layer heterogeneous knowledge storage architecture construction method integrating a large language model. Through in-depth mining of domain knowledge, it constructs an aerospace knowledge ontology with domain semantic depth and a hierarchical knowledge architecture. It also establishes a collaborative extraction and completion method system for aerospace knowledge graphs driven by a large language model, achieving entity relationship extraction and cross-modal knowledge completion. Through multi-dimensional knowledge fusion and standardization, this invention achieves classified and hierarchical storage and management of knowledge metadata, knowledge file entities, and knowledge traceability information through the organic integration of vector databases, graph databases, and relational databases. Furthermore, it designs a retrieval and generation system for question-answering applications, providing a high-precision and highly reliable semantic knowledge infrastructure for downstream tasks such as intelligent question answering and decision support.

[0074] The aerospace application knowledge graph construction device provided by the present invention is described below. The aerospace application knowledge graph construction device described below and the aerospace application knowledge graph construction method described above can be referred to in correspondence.

[0075] Figure 7 This is a schematic diagram of the aerospace application knowledge graph construction device provided by the present invention, as shown below. Figure 7 As shown, this invention provides a space application knowledge graph construction device, including a multimodal data acquisition module 701, a multimodal data processing module 702, a space application knowledge extraction module 703, and a space application knowledge graph construction module 704. The multimodal data acquisition module 701 is used to acquire target space application knowledge multimodal data from various space application knowledge files; the multimodal data processing module 702 is used to process each modality data in the target space application knowledge multimodal data based on the data type of the multimodal data, obtaining lightweight markup language text and lightweight markup language text embedding vectors corresponding to each modality data; the space application knowledge extraction module 703 is used to extract knowledge based on preset space application knowledge... The system extracts prompt words and uses a large language model to perform aerospace knowledge extraction and knowledge completion processing on the lightweight markup language text, obtaining metadata information corresponding to each entity in the aerospace knowledge file. Based on the relationships between the entities in the aerospace knowledge file and the metadata information, it constructs knowledge entity relationship triples. The aerospace application knowledge graph construction module 704 is used to determine the association relationship between the aerospace knowledge file, the lightweight markup language text, the lightweight markup language text embedding vector, and the knowledge entity relationship triples based on the numbering information of the aerospace knowledge file, and constructs an aerospace application knowledge graph based on the association relationship and the knowledge entity relationship triples.

[0076] The aerospace application knowledge graph construction device provided by this invention obtains multimodal data of target aerospace knowledge from various aerospace knowledge files, processes different modal data according to data type, and transforms them into lightweight markup language text; then, using preset aerospace knowledge extraction prompts and a large language model, it extracts metadata information corresponding to each entity from the text, and constructs knowledge entity relationship triples by combining the relationships between entities; finally, based on the aerospace knowledge file number information, it clarifies the association relationship between the file, lightweight markup language text, embedding vector, and knowledge entity relationship triples, thereby constructing an aerospace application knowledge graph that can accurately present entities and relationships in the aerospace application field, providing efficient and accurate knowledge support for aerospace application knowledge question answering.

[0077] The apparatus provided in this embodiment of the invention is used to execute the above-described method embodiments. For specific processes and details, please refer to the above embodiments, which will not be repeated here.

[0078] Figure 8 This is a schematic diagram of the structure of the aerospace application knowledge question-and-answer system provided by the present invention, as shown below. Figure 8 As shown, this invention provides a space application knowledge question-and-answer system, including a question receiving module 801, a question parsing module 802, and a question-and-answer module 803. The question receiving module 801 is used to acquire a target space knowledge question; the question parsing module 802 is used to acquire space application knowledge graph information corresponding to the target space knowledge question from a graph database, wherein the space application knowledge graph information is generated based on the space application knowledge graph construction method described in the above embodiments; the question-and-answer module 803 is used to input the space application knowledge graph information into a large language model to obtain the question-and-answer result corresponding to the target space knowledge question.

[0079] In this invention, the question receiving module 801 serves as the entry point for interaction between the aerospace application knowledge question-and-answer system and the user terminal, and is responsible for acquiring the target aerospace knowledge questions raised by the user terminal. Specifically, the question receiving module 801 monitors the user terminal's input behavior in real time, and can receive the questions promptly and accurately, whether the questions are entered as text via keyboard or converted into text form via speech recognition technology.

[0080] The problem parsing module 802 is used to perform in-depth analysis and processing of the target aerospace knowledge problem proposed by the user, so as to accurately obtain aerospace application knowledge graph information related to the problem from the graph database. In this invention, the problem parsing module 802 first performs semantic analysis on the target aerospace knowledge problem, identifying keywords, entities (such as author names, algorithm names, etc.) and problem types (such as factual problems, reasoning problems, etc.) in the problem. Then, based on its understanding of the problem, the problem parsing module 802 retrieves the aerospace application knowledge graph information corresponding to the problem from the pre-constructed graph database. In this invention, the aerospace application knowledge graph information is generated based on the aerospace application knowledge graph construction method described in the above embodiments, and contains rich and structured knowledge in the aerospace field, such as satellite information, space mission details, and explanations of astronomical phenomena. Through efficient querying and matching in the graph database, the problem parsing module 802 can find the knowledge graph fragment most relevant to the target problem.

[0081] The question-answering module 803 takes the aerospace application knowledge graph information obtained by the question parsing module 802 as input and uses a large language model to generate question-answering results corresponding to the target aerospace knowledge question. Specifically, the question-answering module 803 inputs the aerospace application knowledge graph information retrieved from the graph database into the large language model. The large language model has powerful semantic understanding and knowledge reasoning capabilities, and can process structured and unstructured knowledge information. The large language model performs in-depth analysis of the input knowledge graph, combines the questions raised by the user, and uses its own knowledge reserves and reasoning logic to generate accurate, clear, and easy-to-understand answers.

[0082] In this invention, the aerospace application knowledge question-and-answer system can retrieve information from multiple databases to address different types of questions raised by users, thus achieving a cross-database association mechanism for question-and-answer generation. In this invention, graph databases, vector databases, and relational databases are linked through the numbering information of aerospace knowledge files.

[0083] Specifically, when the aerospace application knowledge question-and-answer system processes questions raised by users about aerospace knowledge itself, it will first perform a retrieval operation on the graph database to obtain the corresponding knowledge node entities. Figure 9 The diagram illustrating the database retrieval and question-answer generation process for different types of questions provided by this invention can be referred to. Figure 9 As shown, taking the query "What papers does author A have?" as an example, after receiving the question, the aerospace application knowledge question answering system first searches the graph database, extracts the target author node and its associated paper nodes, and obtains the document ID attribute contained in the paper nodes. This structured data is then input into the large language model to generate the answer to the question.

[0084] When the aerospace application knowledge question-answering system processes queries involving knowledge document content, it will first trigger the vector database retrieval mechanism. Through semantic matching strategies, it will extract relevant text and its corresponding document IDs (i.e., numbering information), and transmit this structured data to a large language model to generate accurate responses based on the original knowledge documents. (See reference...) Figure 9 As shown, taking "What business applications can be supported by building data extracted using remote sensing intelligent interpretation methods" as an example, the aerospace application knowledge question and answer system will first perform vector database retrieval to obtain relevant chunk information on "remote sensing intelligent interpretation" and "building data" and their corresponding document IDs, and return them to the large language model to generate questions and answers.

[0085] When questions involve both knowledge document content and the knowledge itself, the aerospace application knowledge question-answering system can comprehensively search both graph and vector databases. It combines the knowledge document node information returned by the graph database with the relevant text information returned by the vector database to generate the question and answer. Taking the question "What work has been done by a certain entity in assessing the geographical potential of distributed photovoltaic power after 2023?" as an example, the aerospace application knowledge question-answering system will extract keywords such as "a certain entity" and "2023" to search the graph database, returning corresponding node information and document ID information. Simultaneously, the system will query relevant chunk data on "assessment of geographical potential of distributed photovoltaic power" in the vector database and return the corresponding document ID. Finally, all these elements are input into the large language model to generate the question and answer.

[0086] Figure 10 This is a schematic diagram of the question-and-answer process for spatial regions in aerospace knowledge files provided by the present invention, which can be used as a reference. Figure 10 As shown, in this invention, when the aerospace application knowledge question-answering system processes spatial region-related query tasks, it performs a graph database search based on the geographical range of the target area specified by the user, extracting node entities and their attribute information associated with the geographical range of the target area. Then, the aerospace application knowledge question-answering system uses keywords to perform semantic matching retrieval based on a vector database to obtain relevant data blocks and corresponding document identifiers. Taking a typical interaction scenario: the user defines a target area through a map interaction tool and asks "What remote sensing resource surveys have been conducted in this area?" as an example, the aerospace application knowledge question-answering system first parses the geographical coordinates or spatial range of the area, retrieves matching spatial region nodes and their associated attributes in the graph database, then extracts keywords such as "remote sensing resource survey" from the user's query, drives the vector database to perform a semantic similarity-based search, returns relevant data blocks and document IDs containing the keywords, and finally integrates the two types of search results and inputs them into a large language model to generate an accurate response with spatial association information.

[0087] Figure 11 This is a schematic diagram illustrating the tracing process of the original aerospace knowledge documents provided by the present invention, which can be referred to. Figure 11 When users need to trace the source of cited content, the aerospace application knowledge Q&A system can perform a retrieval operation in the relational database through document identifiers (i.e. numbering information) to obtain the corresponding document URL, providing users with basic information support for original text viewing.

[0088] The aerospace application knowledge question-answering system provided by this invention obtains multimodal data of target aerospace knowledge from various aerospace knowledge files. Based on the data type, it performs targeted processing on different modalities of the data, transforming them into lightweight markup language text. Then, using preset aerospace knowledge extraction prompts and a large language model, it extracts metadata information corresponding to each entity from the text, and constructs knowledge entity relationship triples by combining the relationships between entities. Finally, based on the aerospace knowledge file number information, it clarifies the association between the file, the lightweight markup language text, the embedding vector, and the knowledge entity relationship triples, thereby constructing an aerospace application knowledge graph. This graph accurately presents entities and relationships in the aerospace application domain, providing efficient and accurate knowledge support for aerospace application knowledge question-answering.

[0089] Figure 12 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 12 As shown, the electronic device may include: a processor 1201, a communications interface 1202, a memory 1203, and a communication bus 1204, wherein the processor 1201, the communications interface 1202, and the memory 1203 communicate with each other through the communication bus 1204. The processor 1201 can call logical instructions in the memory 1203 to execute a method for constructing a space application knowledge graph. This method includes: acquiring target space knowledge multimodal data from various space knowledge files; processing each modality of the target space knowledge multimodal data based on its data type to obtain lightweight markup language text and lightweight markup language text embedding vectors corresponding to each modality; performing space knowledge extraction and knowledge completion processing on the lightweight markup language text based on preset space knowledge extraction prompts and a large language model to obtain metadata information corresponding to each entity in the space knowledge file, and constructing knowledge entity relationship triples based on the relationships between entities in the space knowledge file and the metadata information; determining the association between the space knowledge file, the lightweight markup language text, the lightweight markup language text embedding vectors, and the knowledge entity relationship triples based on the numbering information of the space knowledge file, and constructing a space application knowledge graph based on the association and the knowledge entity relationship triples.

[0090] Furthermore, the logical instructions in the aforementioned memory 1203 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0091] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer is able to execute the aerospace application knowledge graph construction method provided by the above methods, the method comprising: acquiring target aerospace knowledge multimodal data from various aerospace knowledge files; based on the data type of the multimodal data, performing data processing on each modality data in the target aerospace knowledge multimodal data respectively to obtain lightweight markup language text and lightweight markup language text embedding vectors corresponding to each modality data; based on... Based on preset aerospace knowledge extraction prompts and a large language model, aerospace knowledge extraction and knowledge completion processing are performed on the lightweight markup language text to obtain metadata information corresponding to each entity in the aerospace knowledge file. Based on the relationships between the entities in the aerospace knowledge file and the metadata information, knowledge entity relationship triples are constructed. Based on the numbering information of the aerospace knowledge file, the association relationship between the aerospace knowledge file, the lightweight markup language text, the lightweight markup language text embedding vector, and the knowledge entity relationship triples is determined. Based on the association relationship and the knowledge entity relationship triples, an aerospace application knowledge graph is constructed.

[0092] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the aerospace application knowledge graph construction method provided in the above embodiments. The method includes: acquiring target aerospace knowledge multimodal data from various aerospace knowledge files; processing each modality of the target aerospace knowledge multimodal data based on the data type of the multimodal data to obtain lightweight markup language text and lightweight markup language text embedding vectors corresponding to each modality; and extracting aerospace knowledge based on preset aerospace knowledge extraction prompts and a large language model. The lightweight markup language text is subjected to aerospace knowledge extraction and knowledge completion processing to obtain metadata information corresponding to each entity in the aerospace knowledge file. Based on the relationships between the entities in the aerospace knowledge file and the metadata information, a knowledge entity relationship triplet is constructed. Based on the numbering information of the aerospace knowledge file, the association relationship between the aerospace knowledge file, the lightweight markup language text, the lightweight markup language text embedding vector, and the knowledge entity relationship triplet is determined. Based on the association relationship and the knowledge entity relationship triplet, an aerospace application knowledge graph is constructed.

[0093] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0094] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing a knowledge graph for aerospace applications, characterized in that, include: Acquire multimodal data of target aerospace knowledge from various aerospace knowledge files; Based on the data type of multimodal data, data processing is performed on each modality of the target aerospace knowledge multimodal data to obtain the lightweight markup language text and lightweight markup language text embedding vector corresponding to each modality of data. Based on preset aerospace knowledge extraction prompts and a large language model, aerospace knowledge extraction and knowledge completion processing are performed on the lightweight markup language text to obtain metadata information corresponding to each entity in the aerospace knowledge file. Based on the relationship between each entity in the aerospace knowledge file and the metadata information, a knowledge entity relationship triplet is constructed. Based on the numbering information of the aerospace knowledge file, the association relationship between the aerospace knowledge file, the lightweight markup language text, the lightweight markup language text embedding vector, and the knowledge entity relationship triple is determined, and an aerospace application knowledge graph is constructed according to the association relationship and the knowledge entity relationship triple.

2. The method for constructing aerospace application knowledge graphs according to claim 1, characterized in that, After determining the association relationship between the aerospace knowledge file, the lightweight markup language text, the lightweight markup language text embedding vector, and the knowledge entity relation triple based on the numbering information of the aerospace knowledge file, and constructing an aerospace application knowledge graph based on the association relationship and the knowledge entity relation triple, the method further includes: A graph database is constructed based on the aforementioned aerospace application knowledge graph; A vector database is constructed based on the lightweight markup language text embedding vectors corresponding to the lightweight markup language texts. A relational database is constructed based on the aerospace knowledge file and the document access path information corresponding to the aerospace knowledge file; The graph database, the vector database, and the relational database are used to perform cross-database association retrieval through the association relationships during the knowledge question-and-answer process for aerospace applications.

3. The method for constructing aerospace application knowledge graphs according to claim 1, characterized in that, The target aerospace knowledge multimodal data includes at least text data, tabular data, and image data; The data type based on multimodal data involves processing each modality of the target aerospace knowledge multimodal data to obtain lightweight markup language text and lightweight markup language text embedding vectors corresponding to each modality, including: Lightweight language conversion processing is performed on the text data to obtain the lightweight markup language text and the lightweight markup language text embedding vector corresponding to the text data; The table data is preprocessed to obtain the lightweight markup language text and the lightweight markup language text embedding vector corresponding to the table data. The data preprocessing includes at least data deduplication, missing value imputation, outlier removal and format normalization. Obtain the string description content corresponding to the image data, and construct the lightweight markup language text and the lightweight markup language text embedding vector corresponding to the image data based on the string description content.

4. The method for constructing aerospace application knowledge graphs according to claim 3, characterized in that, The step of obtaining the string description content corresponding to the image data, and constructing the lightweight markup language text and the lightweight markup language text embedding vector corresponding to the image data based on the string description content, includes: Based on a visual language model, the image data is parsed to obtain the semantic information of the image data. The context information of the image data in the aerospace knowledge file is obtained, and the context information is input into the large language model to obtain the image and text supplementary information corresponding to the image data; A hash code value corresponding to the image data is generated, and the lightweight markup language text and the lightweight markup language text embedding vector corresponding to the image data are constructed based on the image semantic information, the image and text supplementary information and the hash code value.

5. The method for constructing aerospace application knowledge graphs according to claim 3, characterized in that, The method further includes: Based on the large language model, geographic information is identified in the lightweight markup language text. If there is geospatial semantics in the lightweight markup language text, the corresponding geographic unit information in the lightweight markup language text is obtained. Based on a preset geocoding algorithm, the geographic unit information is converted into corresponding geocoding information; The geocoding information is added to the lightweight markup language text to obtain lightweight markup language text with geocoding information.

6. The method for constructing aerospace application knowledge graphs according to any one of claims 1 to 5, characterized in that, The preset aerospace knowledge extraction prompts include aerospace knowledge completion prompts and aerospace knowledge extraction prompts; The process involves extracting and completing space-related knowledge from the lightweight markup language text based on preset space-related knowledge extraction prompts and a large language model. This yields metadata information corresponding to each entity in the space-related knowledge file. Furthermore, based on the relationships between entities in the space-related knowledge file and the metadata information, a knowledge entity relationship triplet is constructed, including: Based on the aforementioned aerospace knowledge completion prompts and the aforementioned large language model, knowledge completion information corresponding to the lightweight markup language text is generated within a three-dimensional ontology framework. The three-dimensional ontology framework includes the aerospace information's business system, technical system, and data system. Based on the knowledge completion information, the lightweight markup language text is completed to obtain the completed lightweight markup language text; Based on the aforementioned aerospace knowledge extraction prompts and the large language model, aerospace knowledge extraction processing is performed on the completed lightweight markup language text to obtain metadata information corresponding to each entity in the completed lightweight markup language text. Based on the relationships between each entity in the completed lightweight markup language text and the metadata information, the knowledge entity relationship triplet is constructed.

7. A device for constructing a knowledge graph for aerospace applications, characterized in that, include: The multimodal data acquisition module is used to acquire multimodal data of target aerospace knowledge from various aerospace knowledge files; The multimodal data processing module is used to process each modal data in the target aerospace knowledge multimodal data according to the data type of the multimodal data, and obtain the lightweight markup language text and lightweight markup language text embedding vector corresponding to each modal data. The aerospace knowledge extraction module is used to perform aerospace knowledge extraction and knowledge completion processing on the lightweight markup language text based on preset aerospace knowledge extraction prompts and a large language model, to obtain metadata information corresponding to each entity in the aerospace knowledge file, and to construct knowledge entity relationship triples based on the relationship between each entity in the aerospace knowledge file and the metadata information. The aerospace application knowledge graph construction module is used to determine the association relationship between the aerospace knowledge file, the lightweight markup language text, the lightweight markup language text embedding vector and the knowledge entity relationship triple based on the numbering information of the aerospace knowledge file, and to construct the aerospace application knowledge graph based on the association relationship and the knowledge entity relationship triple.

8. A space-based application knowledge question-and-answer system, characterized in that, include: The question receiving module is used to acquire target aerospace knowledge questions; The problem parsing module is used to obtain aerospace application knowledge graph information corresponding to the target aerospace knowledge problem from the graph database, wherein the aerospace application knowledge graph information is generated based on the aerospace application knowledge graph construction method according to any one of claims 1 to 6; The question-and-answer module is used to input the aerospace application knowledge graph information into the large language model to obtain the question-and-answer results corresponding to the target aerospace knowledge question.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the aerospace application knowledge graph construction method as described in any one of claims 1 to 6.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the aerospace application knowledge graph construction method as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Information processing device, method, and program for performing pseudospace formation and reference frame inference based on text input including ASCII layout.

    JP7882637B1