A method for constructing aerospace knowledge large model based on progressive knowledge injection and retrieval enhanced generation

By applying progressive knowledge injection and retrieval enhancement generation technology in the aerospace field, an efficient aerospace knowledge model was built, solving the problems of insufficient knowledge coverage and lag in traditional systems, and realizing in-depth knowledge integration and efficient application in professional fields.

CN119808931BActive Publication Date: 2025-05-23NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510309163.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-05-23
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

The traditional knowledge management and intelligent question-and-answer system in the field of aerospace has problems such as insufficient knowledge coverage, difficulty in analyzing complex engineering problems, and lag in timeliness of knowledge updates.

Method used

The aerospace knowledge big model construction method based on progressive knowledge injection and retrieval enhancement generation is adopted. Through a three-stage progressive hybrid course learning framework and dynamic parameter thawing strategy, the efficient internalization of domain knowledge is achieved, and the search enhancement generation module is designed to ensure professional compliance of generated content.

Benefits of technology

It has achieved in-depth integration and efficient application of knowledge in the field of aerospace, improved the consistency of knowledge coverage and facts, and solved the problems of term deviation, knowledge lag and logic weakness in traditional models in professional scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119808931B_ABST
    Figure CN119808931B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing a large model of aerospace knowledge based on progressive knowledge injection and retrieval enhancement generation, including: collecting multi-source knowledge data and preprocessing to construct an aerospace knowledge base; based on the DeepSeek‑R1‑8B model, using a three-stage progressive hybrid course learning framework for continuous pre-training; through instruction data and knowledge guidance mechanism, supervised fine-tuning is performed to complete the basic construction of the large model; constructing a retrieval enhancement generation module to form a "retrieval‑filtering‑generation" process, and through deep integration of real-time retrieval and generative reasoning, the knowledge coverage and factual consistency of the model are improved; through multi-dimensional quantitative indicators and dynamic testing mechanisms, the performance of the model is evaluated to form a complete evaluation ecology from data construction to feedback optimization. The method proposed in the present invention realizes the deep integration and efficient application of knowledge in the aerospace field through multi-stage collaborative optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a method for constructing an aerospace knowledge large model based on progressive knowledge injection and retrieval enhancement generation. Background Art

[0002] In the field of aerospace, traditional knowledge management and intelligent question-answering systems face significant technical bottlenecks and challenges. Although traditional general-purpose large language models perform well in open-domain tasks, they are seriously deficient in covering knowledge in professional fields. Such models usually lack a deep understanding of the aerospace terminology system and have difficulty accurately analyzing complex engineering problems. In addition, knowledge updates are highly dependent on manual annotation and retraining, resulting in delayed timeliness and inability to adapt to the needs of rapid iteration of aerospace technology.

[0003] With the breakthroughs in deep learning and retrieval-enhanced generation technology, the new generation of knowledge models has gradually shown its potential to solve the above problems. The progressive knowledge injection technology gradually decouples and strengthens domain knowledge from general corpus through phased course learning, so that the model can accurately internalize the core concepts and logical connections in the aerospace field while retaining basic language capabilities. Summary of the invention

[0004] In view of the deficiencies of the above-mentioned prior art, the present invention proposes a method for constructing a large model of aerospace knowledge based on progressive knowledge injection and retrieval enhancement generation. By designing a three-stage progressive hybrid course learning framework and combining it with a dynamic parameter unfreezing strategy, efficient internalization of domain knowledge is achieved; a retrieval enhancement generation module is designed to accurately recall technical documents from authoritative knowledge bases, and professional compliance of generated content is ensured through term constraint decoding and physical rule verification. The method proposed in the present invention aims to provide a highly reliable and explainable knowledge service base for aerospace intelligent applications.

[0005] In order to achieve the above technical objectives, the present invention provides the following technical methods:

[0006] A method for constructing an aerospace knowledge large model based on progressive knowledge injection and retrieval enhancement generation specifically comprises the following steps:

[0007] S1. Design an automated collection tool chain for various aerospace knowledge data sources, build a data collection system, and obtain raw data; perform fine-grained cleaning, deduplication, and organization on the acquired raw data to build a high-quality aerospace knowledge database;

[0008] S2. Based on the DeepSeek-R1-8B model and the knowledge database built in step S1, continuous pre-training is performed, a three-stage progressive hybrid course learning framework is designed, the data ratio in the pre-training database is dynamically planned, and the concentration and complexity of domain knowledge are gradually improved;

[0009] S3. Based on the pre-trained model, supervised fine-tuning is performed by fine-tuning the domain-specific instruction data and knowledge guidance mechanism, calibrating the model's generation logic and professional knowledge expression, and transforming the implicit embedded domain knowledge injected in the pre-training stage into explicit reasoning capabilities, completing the basic construction of the aerospace knowledge model.

[0010] S4. Build a retrieval enhancement generation module to form a "retrieval-filtering-generation" process, and improve the knowledge coverage and factual consistency of the big model in the aerospace field by deeply integrating real-time retrieval and generative reasoning;

[0011] S5. Build a professional evaluation benchmark for the aerospace field. Through multi-dimensional quantitative indicators and dynamic testing mechanisms, comprehensively evaluate the performance of the constructed large model in key tasks, and form a complete evaluation ecosystem from data construction to feedback optimization.

[0012] Furthermore, step S1 specifically includes:

[0013] S11. Plan multi-dimensional data sources and improve the knowledge system, including:

[0014] Capture authoritative academic literature in the field of aerospace to form a basic corpus covering multiple sub-fields; collect industry standards in the field of aerospace; obtain authorized versions of engineering manuals in the field of aerospace according to engineering practice needs;

[0015] S12. Build an automated collection system, including:

[0016] Build a distributed crawler cluster, integrate dynamic rendering, recursive crawling and IP proxy pool, and crawl web data legally; build a PDF processing chain, extract text through the PyMuPDF tool, use Tesseract 5.0 to process scans, combine the fine-tuned PP-OCRv3 model to achieve accurate formula recognition, and simultaneously generate MathML and natural language dual encoding;

[0017] S13. Design a structured conversion pipeline; parse the parameter tables in academic literature and technical manuals, generate "entity-attribute-value" structured triples through the template engine, and establish machine-understandable semantic associations;

[0018] S14. Build a data cleaning and standardization pipeline, including:

[0019] Delete and correct erroneous data and outliers to reduce interference with model training; unify data formats and encoding methods to ensure smooth integration and processing of data; use regular expressions to remove non-technical content; train classifiers to score text technical density and filter content with too low a score; perform fine-grained deduplication and redundancy elimination on data, perform entity linking on multi-source descriptions of the same event to avoid overfitting caused by data duplication; construct an aerospace terminology comparison table and perform batch replacement to achieve terminology standardization;

[0020] S15. The cleaned and standardized data are classified and stored according to the source, and the database is accurately annotated using an automated metadata generation pipeline;

[0021] S16. Use the distributed storage system to store the large-scale text knowledge, complete the construction of a high-quality aerospace knowledge database, and build an aerospace knowledge graph based on the knowledge database.

[0022] Furthermore, step S2 specifically includes:

[0023] S21. Design a three-stage progressive hybrid course learning framework, dynamically sample and mix the high-quality aerospace domain knowledge database constructed in step S1 and the general database open source on the Internet, and construct pre-training data sets suitable for three different learning stages; the three different learning stages are the general knowledge retention stage, the domain knowledge injection stage, and the deep specialization stage, which correspond to the early, middle, and late stages of pre-training respectively;

[0024] S22. In the early stage of pre-training, the pre-training dataset contains 90% general corpus and 10% basic data in the aerospace field. Based on the original DeepSeek-R1-8B model, the masked language model task is carried out, and multiple rounds of training are performed, focusing on learning the semantic embedding of core terms in the field and the association of basic concepts. At the same time, a dynamic masking mechanism is designed to mask the field terms in the input text, forcing the model to infer the deep semantics of professional vocabulary from the context.

[0025] S23. In the middle of pre-training, the proportion of aerospace data in the pre-training data set increased to 50%, with a focus on introducing medium-difficulty text content; based on the DeepSeek-R1-8B model that has undergone initial pre-training, the masked language model task continued, with multiple rounds of training to improve the model's ability to model complex associations;

[0026] S24. In the later stage of pre-training, the proportion of aerospace data in the pre-training data set increased to 80%, focusing on high-density professional content, especially including the derivation of complex mathematical formulas; based on the DeepSeek-R1-8B model that has undergone mid-term pre-training, the masked language model and next structure prediction tasks were trained for multiple rounds to enhance the model's logical reasoning, context understanding, and information integration capabilities;

[0027] S25. Introduce a numerical protection layer in the process of generating model results, and automatically mark " <num>"and" <unit>" tag, constraining the physical rationality of the numerical format and avoiding the generation of results that go against common sense.

[0028] Furthermore, step S3 specifically includes:

[0029] S31. Following the principles of "task diversity" and "knowledge completeness", a fine-tuning dataset is constructed based on a high-quality aerospace knowledge database;

[0030] S32. A dynamic parameter unfreezing strategy is adopted to train the model after continuous pre-training on five types of fine-tuning tasks based on five types of fine-tuning data sets. In the initial stage, 60% of the underlying model parameters containing common language patterns are frozen and do not participate in training. As the number of training rounds increases, the parameters are gradually unfrozen, and finally the full model is made adjustable, completing the basic construction of the aerospace knowledge large model.

[0031] Furthermore, step S31 specifically includes:

[0032] S311. Based on typical scenarios in the aerospace field, five core task templates are defined, millions of instruction-response pairs are generated through automated pipelines, and instances are generated by automatically filling in real data; the five core task templates are: interpretation, calculation, diagnosis, comparison, and design;

[0033] S312. To improve the model's anti-interference ability, adversarial enhancement is applied to the data: by randomly replacing some domain terms and confusing the unit system, the model is forced to maintain semantic consistency under noise interference;

[0034] S313, data cleaning stage, calculate the sample perplexity, filter out samples with too high perplexity, and conduct manual review through expert cross-review mechanism. Randomly select some data and have them independently labeled by multiple experts in the field. Only samples that are passed unanimously are retained to ensure the accuracy of technical details.

[0035] Furthermore, step S4 specifically includes:

[0036] S41, building a query enhancement module to rewrite the query input by the user using the large model;

[0037] S42. Build a query checking module to check the rewritten new query in three aspects, including:

[0038] Determine whether the query entered by the user is related to aerospace knowledge; evaluate whether the query entered by the user is suitable for knowledge retrieval enhancement to generate a response; verify whether the user's query meets the specification requirements, intercept restrictive content, and ensure the compliance of the query; if and only if the three conditions are met at the same time, the aerospace knowledge big model will continue to perform knowledge retrieval enhancement;

[0039] S43, constructing a query expansion and decomposition module to handle diverse user queries, and performing hybrid granularity adaptive retrieval according to the complexity of the input user query, generating subqueries through the original query, and documents related to the subqueries retrieved by the large model;

[0040] S44, construct a document filtering module to evaluate the relevance between the retrieved documents and their corresponding subqueries; and comprehensively sort them according to the source authority, timeliness, and relevance scores, and retain the three documents with the highest comprehensive scores; if there is a conflict in the search results, start the voting mechanism and mark the doubtful points;

[0041] S45. The documents that have been filtered through multiple layers are then combined with their corresponding sub-queries and used as the input of the model together with the rewritten user query to guide the response generation of the large model. In addition, when the user query does not obtain the corresponding relevant documents, the large model is prompted to provide appropriate explanations.

[0042] Further, step S43 specifically includes:

[0043] For user queries that only involve a single domain-specific entity, the user query is expanded to fill in the general attributes of the relevant entity to generate a subquery; for user queries that contain multiple domain-specific entities, the query is split, and the association dimensions between entities are extracted based on the aerospace knowledge graph, and the complex original query is rewritten into an easy-to-retrieve form to generate a structured subquery; the subqueries generated by expansion and splitting are used to obtain the documents corresponding to each subquery.

[0044] Furthermore, step S5 specifically includes:

[0045] S51. Following the principles of "domain coverage", "difficulty level" and "dynamic evolution", a test set covering four major task types is constructed based on a high-quality aerospace knowledge database. The difficulty of the test questions is dynamically calibrated through the item response theory model, and the discrimination parameters and difficulty parameters are calculated. Questions with low discrimination and unbalanced difficulty are eliminated to ensure that the test set can effectively distinguish large models of different ability levels.

[0046] S52. A composite weighting scheme is used to form an evaluation index system, and the comprehensive scores obtained include: terminology accuracy, calculation accuracy, logical coherence and response speed;

[0047] The accuracy of terms is verified by knowledge graph entity linking; the calculation accuracy relies on the SymPy symbolic computing engine to parse and generate mathematical expressions in the text and verify the rationality of the values; the logical coherence uses the pre-trained Coherence-BERT model to evaluate the integrity of the causal chain between paragraphs;

[0048] S53. In response to the continuous evolution of domain knowledge, a dynamic assessment platform is built to conduct adaptive difficulty testing: dynamically adjust the difficulty of questions based on the real-time performance of the big model. If the correct rate is >90% after answering multiple questions continuously, a higher-level task will be automatically pushed; if the error rate is >30%, it will be downgraded to the basic question type to prevent assessment bias;

[0049] S54. Introduce a virtual-reality verification mechanism: simulate the digital twin environment in a virtual scene, requiring the large model to output solutions that conform to physical laws; deploy the model in parallel at the engineering site, compare the consistency of its suggestions with the engineers' decisions, and calculate the degree of consistency;

[0050] S55. By continuously absorbing progress in the field and user feedback, a dynamic closed-loop evaluation system of "test-optimize-retest" is formed to ensure that the knowledge model maintains professionalism and practicality during rapid iteration.

[0051] Based on the above technical solution, the present invention has at least the following beneficial effects:

[0052] The method proposed in the present invention aims at the characteristics of aerospace field, which is knowledge-intensive, dynamic and in high demand. It realizes the deep integration and efficient application of aerospace field knowledge based on the collaborative optimization of multiple stages of "pre-training-fine-tuning-knowledge enhanced retrieval". It balances general and domain capabilities through progressive knowledge injection, enhances generation reliability through multi-granularity retrieval fusion, and realizes accurate capability diagnosis through dynamic evaluation system, thus solving the pain points of traditional large models in professional scenarios, such as terminology deviation, knowledge lag, weak logic, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a flow chart of the method for constructing a large model of aerospace knowledge proposed in the present invention;

[0054] Figure 2 This is the retrieval enhancement generation module diagram proposed by the present invention. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0056] Although the steps in the present invention are arranged with numbers, they are not used to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a certain step requires other steps as a basis, the relative order of the steps can be adjusted. It is understood that the term "and / or" used in this article involves and covers any and all possible combinations of one or more of the associated listed items.

[0057] like Figure 1 As shown, a method for constructing an aerospace knowledge large model based on progressive knowledge injection and retrieval enhancement generation specifically includes the following steps:

[0058] S1. Design an automated collection tool chain, build a data collection system, and obtain raw data from a variety of aerospace knowledge data sources (including public data sets, web page data, industry-specific data, and related books). Perform fine-grained cleaning, deduplication, and organization of the acquired raw data to build a high-quality aerospace knowledge database.

[0059] As a preferred implementation, step S1 specifically includes:

[0060] S11. Plan multi-dimensional data sources and improve the knowledge system. In this embodiment, the specific steps are as follows:

[0061] Prioritize crawling authoritative academic literature such as public technical reports, public databases, and public aerospace thesis collections to form a basic corpus covering sub-fields such as propulsion technology, materials science, and aerodynamics;

[0062] Systematically collect industry standards in the aerospace field; including: relevant regulatory documents, aerospace parts of industry standards;

[0063] To meet the needs of engineering practice, obtain authorized versions of engineering manuals such as maintenance manuals, technical manuals, etc. for various types of aerospace vehicles;

[0064] S12. Build an automated collection system, including:

[0065] Build a distributed crawler cluster, integrate dynamic rendering, recursive crawling and IP proxy pool, and crawl web data legally; build a PDF processing chain, extract text through the PyMuPDF tool, use Tesseract 5.0 to process scans, combine the fine-tuned PP-OCRv3 model to achieve accurate formula recognition, and simultaneously generate MathML and natural language dual encoding; for example, convert " ” is converted to “the velocity increment is equal to the specific impulse times the product of the standard gravitational acceleration and the natural logarithm of the mass ratio”;

[0066] S13. Design a structured conversion pipeline; parse the parameter tables in academic literature and technical manuals, generate "entity-attribute-value" structured triples through the template engine, and establish machine-understandable semantic associations; for example, map the parameter table entry "thrust: 4152kN" to (RD-180 engine, vacuum thrust, 4152kN), and establish machine-understandable semantic associations;

[0067] S14. Build a data cleaning and standardization pipeline, including:

[0068] Delete and correct erroneous data and outliers to reduce interference with model training;

[0069] Unify data formats and coding methods to ensure smooth integration and processing of data;

[0070] Use regular expressions to remove non-technical content; train a classifier to score the technical density of the text and filter out content with low technical density scores; in this embodiment, use the open source RoBERTa-base classifier neural network model to evaluate the technical density of the text (the technical density of the text is mainly used to measure the density of professional terms, technical parameters and logical complexity in the aerospace field in the text), calculate the proportion of professional vocabulary in the text, output the density probability of the text, give the technical density score of the text, and filter out low-quality content with a score less than 0.85;

[0071] Fine-grained data deduplication and redundancy elimination, entity linking of multi-source descriptions of the same event to avoid overfitting caused by data duplication; building an aerospace terminology comparison table and performing batch replacement to achieve terminology standardization (e.g. "liquid oxygen kerosene" → "LOX / Kerosene");

[0072] S15. The cleaned and standardized data are classified and stored according to the source, and the database is accurately annotated using an automated metadata generation pipeline;

[0073] S16. Use the distributed storage system to store the large-scale text knowledge, complete the construction of a high-quality aerospace knowledge database, and build an aerospace knowledge graph based on the knowledge database.

[0074] S2. Based on the DeepSeek-R1-8B model, continuous pre-training is carried out, and a three-stage progressive hybrid course learning framework is designed. The data ratio in the pre-training database is dynamically planned to gradually improve the concentration and complexity of domain knowledge. In this way, the progressive hybrid course learning framework deeply integrates the professional knowledge of the aerospace field into the parameter space of the basic model, ensuring that the model has a preliminary understanding of the aerospace field.

[0075] As a preferred implementation, step S2 specifically includes:

[0076] S21. Design a three-stage progressive hybrid course learning framework, dynamically sample and mix the high-quality aerospace domain knowledge database constructed in step S1 and the general database open source on the Internet, and construct pre-training data sets suitable for three different learning stages; the three different learning stages are the general knowledge retention stage, the domain knowledge injection stage, and the deep specialization stage, which correspond to the early, middle, and late stages of pre-training respectively;

[0077] S22. In the early stage of pre-training, the pre-training dataset contains 90% general corpus and 10% basic data in the aerospace field, such as popular science articles and textbook glossaries; the masked language model task is performed based on the original DeepSeek-R1-8B model, and 100,000 rounds of training are performed, focusing on the semantic embedding of core terms in the learning field and the association of basic concepts;

[0078] It should be noted that the Masked Language Model (MLM) is a common pre-training task for large model training, which is mainly used to train large models to understand and generate natural language text. The core idea is to randomly mask some words in the input text and let the model predict these masked words based on the context. This training method enables the model to learn the meaning of each word in a bidirectional context, thereby grasping deeper semantic information;

[0079] Here, a dynamic masking mechanism is designed to impose a masking probability of up to 45% on domain terms in the input text (such as "specific impulse" and "Mach number"), forcing the model to infer the deep semantics of professional vocabulary from the context;

[0080] S23. In the middle of pre-training, the proportion of aerospace data in the pre-training data set increased to 50%, with a focus on introducing medium-difficulty text content; based on the DeepSeek-R1-8B model that had undergone initial pre-training, the masked language model task continued, with 200,000 rounds of training, focusing on improving the model's ability to model complex relationships such as "the composition of rocket fuel-the thrust generated by rocket fuel";

[0081] S24. In the later stage of pre-training, the proportion of aerospace data in the pre-training data set increased to 80%, focusing on high-density professional content, especially including the derivation of complex mathematical formulas; based on the DeepSeek-R1-8B model that has undergone mid-term pre-training, the masked language model and next structure prediction tasks were trained for 100,000 rounds respectively, to enhance the model's logical reasoning, context understanding, and information integration capabilities;

[0082] Next Sentence Prediction (NSP) is a common pre-training task for large model training. It aims to help the model understand the relationship between sentences and enhance its ability to model text structure and semantic coherence. The task goal is to determine whether two sentences are semantically coherent or continuous in the original text. Specifically, the model needs to predict whether sentence B is the next sentence of sentence A. This training method enables the model to learn sentence-level semantic relationships, thereby better understanding the logical structure of the text.

[0083] S25. Introduce a numerical protection layer in the process of generating model results, and automatically annotate through regular expressions" <num>"and" <unit>" tag, constraining the physical rationality of the numerical format and avoiding the generation of results that go against common sense.

[0084] S3. Based on the pre-trained model, supervised fine-tuning is performed by fine-tuning the domain-specific instruction data and knowledge guidance mechanism, calibrating the model's generation logic and professional knowledge expression, and transforming the implicit embedded domain knowledge injected in the pre-training stage into explicit reasoning capabilities, completing the basic construction of the aerospace knowledge model.

[0085] As a preferred implementation, step S3 specifically includes:

[0086] S31. Following the principles of "task diversity" and "knowledge completeness", a fine-tuning dataset is constructed based on a high-quality aerospace knowledge database;

[0087] More specifically, step S31 includes:

[0088] S311. Based on typical scenarios in the aerospace field, five core task templates are defined, millions of instruction-response pairs are generated through automated pipelines, and instances are generated by automatically filling in real data; the five core task templates are: interpretation, calculation, diagnosis, comparison, and design, as shown in Table 1 below:

[0089] Table 1 Five core task templates

[0090] ;

[0091] S312. To improve the model's anti-interference ability, 20% of the data is subjected to adversarial enhancement: by randomly replacing 15% of the domain terms (such as replacing "specific impulse" with "specific thrust") and confusing the unit system (such as alternating between "kN" and "lbf"), the model is forced to maintain semantic consistency under noise interference;

[0092] S313, data cleaning stage, calculate the sample perplexity, filter out samples with perplexity > 2.5, and conduct manual review through expert cross-review mechanism. Randomly select some data and have multiple experts in the field independently annotate them, and only keep samples that pass unanimously to ensure the accuracy of technical details;

[0093] In this embodiment, the GPT-2 Small model is used to calculate the sample perplexity (used to measure the prediction ability of the language model for a text, reflecting the "understandability" of the text; the lower the perplexity, the more the text conforms to the language rules of the model and the higher the quality);

[0094] The specific method for calculating the perplexity is as follows: input the text sample into the open source model GPT-2Small (124M parameters) trained on the general corpus, divide the long text into a sequence of many words, obtain the predicted probability of each word output by the model, use the cross entropy loss function to calculate the prediction loss of each word, average the loss values ​​and exponentially obtain the sample perplexity PPL. The specific calculation formula is as follows:

[0095] ;

[0096] in, is the text sample to be evaluated, is the probability of the language model predicting the i-th word, and N is the total number of words in the text;

[0097] S32. In order to take into account both the retention of pre-trained knowledge and domain adaptability, a dynamic parameter unfreezing strategy is adopted to train the model after continuous pre-training for five types of fine-tuning tasks, with each training round lasting 50,000 rounds. In the initial stage, 60% of the underlying model parameters containing common language patterns are frozen and do not participate in training. After every 1,000 rounds of training, 5% of the parameters are unfrozen, ultimately achieving full model adjustability and completing the basic construction of the aerospace knowledge model.

[0098] The dynamic parameter unfreezing strategy enables the large model to focus on high-level semantic adaptation (such as contextual reasoning of professional terms) in the early stages of fine-tuning. As fine-tuning progresses, the underlying parameters can be gradually released to capture fine-grained domain features (such as rocket model naming rules).

[0099] In addition, this embodiment also introduces a knowledge constraint mechanism in the process of generating model results: that is, a domain term tree containing 27,000 nodes is constructed, and hard constraints are imposed on key terms such as "specific impulse" and "Mach number" to ensure that the generated content strictly complies with the predefined dictionary; for numerical parameters, regular expressions are used to force matching. <num> value< / num> <unit> unit< / unit> " format, intercepting possible physical dimension errors such as "thrust = 5000kN·m".

[0100] S4. Build a retrieval enhancement generation module to form a "retrieval-filtering-generation" process, and improve the knowledge coverage and factual consistency of the big model in the aerospace field by deeply integrating real-time retrieval and generative reasoning;

[0101] As a preferred implementation, step S4 specifically includes:

[0102] S41. Build a query enhancement module and use the big model to perform query rewriting. When a user is talking to the big model, the user will enter a query and the big model will output the corresponding answer. However, the user's query description is not always accurate, so some rewriting enhancement should be performed first to accurately identify the user's intention and provide more satisfactory results.

[0103] Specifically, if Figure 2 As shown in the figure, the big model will forcibly replace non-standard expressions (such as "specific thrust" → "specific impulse") based on the pre-built aerospace term comparison table to ensure accurate interpretation of user queries; when users have conversations with the big model, the multiple rounds of dialogues often contain a large number of references and entity references. The aerospace knowledge big model will integrate the aerospace knowledge graph and build a context entity chain based on the previous dialogue history to map the fuzzy expressions in the user query to specific entities to eliminate the reference ambiguity and clarify the reference in the user query. For example, when the user asks "the thrust of a certain rocket" and "its fuel type" in succession, the system automatically binds "it" to the entity "a certain rocket";

[0104] S42. Build a query checking module to check the rewritten new query in three aspects, including:

[0105] Determine whether the query entered by the user is related to aerospace knowledge; evaluate whether the query entered by the user is suitable for knowledge retrieval enhancement to generate a response; verify whether the user's query meets the specification requirements, intercept restrictive content, and ensure the compliance of the query; if and only if the three conditions are met at the same time, the aerospace knowledge big model will continue to perform knowledge retrieval enhancement;

[0106] S43, constructing a query expansion and decomposition module to handle diverse user queries, and performing hybrid granularity adaptive retrieval according to the complexity of the input user query, generating subqueries through the original query, and documents related to the subqueries retrieved by the large model;

[0107] The complexity of the query input by the user is positively correlated with the number of domain-specific knowledge entities contained in the query. Since a single query may contain multiple entities, it will affect the large model's ability to generate a satisfactory response and interfere with the large model's ability to accurately retrieve documents related to the question. Therefore, this application constructs a query expansion and decomposition module, namely

[0108] For user queries involving only a single domain-specific entity, the user query is expanded to fill in the general attributes of the relevant entity and generate subqueries; for user queries involving multiple domain-specific entities, the query is split, and the associated dimensions between entities are extracted based on the aerospace knowledge graph, and the complex original query is rewritten into an easy-to-retrieve form to generate structured subqueries for subsequent retrieval, thereby improving the recall rate of relevant document retrieval;

[0109] Through the processing of step S43, when the user enters a query, the result output by the large model is a group of: sub-queries and keywords; wherein the sub-query field contains a list of sub-query strings obtained by rewriting the user input, and the keyword field provides a list of keywords corresponding to each sub-query.

[0110] For the subqueries generated by expanding or decomposing the original query, this query expansion and decomposition module will sequentially perform two steps: multi-path parallel recall and two-stage re-ranking screening, so as to obtain the documents corresponding to each subquery.

[0111] S44. Build a document filtering module to evaluate the relevance between the retrieved documents and their corresponding subqueries; sort them by source authority, timeliness (weight for the past three years × 1.5), and relevance score (DPR similarity), and retain the three documents with the highest comprehensive scores; if there is a conflict in the search results (such as a difference of >10% in the thrust value of a certain engine), start the voting mechanism and mark the doubtful points;

[0112] S45. The documents that have been filtered through multiple layers are then combined with their corresponding sub-queries and used as the input of the model together with the rewritten user query to guide the response generation of the large model. In addition, when the user query does not obtain the corresponding relevant documents, the large model is prompted to provide appropriate explanations.

[0113] S5. Build a professional evaluation benchmark for the aerospace field. Through multi-dimensional quantitative indicators and dynamic testing mechanisms, comprehensively evaluate the performance of the built large model in key tasks, and form a complete evaluation ecosystem from data construction to feedback optimization;

[0114] As a preferred implementation, step S5 specifically includes:

[0115] S51. Following the principles of "domain coverage", "difficulty level" and "dynamic evolution", a test set covering four major task types is constructed based on a high-quality aerospace knowledge database;

[0116] In this embodiment, the four major task types are: term explanation, numerical calculation, multi-hop reasoning and fault diagnosis, and the proportion of each type of data is 20%, 30%, 35% and 15% respectively; the data sources include theoretical derivations in authoritative academic papers, operational cases in technical manuals, records of real engineering problems and propositions by experts in the field. For example, the numerical calculation question "Given that the vacuum specific impulse of a rocket engine is 320s and the fuel flow rate is 50kg / s, calculate the theoretical thrust" must be strictly associated with the Tsiolkovsky equation and the data source must be marked; the difficulty of the test questions is dynamically calibrated through the item response theory model, and the discrimination parameter a (target>0.8) and the difficulty parameter b (interval [-2,2]) are calculated to eliminate questions with low discrimination (a < 0.5) or unbalanced difficulty ( ) to ensure that the test set can effectively distinguish models of different ability levels;

[0117] S52. Adopt a composite weighting scheme to form an evaluation index system. The comprehensive scores obtained include: term accuracy rate, calculation accuracy rate, logical coherence degree, and response speed. In this embodiment, the weight settings for each part of the scores are 40%, 30%, 20%, and 10%, giving priority to being accurate and error-free.

[0118] Among them, the term accuracy rate is verified through knowledge graph entity linking (such as detecting whether "specific impulse" is accurately mapped to the Specific Impulse node); the calculation accuracy rate relies on the SymPy symbolic calculation engine to parse and generate mathematical expressions in the text and verify the numerical rationality (such as judging the correctness of " "); the logical coherence degree uses the pre-trained Coherence-BERT model to evaluate the integrity of the causal chain between paragraphs (such as whether the fault diagnosis steps cover the entire process of "anomaly detection - root cause analysis - solution plan").

[0119] S53. To cope with the continuous evolution of domain knowledge, build a dynamic evaluation platform for difficulty adaptive testing: Dynamically adjust the question difficulty based on the real-time performance of the large model. If the accuracy rate > 90% after continuously answering 5 questions, automatically push higher-order tasks (such as orbital transfer optimization design); if the error rate > 30%, degrade to basic question types (such as term matching) to prevent evaluation bias.

[0120] S54. Introduce a virtual-real combination verification mechanism: Simulate digital twin environments such as hypersonic aerodynamic heating and extreme low-temperature propellant storage in a virtual scenario, and require the model to output a thermal protection plan that conforms to physical laws (such as the selection of thermal protection materials); Deploy the model to run in parallel at the engineering site, compare the consistency between its suggestions and the decisions of engineers, and calculate the coincidence degree. In this embodiment, the coincidence degree is set to reach 85%.

[0121] S55. Through continuously absorbing domain progress and user feedback, form a dynamic closed-loop evaluation system of "test - optimization - retest" to ensure that the knowledge large model maintains professionalism and practicality during rapid iteration.

[0122] In summary, the construction and training of the aerospace knowledge large model based on progressive knowledge injection and retrieval-enhanced generation are realized. The present invention balances general and domain capabilities through progressive knowledge injection, enhances the reliability of generation through multi-granularity retrieval fusion, and realizes accurate ability diagnosis through a dynamic evaluation system, solving the pain points of traditional large models such as term deviation, knowledge lag, and weak logic in professional scenarios, and providing a feasible technical path for the intelligent upgrade of aerospace.

[0123] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

[0124] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.< / unit> < / num> < / unit> < / num>

Claims

1. A method for constructing aerospace knowledge big model based on progressive knowledge injection and retrieval enhancement generation, characterized in that: The specific steps include: S1. Design an automated collection tool chain for various aerospace knowledge data sources, build a data collection system, and obtain raw data; perform fine-grained cleaning, deduplication, and organization on the acquired raw data to build a high-quality aerospace knowledge database; S2. Based on the DeepSeek-R1-8B model and the knowledge database built in step S1, continuous pre-training is performed, a three-stage progressive hybrid course learning framework is designed, the data ratio in the pre-training database is dynamically planned, and the concentration and complexity of domain knowledge are gradually improved; Step S2 specifically includes: S21. Design a three-stage progressive hybrid course learning framework, dynamically sample and mix the high-quality aerospace domain knowledge database constructed in step S1 and the general database open source on the Internet, and construct pre-training data sets suitable for three different learning stages; the three different learning stages are the general knowledge retention stage, the domain knowledge injection stage, and the deep specialization stage, which correspond to the early, middle, and late stages of pre-training respectively; S22. In the early stage of pre-training, the pre-training dataset contains 90% general corpus and 10% basic data in the aerospace field. Based on the original DeepSeek-R1-8B model, the masked language model task is carried out, and multiple rounds of training are performed, focusing on learning the semantic embedding of core terms in the field and the association of basic concepts. At the same time, a dynamic masking mechanism is designed to mask the field terms in the input text, forcing the model to infer the deep semantics of professional vocabulary from the context. S23. In the middle of pre-training, the proportion of aerospace data in the pre-training data set increased to 50%, with a focus on introducing medium-difficulty text content; based on the DeepSeek-R1-8B model that has undergone initial pre-training, the masked language model task continued, with multiple rounds of training to improve the model's ability to model complex associations; S24. In the later stage of pre-training, the proportion of aerospace data in the pre-training data set increased to 80%, focusing on high-density professional content, especially including the derivation of complex mathematical formulas; based on the DeepSeek-R1-8B model that has undergone mid-term pre-training, the masked language model and next structure prediction tasks were trained for multiple rounds to enhance the model's logical reasoning, context understanding, and information integration capabilities; S25. Introduce a numerical protection layer in the process of generating model results, and automatically annotate through regular expressions" <num>"and" <unit> " tag, constrains the physical rationality of the numerical format and avoids generating results that go against common sense;< / unit> < / num> S3. Based on the pre-trained model, supervised fine-tuning is performed by fine-tuning the domain-specific instruction data and knowledge guidance mechanism, calibrating the model's generation logic and professional knowledge expression, and transforming the implicit embedded domain knowledge injected in the pre-training stage into explicit reasoning capabilities, completing the basic construction of the aerospace knowledge model. S4: Construct a search enhancement generation module to form a "search-filter-generate" process, and improve the knowledge coverage and factual consistency of the big model in the aerospace field by deeply integrating real-time search and generative reasoning; Step S4 specifically includes: S41, building a query enhancement module to rewrite the query input by the user using the large model; S42. Build a query checking module to check the rewritten new query in three aspects, including: Determine whether the query entered by the user is related to aerospace knowledge; evaluate whether the query entered by the user is suitable for knowledge retrieval enhancement to generate a response; verify whether the user's query meets the specification requirements, intercept restrictive content, and ensure the compliance of the query; if and only if the three conditions are met at the same time, the aerospace knowledge big model will continue to perform knowledge retrieval enhancement; S43, constructing a query expansion and decomposition module to handle diverse user queries, and performing hybrid granularity adaptive retrieval according to the complexity of the input user query, generating subqueries through the original query, and documents related to the subqueries retrieved by the large model; S44, construct a document filtering module to evaluate the relevance between the retrieved documents and their corresponding subqueries; and comprehensively sort them according to the source authority, timeliness, and relevance scores, and retain the three documents with the highest comprehensive scores; if there is a conflict in the search results, start the voting mechanism and mark the doubtful points; S45. The documents that have been filtered through multiple layers are then combined with their corresponding subqueries and used as the input of the model together with the rewritten user query to guide the generation of the response of the large model; in addition, when the user query does not obtain the corresponding relevant documents, the large model is prompted to provide appropriate explanations; S5. Build a professional evaluation benchmark for the aerospace field. Through multi-dimensional quantitative indicators and dynamic testing mechanisms, comprehensively evaluate the performance of the constructed large model in key tasks, and form a complete evaluation ecosystem from data construction to feedback optimization.

2. The method for constructing aerospace knowledge big model according to claim 1, characterized in that: Step S1 specifically includes: S11. Plan multi-dimensional data sources and improve the knowledge system, including: Capture authoritative academic literature in the field of aerospace to form a basic corpus covering multiple sub-fields; collect industry standards in the field of aerospace; obtain authorized versions of engineering manuals in the field of aerospace according to engineering practice needs; S12. Build an automated collection system, including: Build a distributed crawler cluster, integrate dynamic rendering, recursive crawling and IP proxy pool, and crawl web data legally; build a PDF processing chain, extract text through the PyMuPDF tool, use Tesseract 5.0 to process scans, combine the fine-tuned PP-OCRv3 model to achieve accurate formula recognition, and simultaneously generate MathML and natural language dual encoding; S13. Design a structured conversion pipeline; parse the parameter tables in academic literature and technical manuals, generate "entity-attribute-value" structured triples through the template engine, and establish machine-understandable semantic associations; S14. Build a data cleaning and standardization pipeline, including: Delete and correct erroneous data and outliers to reduce interference with model training; unify data formats and encoding methods to ensure smooth integration and processing of data; use regular expressions to remove non-technical content; train classifiers to score text technical density and filter content with too low a score; perform fine-grained deduplication and redundancy elimination on data, perform entity linking on multi-source descriptions of the same event to avoid overfitting caused by data duplication; construct an aerospace terminology comparison table and perform batch replacement to achieve terminology standardization; S15. The cleaned and standardized data are classified and stored according to the source, and the database is accurately annotated using an automated metadata generation pipeline; S16. Use the distributed storage system to store the large-scale text knowledge, complete the construction of a high-quality aerospace knowledge database, and build an aerospace knowledge graph based on the knowledge database.

3. The method for constructing aerospace knowledge big model according to claim 1, characterized in that: Step S3 specifically includes: S31. Following the principles of "task diversity" and "knowledge completeness", a fine-tuning dataset is constructed based on a high-quality aerospace knowledge database; S32. Adopt a dynamic parameter unfreezing strategy, and train the model after continuous pre-training on five types of fine-tuning tasks based on five types of fine-tuning data sets. In the initial stage, freeze 60% of the underlying model parameters containing common language patterns and do not participate in training. As the number of training rounds increases, gradually unfreeze the parameters, and finally make the whole model adjustable, completing the basic construction of the aerospace knowledge large model.

4. The method for constructing aerospace knowledge big model according to claim 3, characterized in that: Step S31 specifically includes: S311. Based on typical scenarios in the aerospace field, five core task templates are defined, millions of instruction-response pairs are generated through automated pipelines, and instances are generated by automatically filling in real data; the five core task templates are: interpretation, calculation, diagnosis, comparison, and design; S312. To improve the model's anti-interference ability, adversarial enhancement is applied to the data: by randomly replacing some domain terms and confusing the unit system, the model is forced to maintain semantic consistency under noise interference; S313, data cleaning stage, calculate the sample perplexity, filter out samples with too high perplexity, and conduct manual review through expert cross-review mechanism. Randomly select some data and have them independently labeled by multiple experts in the field. Only samples that are passed unanimously are retained to ensure the accuracy of technical details.

5. The method for constructing aerospace knowledge big model according to claim 1, characterized in that: Step S43 is: For user queries that only involve a single domain-specific entity, the user query is expanded to fill in the general attributes of the relevant entity to generate a subquery; for user queries that contain multiple domain-specific entities, the query is split, and the association dimensions between entities are extracted based on the aerospace knowledge graph, and the complex original query is rewritten into an easy-to-retrieve form to generate a structured subquery; the subqueries generated by expansion and splitting are used to obtain the documents corresponding to each subquery.

6. The method for constructing aerospace knowledge big model according to claim 1, characterized in that: Step S5 specifically includes: S51. Following the principles of "domain coverage", "difficulty level" and "dynamic evolution", a test set covering four major task types is constructed based on a high-quality aerospace knowledge database. The difficulty of the test questions is dynamically calibrated through the item response theory model, and the discrimination parameters and difficulty parameters are calculated. Questions with low discrimination and unbalanced difficulty are eliminated to ensure that the test set can effectively distinguish large models of different ability levels. S52. A composite weighting scheme is used to form an evaluation index system, and the comprehensive scores obtained include: terminology accuracy, calculation accuracy, logical coherence and response speed; The accuracy of terms is verified by knowledge graph entity linking; the calculation accuracy relies on the SymPy symbolic computing engine to parse and generate mathematical expressions in the text and verify the rationality of the values; the logical coherence uses the pre-trained Coherence-BERT model to evaluate the integrity of the causal chain between paragraphs; S53. In response to the continuous evolution of domain knowledge, a dynamic assessment platform is built to conduct adaptive difficulty testing: dynamically adjust the difficulty of questions based on the real-time performance of the big model. If the correct rate is >90% after answering multiple questions continuously, a higher-level task will be automatically pushed; if the error rate is >30%, it will be downgraded to the basic question type to prevent assessment bias; S54. Introduce a virtual-reality verification mechanism: simulate the digital twin environment in a virtual scene, requiring the large model to output solutions that conform to physical laws; deploy the model in parallel at the engineering site, compare the consistency of its suggestions with the engineers' decisions, and calculate the degree of consistency; S55. By continuously absorbing progress in the field and user feedback, a dynamic closed-loop evaluation system of "test-optimize-retest" is formed to ensure that the knowledge model maintains professionalism and practicality during rapid iteration.

Citation Information

Patent Citations

  • Large model programming review system based on combination of knowledge graph and RAG

    CN119106173A

  • Aviation standard question and answer optimization method and system based on atlas and document data

    CN119621894A