A method for constructing a safety interpretable air traffic control instruction generation large model based on retrieval enhancement generation in the aviation field
By constructing a multimodal aviation instruction corpus and fine-tuning a large language model, combined with ViDoRAG's multimodal hybrid retrieval and multi-agent collaborative generation, the problems of professional knowledge understanding and decision transparency in the aviation field were solved, and professional instructions for aviation safety regulations were generated.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP
- Filing Date
- 2026-01-21
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies lack professional knowledge in the aviation field, generate results that lack interpretability, and are difficult to achieve effective semantic alignment and fusion of multimodal data, thus failing to meet the requirements for transparent decision-making in aviation safety control.
By constructing a multimodal aviation command corpus, a low-rank adaptive method and a contrastive loss function are used to fine-tune the large language model. Combined with ViDoRAG's multimodal hybrid retrieval and multi-agent collaborative generation, semantic alignment and decision traceability of the large air traffic control command generation model are achieved.
It generates professional instructions that comply with aviation safety regulations, possesses a complete decision-making chain, and improves the semantic alignment and decision transparency of multimodal data.
Smart Images

Figure CN122133789A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for constructing a large model for generating air traffic control instructions, and more particularly to a method for constructing a large model for generating safe and interpretable air traffic control instructions in the aviation field based on retrieval-enhanced generation. Background Technology
[0002] This section provides only background information relevant to this disclosure and is not necessarily prior art.
[0003] With the rapid development of intelligent aviation systems, air traffic control, flight operations, and safety management are facing an urgent need for massive multimodal data processing and intelligent decision-making. Achieving accurate perception of aviation operational situations, risk warnings, and decision support through artificial intelligence technology has become a core direction for the industry's digital transformation. However, the aviation field is characterized by its intensive professional knowledge, stringent safety requirements, and complex decision-making chains. Traditional large language models face three core challenges in their application: First, general-purpose large language models lack a deep understanding of aviation professional knowledge, making it difficult to accurately analyze professional content such as flight parameters, airspace structure, and control procedures; second, the model-generated results have a "black box" characteristic, lacking the interpretability and logical traceability required for aviation safety; and third, traditional full-parameter fine-tuning methods suffer from high computational costs and low efficiency in injecting domain knowledge, severely restricting the practical application of models in real-world aviation scenarios.
[0004] Currently, some research has attempted to improve the domain adaptability of models through techniques such as instruction fine-tuning and knowledge graph embedding. These methods fine-tune pre-trained models by constructing domain corpora or enhance the model's semantic understanding capabilities through external knowledge bases, thus improving the model's mastery of professional terminology and basic knowledge to some extent. However, existing methods face significant limitations in complex aviation scenarios: on the one hand, it is difficult to achieve effective semantic alignment and fusion reasoning of multimodal data (including radar images, flight track data, meteorological information, etc.); on the other hand, it lacks a complete and interpretable chain from data input to decision output, failing to meet the mandatory requirements of aviation safety control for transparency in the decision-making process. In addition, traditional methods still have significant gaps in the accuracy, compliance, and practicality of the content generated when dealing with dynamic and complex scenarios such as handling aviation emergencies and multi-role collaborative decision-making.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] Purpose of the invention: The technical problem to be solved by the present invention is to address the shortcomings of the existing technology by providing a method for constructing a large model of air traffic control command generation that is safe and interpretable in the aviation field based on retrieval enhancement.
[0007] To address the aforementioned technical problems, this invention discloses a method for constructing a large-scale, interpretable air traffic control instruction generation model based on retrieval enhancement, comprising the following steps:
[0008] Step 1: Collect multimodal aviation data in the air traffic control field and construct a multimodal aviation command corpus;
[0009] Step 2: Based on the multimodal aviation command corpus constructed in Step 1, a low-rank adaptive method and a contrastive loss function are used to fine-tune the Transformer encoding layer of the large language model. The latent space encoding of the large language model is used to encode images and text in the aviation domain to achieve semantic alignment and retrieval.
[0010] Step 3: Combining ViDoRAG's multimodal hybrid retrieval and multi-agent collaborative generation capabilities, the heterogeneous data and business logic in the air traffic control command generation model retrieval process are projected into a fully traceable decision space, thus completing the construction of a safe and interpretable air traffic control command generation model based on retrieval enhancement.
[0011] Furthermore, the construction of the multimodal aviation command corpus described in step 1 includes:
[0012] Step 1-1: Convert the multimodal aviation data stream Mapped to a dynamically updated set of basic morphemes The mapping rules are as follows:
[0013]
[0014] in, Represents multimodal data stream Chinese analytic function The parsed candidate basic morphemes, if the currently parsed morpheme With historical elements The change exceeds the threshold Then the current morpheme will be added to the basic morpheme set. Otherwise, retain the old morphemes. ; Indicates the subscripts of different morpheme sets. Indicates a specific moment to distinguish data streams;
[0015] Steps 1-2 involve semantic fusion based on logical operations, which integrates the discrete set of basic morphemes. Mapped to a set of complex semantic scene morphemes conforming to aviation domain rules The mapping rules are as follows:
[0016]
[0017] in, The basic set of morphemes for input The basic morphemes in Chinese, This indicates that the morpheme combination is validated based on the knowledge rule base in the aviation field. Indicates logical conflict detection; Represents logical combination of morphemes based on the AND relation; Indicates taking the complement;
[0018] Steps 1-3 involve using a pre-defined prompting engineering strategy to combine the set of morphemes from complex semantic scenarios. Mapped to natural language descriptions that conform to air traffic control expression conventions The mapping process is as follows:
[0019]
[0020] in, This indicates an input suggestion built based on a domain-adaptive suggestion template. This represents the text generation process of a large language model. This represents an iterative optimization function based on aviation professional dictionaries and expression standards;
[0021] Steps 1-4, based on natural language description Combined with a dialogue scenario database in the aviation field Generate a multimodal aviation command corpus , means as follows:
[0022]
[0023] in, Represents a dialogue scenario library, For scenario-based Dialogue simulation function, For random disturbance factors, It is to construct the index of the specified perturbation factor. It is the size of the set used to construct the perturbation factors.
[0024] Furthermore, in step 1-1, the analytical function Based on a pre-trained large language model, it is represented as follows:
[0025]
[0026] in, Indicates based on pre-trained parameter set The Great Prophecy Model Data stream processing functions typically use language models to summarize the data stream to filter out noise. Indicates that for the first Specific prompt words designed based on basic morphemes, This indicates a text token concatenation operation.
[0027] Furthermore, in steps 1-3, the iterative optimization function is defined as follows:
[0028]
[0029]
[0030] in, For compliance verification functions, This indicates a feedback generation function. The maximum number of iterations is . .
[0031] Furthermore, the fine-tuning of the Transformer encoding layer of the large language model described in step 2 involves using low-rank adaptive tuning combined with a contrastive loss function to fine-tune the Transformer encoding layer of the large language model PaliGemma.
[0032] Furthermore, the fine-tuning of the Transformer encoding layer of the large language model PaliGemma using low-rank adaptive combined with a contrastive loss function specifically includes:
[0033] Keep the SigLIP visual encoder parameters frozen in the large language model PaliGemma, and inject a trainable low-rank decomposition matrix into the self-attention mechanism module of the Transformer layer. and Set the rank and scaling factor of the low-rank adaptive model, while keeping the projection layer in a fully trainable state, and train it to obtain the empty control command to generate a large model.
[0034] During training, the loss function is constructed with the goal of maximizing the later interaction similarity of positive sample pairs. .
[0035] Furthermore, the construction of the loss function The details are as follows:
[0036]
[0037] in, For similarity function, For the set of negative samples within the batch, Indicates the training batch size. Refers to the power function. For querying text, To retrieve the collection of documents corresponding to the text, This represents a collection of documents that are unrelated to the query text. Indicates the subscript of the training sample. This indicates the index of irrelevant documents in the training samples.
[0038] Furthermore, the similarity function mentioned in step 2 is calculated as follows:
[0039] Step 2-1: Input a document page image, standardize it according to a fixed resolution, and divide it into several non-overlapping fixed-size image blocks using a sliding window; input each image block into a pre-trained SigLIP visual encoder to obtain the visual encoding result token. , means as follows:
[0040]
[0041] in, For the dimensions of visual features, This represents the encoding results of each image block. This indicates the number of image blocks that the current document has been divided into;
[0042] Step 2-2, visual encoding result token A mixed sequence is formed with a fixed text prefix token and fed into the Transformer encoding layer of the large language model PaliGemma. The output is a cross-modal joint representation of the visual encoding result token. , means as follows:
[0043]
[0044] Where h is the dimension of the hidden layer of the Transformer encoding layer; This represents the joint representation of the various image patches.
[0045] Steps 2-3: For text queries entered in plain text format The query is directly encoded using the language Transformer layer of the large language model PaliGemma, and the resulting text token sequence is output after encoding. , means as follows:
[0046]
[0047] in, To query the total number of tokens, This represents the combined representation of the query token;
[0048] Steps 2-4: Add a low-dimensional projection layer after the output of the language Transformer layer, assuming the projection matrix is... The embedding mapping process is as follows:
[0049]
[0050]
[0051] in, This is the embedded representation of the projected query. Embedding representations for projected documents;
[0052] Finally, the query embedding matrix is obtained. With document embedding matrix They are represented as follows:
[0053]
[0054]
[0055] Steps 2-5: Calculate the query using the post-interaction mechanism. and documents similarity The similarity function is:
[0056]
[0057] in, This represents the inner product operation, used to calculate the semantic similarity between two low-dimensional embedding vectors.
[0058] Furthermore, step 3, which involves projecting heterogeneous data and business logic from the air traffic control instruction generation model retrieval process into a fully traceable decision space, includes:
[0059] Step 3-1: During the retrieval process, a source tracing tag is added to each candidate knowledge unit. This tag contains original source information, processing information, and credibility information. When the retrieval results are returned, the source tracing tag is output synchronously. The tag information is directly called during the generation stage to achieve full-chain traceability. The knowledge unit mentioned is the document embedding matrix constructed in Step 2-4. Embedded representation vectors in and corresponding block content ;
[0060] Step 3-2: Embed the multimodal knowledge units, i.e., documents, into the matrix. Constructing a knowledge graph for the aviation field;
[0061] Step 3-3: Based on the aviation knowledge graph, intelligently resolve knowledge conflicts;
[0062] Steps 3-4 involve using a pre-defined recall strategy to identify and resolve ambiguities in user questions, extracting core terms as search filtering conditions, and scoring the initial search results to obtain the optimal result.
[0063] Steps 3-5: Based on the ViDoRAG multi-agent architecture, generate air traffic control instructions based on the optimal results.
[0064] Furthermore, the multimodal aeronautical data stream described in step 1-1 This includes flight parameters, airport status, meteorological information, airspace structure, aviation safety management, and aviation regulations text data collected from the air traffic control domain.
[0065] Beneficial effects:
[0066] This invention can effectively integrate multimodal aviation data and generate professional instructions that comply with aviation safety regulations and have a complete decision-making chain through an interpretability-enhanced architecture and dynamic compliance verification mechanism. Attached Figure Description
[0067] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.
[0068] Figure 1 This is a flowchart of the method of the present invention.
[0069] Figure 2 This is a schematic diagram of the large model architecture provided by the present invention. Detailed Implementation
[0070] This invention proposes a method for constructing a large-scale model for generating safe and interpretable air traffic control instructions in the aviation field based on retrieval enhancement. This large-scale model is used to fuse multimodal aviation data and generate professional instructions that comply with aviation safety regulations and have a complete decision-making chain through an interpretability enhancement architecture and dynamic compliance verification mechanism.
[0071] This invention provides the following technical solution:
[0072] A method for constructing a large-scale model of air traffic control command generation that is safety-interpretable in the aviation field based on retrieval-enhanced generation, such as... Figure 1 As shown, the method includes the following steps:
[0073] Step 1: Given multimodal aviation data, key technologies such as structured language prompt templates, multi-turn dialogue enhancement, and multimodal reasoning are used to unify the multi-source heterogeneous data into a large-scale, high-quality multimodal aviation command corpus.
[0074] Step 1-1: Given a multimodal aeronautical data stream that integrates flight parameters, airport status, meteorological information, airspace structure, aviation safety management, and aviation regulations. Map it to a dynamically updated set of basic morphemes The mapping rules are as follows:
[0075]
[0076]
[0077] Among them, analytic function Based on pre-trained large language models Indicates based on pre-trained parameter set For this large oracle model, this paper selects Qwen2.5-VL-7B as the basic large language model (Reference: Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., & Lin, J. (2025). Qwen2.5-VL Technical Report. ArXiv, abs / 2502.13923.). Indicates that for the first Specific prompt words designed based on basic morphemes, This indicates a text token concatenation operation. Represents multimodal data stream Chinese analytic function The parsed candidate basic morphemes, if the currently parsed morpheme With historical elements The change exceeds the threshold If the current morpheme is found, it is added to the set; otherwise, the old morpheme is retained. .
[0078] Steps 1-2: The semantic fusion engine based on logical operations maps discrete basic morphemes into a set of composite semantic scene morphemes that conform to aviation domain rules. The generation rules are as follows:
[0079]
[0080] in This indicates a logical combination of morphemes based on the "AND" relationship. This represents the set of generated compound semantic scene morphemes. The basic morphemes for input, The rationality of morpheme combinations was verified using a large language model based on a knowledge rule base in the aviation field. This indicates a logical conflict between morphemes such as "speed = 0 km / h" and "altitude continues to increase". This indicates taking the complement set.
[0081] Steps 1-3: Given structured morpheme data processed by the semantic fusion engine It maps the pre-set large language model prompts into natural language descriptions that conform to the expression habits of air traffic control. The generation and optimization process is as follows:
[0082]
[0083] in This indicates an input suggestion built based on a manually constructed domain-adaptive suggestion template. This represents the text generation process of a large language model. This represents an iterative optimization function based on aviation professional dictionaries and expression standards. It improves the quality of substandard samples through multiple rounds of refinement and example correction. Specifically, the iterative optimization function is defined as follows:
[0084]
[0085]
[0086] in, This compliance verification function, by introducing a language model from an aviation rule base, can perform natural language processing. Automatic detection is performed. If there are errors in terminology, non-standard abbreviations, or sentence structure violations, a verification failure status and a set of errors are returned. This represents the feedback generation function, implemented using a pre-trained language model. When validation fails, this function constructs correction instructions based on the error set, guiding the language model to modify the text. Iterative optimization is performed using the above method until compliance validation passes or the maximum number of iterations is exceeded. .
[0087] Steps 1-4: Based on natural language descriptions that conform to air traffic control expression habits Dialogue Scenario Library with the Aviation Industry A multimodal aviation command corpus was generated using multi-turn dialogue enhancement technology. The process is described as follows:
[0088]
[0089] in It represents a dialogue scenario library containing over a hundred typical scenarios across eight categories, including routine command and emergency response. This is a dialogue simulation function that is context-based. Role configuration and environmental information drive the large language model to generate multi-scene corpora through multi-turn dialogues. This is a random disturbance factor. (In a more technical description) Based on semantics, contextualized dialogue simulation and controllable mutation mechanism are used to effectively expand the data diversity of the aviation multimodal command corpus.
[0090] Step 2: Building upon Step 1, a high-efficiency document retrieval model is constructed using a visual language model. This model achieves cross-modal semantic alignment and efficient retrieval through end-to-end visual language modeling and multi-vector matching mechanisms.
[0091] Step 2-1: Input the supporting document page image used as the answer to the query. Standardize it to a fixed resolution and divide it into several non-overlapping fixed-size image patches using a sliding window. Each image patch is input to a pre-trained SigLIP visual encoder (Reference: Zhai, X., Mustafa, B., Kolesnikov, A., & Beyer, L.(2023). Sigmoid loss for language image pre-training. In Proceedings of the IEEE / CVF international conference on computer vision (pp. 11975-11986).). The visual encoding result token is represented as:
[0092]
[0093] in, For the dimensions of visual features, This is the encoded token sequence. For text sequence numbers, This is the preset number of blocks.
[0094] Step 2-2: The visual encoding result token is concatenated with the fixed task cue prefix token to form a mixed sequence, which is then input into the PaliGemma language Transformer encoding layer (reference: Beyer, L., Steiner, A., Pinto, AS, Kolesnikov, A., Wang, X., Salz, D., ... & Zhai, X. (2024). Paligemma: Aversatile 3b vlm for transfer. arXiv preprint arXiv:2407.07726.). The model establishes cross-modal dependencies through a multi-head self-attention mechanism and outputs the cross-modal joint representation corresponding to the visual token.
[0095]
[0096] in, For the hidden layer dimension of the language model, Indicates joint visual embedding. This represents the embedding of the corresponding block, which integrates the structural information of the visual image block with the semantic understanding capability of the language model.
[0097] Steps 2-3: For text query q, the model inputs it as plain text, without needing a visual encoding process. It directly encodes the query through PaliGemma's Language Transformer layer (reference: Beyer, L., Steiner, A., Pinto, AS, Kolesnikov, A., Wang, X., Salz, D., ... & Zhai, X. (2024). Paligemma:A versatile 3b vlm for transfer. arXiv preprint arXiv:2407.07726.). The output query encoding yields a text token sequence representation.
[0098]
[0099] in, To query the total number of tokens (including special extended tokens), This indicates a joint query for Embedding. This indicates the embedding of the corresponding word segment.
[0100] Steps 2-4, to unify the alignment of visual and textual features in the retrieval space and reduce storage and computation costs, add a low-dimensional projection layer after the language model output to map the high-dimensional semantic representation to a compact shared retrieval space. Let the projection matrix be... The embedding mapping process is as follows:
[0101]
[0102]
[0103] Finally, the query embedding matrix is obtained. With document embedding matrix .
[0104] Steps 2-5 involve using a post-interaction mechanism to calculate the similarity between the two pairs. The similarity function is defined as follows:
[0105]
[0106] in This represents the inner product operation, used to calculate the semantic similarity between two low-dimensional embedding vectors. This function is used in the retrieval process to retrieve the set of documents most relevant to the query. Specifically, for each query... The system will use a similarity function to calculate the similarity between the query and each document. The similarity is used to retrieve the most relevant set of documents as the retrieval result, which will directly serve as an external knowledge supplement to the Big Prophecy model.
[0107] Steps 2-6: In order to adapt to the professional context of the aviation field, based on the multimodal aviation command corpus constructed in step 1, the Transformer encoding layer of the large language model PaliGemma is fine-tuned using low-rank adaptive (LoRA) technology combined with a contrastive loss function.
[0108] Specifically, the model keeps the SigLIP visual encoder parameters frozen and injects a trainable low-rank decomposition matrix into the self-attention mechanism module of the Transformer layer of the PaliGemma language model. and Set the rank of LoRA scaling factor Simultaneously maintain the projection layer This is a fully trainable state, adapted to the feature mapping of the retrieval space. During model training, the objective is to maximize the later interaction similarity of positive sample pairs, and the loss function is... The following structure is constructed to establish the intrinsic semantic relationship between the input image patch and the text query:
[0109]
[0110] in, The similarity function described in steps 2-5 For the set of negative samples within the batch, Indicates the training batch size. Refers to a power function.
[0111] Step 3: Combining ViDoRAG's multimodal hybrid retrieval and multi-agent collaborative generation capabilities, collaborative processing and semantic fusion are performed on aviation multimodal data. Through a "three-layer, four-dimensional" interpretability enhancement architecture, heterogeneous data and business logic are projected into a fully traceable decision space.
[0112] Step 3-1: The source tracing mechanism adds a "source tracing tag" to each candidate knowledge unit during the retrieval process. The tag includes original source information, processing information, and credibility information. The knowledge unit is the document embedding matrix constructed in Step 2-4. The system includes embedding vectors and corresponding block content. When retrieval results are returned, traceability tags are output synchronously. Tag information can be directly accessed during the generation phase, enabling full-chain traceability from "generated content to retrieval results to original knowledge." Simultaneously, blockchain technology is used to store the traceability tags, ensuring the tag information is tamper-proof and meeting the auditing requirements of the aviation industry.
[0113] Step 3-2, Visual Retrieval of Multi-hop Reasoning Paths, constructs a "Knowledge Graph in the Aviation Field" from multimodal knowledge units. The graph contains 8 types of core entities (flights, aircraft types, personnel, facilities, etc.) and 12 types of relationships (process relationships, component relationships, legal basis, etc.).
[0114] Step 3-3: Intelligent Knowledge Conflict Resolution. For potential knowledge conflicts within the retrieved knowledge units, firstly, rule-based resolution is performed, establishing rules prioritizing authoritative sources, recentest time, and scenario suitability based on the characteristics of the aviation field. Secondly, model-based resolution is performed, using a pre-trained conflict resolution model for judgment. The model is based on BERT, fine-tuned with conflict cases in the aviation field. Inputting conflicting knowledge units and problem scenarios, it outputs the optimal knowledge unit and selection criteria. After resolution is complete, the system automatically generates a "Conflict Resolution Report."
[0115] Steps 3-4 employ a three-level recall strategy: "terminology filtering - hybrid retrieval - re-ranking." An aviation terminology dictionary is used to identify and resolve ambiguities in user questions, extracting core terms as filtering conditions. Keyword retrieval and semantic retrieval are integrated; keyword retrieval ensures accurate matching of professional terms, while semantic retrieval captures implicit semantic relationships. A Cross-Encoder model is used to score the initial search results a second time, with scoring dimensions including terminology matching degree, semantic relevance, credibility, and scenario adaptability, ultimately returning the Top-N optimal results.
[0116] Steps 3-5, based on the ViDoRAG multi-agent architecture, enhance the functionality of the three agents by adding an explanation generation module and a process visualization module, forming an interpretable generation system of "three agents + two modules". First, the explorer agent is enhanced to generate a "basis summary," recording key decision points in the screening process as part of the process explanation. Second, the user or reviewer agent is enhanced to record the reasoning process using "if-then" rules. For complex reasoning, flowcharts are used to record logical relationships to ensure the reproducibility of the reasoning process. Finally, the responder agent is enhanced to adopt a four-stage generation structure of "conclusion + basis + reasoning + confidence". The specific architecture of the large model proposed in this invention is as follows: Figure 2 As shown.
[0117] Example:
[0118] This invention uses the decision-making process during aircraft climb as an example to describe in detail the execution flow of corpus construction, model training, and retrieval methods. Specifically, at time t, during the final stage of takeoff roll, aircraft B-7677 is at an altitude of 2500 meters and a speed of 110 km / h. 800 meters directly ahead of the runway extension, a vertical man-made object approximately 60 meters high is detected. According to the pilot's operating manual, the optimal climb rate of 180° should be used to climb as quickly as possible. In the corpus construction stage, according to step 1-1, morphemes are first extracted and updated. Since the area ahead of the runway is considered "clear" at time t-1, the morpheme at the current time is modified to "runway extension obstacle". Finally, the basic morphemes at time t can be described as: "Aircraft B-7677", "Speed 110°", "Altitude 2500 meters", "Runway extension obstacle", "Optimal climb rate", "Speed correction 120°".
[0119] Steps 1-2 construct compliant morpheme combinations based on corpus materials. For the morpheme combination: “Aircraft B-660”, “Speed 110”, “Altitude 2500”, “Runway clearance”, “Optimal climb rate speed”, it is considered compliant after inspection. For the morpheme combination: “Aircraft B-7677”, “Speed 110”, “Altitude 2500”, “Runway extension obstacle”, “Optimal climb rate speed”, it is considered non-compliant after inspection.
[0120] Steps 1-3 iteratively generate natural language mappings based on compliant morpheme combinations. For the morpheme combination: "Aircraft B-7677", "Speed 110", "Altitude 2500", "Runway extension obstacle", "Optimal climb angular velocity", "Speed correction 120", the compliant natural language generated iteratively is: "B-7677, there is an obstacle at the departure end, maintain the optimal climb angular velocity, speed 120, until the obstacle is cleared."
[0121] Steps 1-4 generate simulated dialogue scenarios based on the aforementioned generated natural language text. For the compliant natural language text in steps 1-3, a scenario dialogue is generated using the controller-pilot dialogue template: "Pilot: B-7677, current altitude 2500, speed 110, obstacle at departure end. Controller: Received, maintain optimal climb angle, speed 120, avoid obstacle, B-7677."
[0122] The large dataset constructed using the above method will be used for model training in step 2 to implement a dedicated encoder adapted for the aviation field. Then, in the indexing phase, the encoder first discretizes each page of the visual document in the operation manual into a token sequence, which is then encoded into a representation vector and stored. In the retrieval phase, the encoder encodes the pilot's request into a representation vector and retrieves the most relevant documents as external references. Subsequently, the query and relevant documents are input into a large language model to obtain the appropriate decision.
[0123] The present invention provides a method for constructing a safe and interpretable large model for the aviation domain based on retrieval enhancement generation. The model is trained on three datasets with rich corpora for cross-document retrieval tasks in the aviation domain: arXivQA (paper graph dataset), DocVQA (academic scanned document dataset), and TAT-DQA (financial report dataset). Experiments were conducted on general scenarios and aviation domain datasets, respectively.
[0124] The experiment uses nDCG@5 as the core accuracy indicator, supplemented by Recall@1 to verify precise positioning capabilities. Both are classic evaluation indicators in the field of information retrieval and are highly suitable for the actual needs of page-level document retrieval. This experiment adopts a three-level scoring system: 2 for complete relevance, 1 for partial relevance, and 0 for irrelevantness, with a discount factor. This indicates the degree of attention a user pays to preceding results during the search process; the further back in the search results are located, the more significantly the attractiveness of the results diminishes for the user.
[0125] Table 1. Results of the comparative experiment on document retrieval tasks
[0126]
[0127] As shown in Table 1, the model of this invention significantly outperforms traditional BM25 and its enhancement methods in terms of retrieval accuracy and F1 score on multiple test sets. Compared with BM25, which relies on plain text matching, and BM25+OCR, which is based on OCR post-processing, the end-to-end multimodal retrieval model proposed in this invention can directly understand the visual semantics (such as chart structure and layout) in document images, effectively overcoming the information loss caused by OCR recognition errors and layout parsing loss in traditional methods. The results show that this invention achieves more accurate cross-modal semantic alignment while preserving complete visual information, significantly improving the robustness and accuracy of complex document retrieval.
[0128] To verify the effectiveness of the method in step 3, this experiment was conducted on several basic large language models, and the performance was further compared with that of traditional RAG. Specifically, this experiment used the general large language models Qwen2.5-1.5B-Instruct and Llama3.2-1B-Instruct as the base large models, and compared the proposed method, the method without RAG, and the traditional RAG method. The results are shown in Table 2:
[0129] Table 2. Results of the Model Inference Comparison Experiment
[0130]
[0131] in, This describes the lexical overlap between the model output and the standard answer. It is a soft matching metric based on contextual semantic embedding, through The maximum cosine similarity between the prediction and the reference is used to measure the semantic closeness. These two metrics together reflect the consistency between the model's generation and the standard answer, thus assessing the interpretability and accuracy of the model-generated content. Compared to direct model inference or traditional semantic similarity-based RAG methods, the traceable retrieval scheme proposed in this invention can effectively alleviate the illusion of large models and improve model performance in the aviation field.
[0132] In its specific implementation, this application provides a computer storage medium and a corresponding data processing unit. The computer storage medium is capable of storing a computer program, which, when executed by the data processing unit, can run the invention's content regarding a method for constructing a large model of air traffic control instructions for aviation safety based on retrieval enhancement, as well as some or all of the steps in various embodiments. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0133] Those skilled in the art will clearly understand that the technical solutions in the embodiments of the present invention can be implemented using computer programs and their corresponding general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of computer programs, i.e., software products. These computer program software products can be stored in a storage medium and include several instructions to cause a device containing a data processing unit (which may be a personal computer, server, microcontroller, MCU, or network device, etc.) to execute the methods described in various embodiments or certain parts of the embodiments of the present invention.
[0134] This invention provides a method for constructing a large-scale model of air traffic control instructions that is safe and interpretable in the aviation field based on retrieval-enhanced generation. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment of the invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.
Claims
1. A method for constructing a large-scale model of air traffic control command generation that is safety-interpretable in the aviation field based on retrieval-enhanced generation, characterized in that, Includes the following steps: Step 1: Collect multimodal aviation data in the air traffic control field and construct a multimodal aviation command corpus; Step 2: Based on the multimodal aviation command corpus constructed in Step 1, a low-rank adaptive method and a contrastive loss function are used to fine-tune the Transformer encoding layer of the large language model. The latent space encoding of the large language model is used to encode images and text in the aviation domain to achieve semantic alignment and retrieval. Step 3: Combining ViDoRAG's multimodal hybrid retrieval and multi-agent collaborative generation capabilities, the heterogeneous data and business logic in the air traffic control command generation model retrieval process are projected into a fully traceable decision space, thus completing the construction of a safe and interpretable air traffic control command generation model based on retrieval enhancement.
2. The method for constructing a large-scale model of air traffic control command generation that is safety-interpretable in the aviation field based on retrieval enhancement generation as described in claim 1, characterized in that, The construction of the multimodal aviation command corpus mentioned in step 1 includes: Step 1-1: Convert the multimodal aviation data stream Mapped to a dynamically updated set of basic morphemes The mapping rules are as follows: ; in, Represents multimodal data stream Chinese analytic function The parsed candidate basic morphemes, if the currently parsed morpheme With historical elements The change exceeds the threshold Then the current morpheme will be added to the basic morpheme set. Otherwise, retain the old morphemes. ; Indicates the subscripts of different morpheme sets. Indicates a specific moment to distinguish data streams; Steps 1-2 involve semantic fusion based on logical operations, which integrates the discrete set of basic morphemes. Mapped to a set of complex semantic scene morphemes conforming to aviation domain rules The mapping rules are as follows: ; in, The basic set of morphemes for input The basic morphemes in Chinese, This indicates that the morpheme combination is validated based on the knowledge rule base in the aviation field. Indicates logical conflict detection; Represents logical combination of morphemes based on the AND relation; Indicates taking the complement; Steps 1-3 involve using a pre-defined prompting engineering strategy to combine the set of morphemes from complex semantic scenarios. Mapped to natural language descriptions that conform to air traffic control expression conventions The mapping process is as follows: ; in, This indicates an input suggestion built based on a domain-adaptive suggestion template. This represents the text generation process of a large language model. This represents an iterative optimization function based on aviation professional dictionaries and expression standards; Steps 1-4, based on natural language description Combined with a dialogue scenario database in the aviation field Generate a multimodal aviation command corpus , means as follows: ; in, Represents a dialogue scenario library, For scenario-based Dialogue simulation function, For random disturbance factors, It is to construct the index of the specified perturbation factor. It is the size of the set used to construct the perturbation factors.
3. The method for constructing a large-scale model of air traffic control command generation that is safety-interpretable in the aviation field based on retrieval enhancement generation according to claim 2, characterized in that, In step 1-1, the analytical function Based on a pre-trained large language model, it is represented as follows: ; in, Indicates based on pre-trained parameter set The Great Prophecy Model Data stream processing functions typically use language models to summarize the data stream to filter out noise. Indicates that for the first Specific prompt words designed based on basic morphemes, This indicates a text token concatenation operation.
4. The method for constructing a large-scale model for generating air traffic control instructions in the aviation field based on retrieval enhancement generation, as described in claim 3, is characterized in that... In steps 1-3, the iterative optimization function is defined as follows: ; in, For compliance verification functions, This indicates a feedback generation function. The maximum number of iterations is . .
5. The method for constructing a large-scale model for generating air traffic control instructions in the aviation field based on retrieval enhancement generation, as described in claim 4, is characterized in that... The fine-tuning of the Transformer encoding layer of the large language model described in step 2 involves using low-rank adaptive tuning combined with a contrastive loss function to fine-tune the Transformer encoding layer of the large language model PaliGemma.
6. The method for constructing a large-scale model for generating air traffic control instructions in the aviation field based on retrieval enhancement generation according to claim 5, characterized in that, The fine-tuning of the Transformer encoding layer of the large language model PaliGemma using low-rank adaptive combined with a contrastive loss function specifically includes: Keep the SigLIP visual encoder parameters frozen in the large language model PaliGemma, and inject a trainable low-rank decomposition matrix into the self-attention mechanism module of the Transformer layer. and Set the rank and scaling factor of the low-rank adaptive model, while keeping the projection layer in a fully trainable state, and train it to obtain the empty control command to generate a large model. During training, the loss function is constructed with the goal of maximizing the later interaction similarity of positive sample pairs. .
7. The method for constructing a large-scale model of air traffic control command generation that is safety-interpretable in the aviation field based on retrieval enhancement generation according to claim 6, characterized in that, The construction loss function The details are as follows: ; in, For similarity function, For the set of negative samples within the batch, Indicates the training batch size. Refers to the power function. For querying text, To retrieve the collection of documents corresponding to the text, This represents a collection of documents that are unrelated to the query text. Indicates the subscript of the training sample. Indicates the index of irrelevant documents in the training samples.
8. The method for constructing a large-scale model of air traffic control command generation that is safety-interpretable in the aviation field based on retrieval enhancement generation according to claim 7, characterized in that, The similarity function mentioned in step 2 is calculated as follows: Step 2-1: Input a document page image, standardize it according to a fixed resolution, and divide it into several non-overlapping fixed-size image blocks using a sliding window; input each image block into a pre-trained SigLIP visual encoder to obtain the visual encoding result token. , means as follows: ; in, For the dimensions of visual features, This represents the encoding results of each image block. This indicates the number of image blocks that the current document has been divided into; Step 2-2, visual encoding result token A mixed sequence is formed with a fixed text prefix token and fed into the Transformer encoding layer of the large language model PaliGemma. The output is a cross-modal joint representation of the visual encoding result token. , means as follows: ; Where h is the dimension of the hidden layer of the Transformer encoding layer; Represents the joint representation of the various image patches; Steps 2-3: For text queries entered in plain text format The query is directly encoded using the language Transformer layer of the large language model PaliGemma, and the resulting text token sequence is output after encoding. , means as follows: ; in, To query the total number of tokens, This represents the combined representation of the query token; Steps 2-4: Add a low-dimensional projection layer after the output of the language Transformer layer, assuming the projection matrix is... The embedding mapping process is as follows: ; ; in, This is the embedded representation of the projected query. Embedding representations for projected documents; Finally, the query embedding matrix is obtained. With document embedding matrix They are represented as follows: ; ; Steps 2-5: Calculate the query using the post-interaction mechanism. and documents similarity The similarity function is: ; in, This represents the inner product operation, used to calculate the semantic similarity between two low-dimensional embedding vectors.
9. The method for constructing a large-scale model of air traffic control command generation that is interpretable for aviation safety based on retrieval enhancement, as described in claim 8, is characterized in that... Step 3, which involves projecting heterogeneous data and business logic from the air traffic control instruction generation model retrieval process into a fully traceable decision space, includes: Step 3-1: During the retrieval process, a source tracing tag is added to each candidate knowledge unit. This tag includes original source information, processing information, and credibility information. When the retrieval results are returned, the source tracing tag is output synchronously. The tag information is directly called during the generation stage to achieve full-chain traceability. The knowledge unit referred to here is the document embedding matrix constructed in steps 2-4. Embedded representation vectors in and corresponding block content ; Step 3-2: Embed the multimodal knowledge units, i.e., documents, into the matrix. Constructing a knowledge graph for the aviation field; Step 3-3: Based on the aviation knowledge graph, intelligently resolve knowledge conflicts; Steps 3-4 involve using a pre-defined recall strategy to identify and resolve ambiguities in user questions, extracting core terms as search filtering conditions, and scoring the initial search results to obtain the optimal result. Steps 3-5: Based on the ViDoRAG multi-agent architecture, generate air traffic control instructions based on the optimal results.
10. The method for constructing a large-scale model of air traffic control command generation that is safety-interpretable in the aviation field based on retrieval enhancement generation according to claim 9, characterized in that, The multimodal aeronautical data stream described in step 1-1 , This includes flight parameters, airport status, meteorological information, airspace structure, aviation safety management, and aviation regulations data collected from the air traffic control domain.