Method for Generating Instruction Set of Large Model in Aviation Field with Multi-Dimensional Expansion
By adopting a multi-dimensional expansion method in the aviation field big model instruction set generation technology, using aviation scientific research literature and big models to generate instruction sets, and selecting instruction statements through multi-level optimization and diversity scores, the problems of low efficiency and difficult to control in the existing technology are solved, and efficient and accurate instruction generation is achieved.
Patent Information
- Application Number
- CN202411637124.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-11-15
AI Technical Summary
The existing large-model instruction set generation technology in the aviation field is difficult to meet the needs of complex tasks, the template conversion method is inefficient, and the automatic generation method based on the large-model is difficult to ensure the quality of the instruction set, and it is necessary to rely on the filtering method to obtain high-quality instruction sets.
The multi-dimensional expansion method of aviation field big model instruction set generation method is adopted. The input statement and output statement of instructions generated based on aviation research literature is used to generate multi-dimensionally expanded instruction sets, and the optimal instruction statement is selected through multi-level optimization and diversity scores. Finally, a high-quality instruction set is generated through multi-model cross-validation and dynamic stop conditions.
It realizes efficient, accurate and comprehensive instruction generation, can quickly respond to diverse task needs, improve the practicality and adaptability of instruction data, and ensures the high quality and accuracy of instruction sets.
Smart Images

Figure CN119537546B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of aviation instruction set generation, and specifically relates to a method for generating an instruction set for a large model in the aviation field with multi-dimensional expansion. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, the application boom of large models has swept through all walks of life and also shown great potential in the aviation field. Fine-tuning training of large models in the aviation field has become an important means to improve the intelligent level. By performing targeted fine-tuning on specific tasks or data sets, large models can better adapt to complex aviation application scenarios and optimize functions such as decision support, fault diagnosis, and flight control. In this process, the construction of high-quality instruction sets is crucial for the effect of fine-tuning training.
[0003] In the application scenario of aviation instruction set generation, users hope to generate corresponding high-quality instruction sets by inputting specific task requirements. This process not only involves natural language understanding but also needs to consider the specific professional terms and knowledge systems in the aviation field to ensure the accuracy and practicality of the generated instruction sets. Each instance in a typical instruction set consists of three elements: an instruction statement, an optional input that provides supplementary information for the context, and an expected output based on the instruction and the input. There are two methods for constructing instruction sets: one is to construct them manually from annotated natural language data sets. In this method, text label pairs can be converted into instruction sets by using templates and other means. The other is to use large models to automatically generate instruction sets. The method based on template conversion relies on annotated domain data. The advantage of this method is that it is simple and fast to implement, but the task sources are fixed and it is difficult to adapt to complex task requirements, while the method of automatically generating by large models has better adaptability.
[0004] In summary, for the aviation field, the current large model instruction set generation technology faces the following challenges:
[0005] (1) When dealing with complex or irregular tasks, the method based on template conversion often seems inadequate and cannot meet the personalized needs of users;
[0006] (2) The method of automatically generating by large models is difficult to guarantee the quality of instruction set generation and needs to rely on instruction set screening methods to obtain high-quality instruction sets.
[0007] (3) The aviation field has strong professionalism and involves a wide range of knowledge. The generated instruction sets must comprehensively cover the details of related fields. Summary of the Invention
[0008] To address the deficiencies of the above-mentioned existing technologies, the objective of the present invention is to provide a method for generating an instruction set for a large model in the aviation field with multi-dimensional expansion, aiming to achieve efficient, accurate, and comprehensive instruction generation to meet the ever-changing aviation application requirements.
[0009] Specifically, the present invention provides a method for generating an instruction set for a large model in the aviation field with multi-dimensional expansion, which includes the following steps:
[0010] S1. Generate input statements and output statements of instructions based on aviation research literature;
[0011] S2. Generate an instruction set for the aviation field with multi-dimensional expansion, specifically including the following sub-steps:
[0012] S21. Generate input statements and output statements of instructions for each task respectively to obtain multiple few-shot instruction input and output pair sets;
[0013] S22. Generate instruction statements based on multi-level optimization, specifically including the following sub-steps:
[0014] S221. Mix the above multiple few-shot instruction input and output pair sets. For one of the few-shot instruction input and output pair sets where the number of input and output pairs in the set is γ, use the large model to generate the missing instruction statement ρ based on the input statement and output statement, so that the instruction statement ρ can generate the output that i is most matched with the expected output Y i under the condition of the given input X;
[0015] S222. Minimize the semantic deviation of the instruction statement ρ;
[0016] S223. Optimize the instruction statement ρ by introducing a domain relevance constraint to optimize the objective function;
[0017] S224. Determine the optimal instruction statement and the sub-optimal instruction statement:
[0018] Execute the above steps. For each few-shot instruction input and output set, sort the candidate instruction statements in descending order. After the descending order sorting, take the instruction statement ranked first as the optimal instruction statement, that is, the one that satisfies the following formula
[0019]
[0020] After that, introduce a diversity score to select the sub-optimal instruction statement
[0021]
[0022] Among them, S(ρ,X i ,Y i ) is the similarity of the instruction statements, is the diversity of instruction statements, and λ is a tuning parameter used to adjust the weight between similarity and diversity;
[0023] S225, after obtaining the optimal instruction statement and the suboptimal instruction statement, for each set of few-sample instruction input and output pairs, randomly select the optimal instruction statement or the suboptimal instruction statement as the instruction statement; for all sets of few-sample instruction input and output pairs, continuously repeat the above steps until each instruction contains the complete three parts of {instruction statement, input statement, output statement};
[0024] S226, randomly extract an instruction from each task as a seed instruction {ρ 0 ,X 0 ,Y 0} and set the large model to be used as a prompt for the extended command;
[0025] S23, multi-model instruction expansion and cross-validation, specifically includes the following sub-steps:
[0026] S231, expand new instructions: use two large models to expand the original task instructions in depth and breadth, and generate new instructions {ρ new1 ,X new1 ,Y new1} and {ρ new2 ,X new2 ,Y new2}, calculate the similarity between each generated instruction statement and the seed instruction statement, and based on the similarity value, select multiple instructions with low similarity as backup instructions. new ,ρ 0 ) is calculated as:
[0027]
[0028] Where V ρnew is the vector representation of the newly generated instruction, V ρ0 is the vector representation of the seed instruction,
[0029] {ρ new1 ,X new1 ,Y new1} is the instruction generated for the first large model, {ρ new2 ,X new2 ,Y new2}Generated instructions for the second model;
[0030] S232, multi-model cross-validation: cross-validate the backup instructions generated by different models and eliminate duplicate content;
[0031] S233. Set dynamic stop and generation termination conditions. After reaching the dynamic stop and generation termination conditions, stop generating new instructions, and mix the extended instructions with all the original instructions to obtain an instruction set for the aviation field.
[0032] S3. Screen the high-quality instruction set for the aviation field, specifically including the following sub-steps:
[0033] S31. Preliminary screening of the high-quality instruction set for the aviation field: Score each instruction in the instruction set for the aviation field obtained after extension, sort the comprehensive results of the scoring, retain the instruction data with both top rankings, remove the instructions with too low scores or inconsistent rankings, and finally form a preliminary instruction set.
[0034] S32. Fine screening of the high-quality instruction set for the aviation field. Screen the data of the preliminary instruction set through two indicators: quality assessment and knowledge coverage to obtain the final high-quality instruction set for the aviation field.
[0035] Preferably, step S32 specifically includes the following sub-steps:
[0036] S321. Quality assessment based on the reward model: First, splice the instruction statement, input, and output into a complete instruction expression, and then send the spliced instruction set into the reward model for scoring. When the score exceeds the scoring threshold, it is considered that the data quality meets the standard.
[0037] S322. Screening based on the coverage range, specifically including the following sub-steps:
[0038] S3221. Define the semantic embedding representation V of the instruction data set (ρ,X,Y) , and each instruction triple (ρ, X, Y) is represented as its embedding vector;
[0039] S3222. Use the K-Center-Greedy algorithm to gradually select the instruction triples with the greatest diversity from the data set that has passed the quality assessment until the preset coverage range is met or the scale of the data set is minimized, complete the data screening, and obtain the final high-quality instruction set for the aviation field.
[0040] Preferably, step S3222 specifically includes the following sub-steps:
[0041] First, randomly select a point from the data set as the center v c ;
[0042] Second, for each unselected data point v i , calculate the minimum distance between the data point v i and the set of selected center points:
[0043] d min (v i ) = min(v i , v c ) · u(v i , v c )
[0044] Among them, u(v i , v c ) is the penalty factor, and the formula is as follows:
[0045]
[0046] Among them, C is the normalization constant, which is used to approximate the maximum nearest neighbor distance;
[0047] Finally, select the point with the largest distance from the center point set as the instruction triple with the greatest diversity:
[0048]
[0049] Repeat this process until the preset coverage range is met or the scale of the data set is minimized.
[0050] Preferably, step S21 specifically includes the following sub-steps:
[0051] S211. Extract and filter the key features of the input statement and output statement of each instruction:
[0052] Use the pre-trained model BERT to perform semantic representation on the input and output:
[0053]
[0054] Among them, X i is the instruction input, Y i is the expected instruction output, is the semantic representation of the instruction input, is the semantic representation of the instruction output;
[0055] S212. For the vector representations QX i , QY i of the input statement and output statement of each instruction in the aviation knowledge Q&A task, calculate the semantic similarity S(QX i , QY i ):
[0056]
[0057] S213. Set the similarity threshold θ, and filter out the high-correlation input statements and output statements that meet the conditions. The formula is as follows:
[0058] Z filtered = {(QX i , QY i ) | S(QX i , QY i ) > θ, i = 1, 2,.., N}
[0059] Where N is the number of extracted aviation knowledge Q&A task instructions, and θ is the similarity threshold;
[0060] S214. Introduce a feature grouping strategy based on the K-Means clustering algorithm for each task; concatenate the semantic representations of the input-output pairs to obtain a joint vector representation Perform K-Means clustering on all vectors V i :
[0061] C k = {v i | Cluster(v i ) = k, i = 1, 2,.., N}
[0062] Where C k represents the k-th clustering cluster, and through K-Means clustering, the input-output pairs are divided into different topic clusters;
[0063] S215. Each task generates a large number of clustering clusters, and within the clusters are instruction inputs and output pairs, finally obtaining a large set of few-shot instruction inputs and output pairs.
[0064] Preferably, after mixing in step S221, we get:
[0065]
[0066] Where T is the number of topic / domain clusters, R is a certain task, R ∈ {aviation knowledge Q&A, literature abstract generation, keyword generation, literature bibliography generation, paragraph generation, paragraph optimization, paragraph continuation}, and the number of all few-shot instruction input and output pair sets is T i represents the number of topic / domain clusters in the i-th task, is the number of input and output pairs under the j-th topic / domain cluster of the n-th task;
[0067] In step S222, use the maximum similarity S(ρ, X i , Y i ) to score the similarity of generating Y i based on the instruction ρ for the input X i , quantifying the similarity degree between the generated output and the expected output:
[0068]
[0069] Among them, is the vector representation of the output generated by instruction ρ for input X i ;
[0070] In step S223, the optimization objective function L2(ρ) is:
[0071]
[0072] where w i is the weight of the task priority, used to reflect the relevance or importance of different tasks, and the weight w i is dynamically adjusted according to the actual requirements of the task.
[0073] Preferably, in step S11, first retrieve scientific research papers related to the aviation field from the open paper bibliographic query interface as the input data for constructing the aviation field instruction set, and then define the aviation field instruction set tasks based on the obtained aviation scientific research literature data and common tasks for constructing the instruction set.
[0074] Preferably, step S233 is specifically: to avoid generating too many similar instructions, set a generation termination condition. After each round of generation, calculate the overall similarity S of the newly generated instructions, inputs, and output instances with the existing set overall , and the calculation formula is as follows:
[0075]
[0076] where D is the existing set, and D new is the newly generated set. When S overall remains high continuously for multiple times, stop generating to form the final instruction set.
[0077] Preferably, step S31 specifically includes the following sub-steps:
[0078] S311. Score each instruction in dimensions such as accuracy, integrity, and clarity through model M 1 and model M 2 respectively. The scoring formula is:
[0079] Score(ρ) = min(35, A(ρ)) + min(35, C(ρ)) + min(30, Q(ρ))
[0080] where A(ρ) refers to accuracy, C(ρ) refers to integrity, and Q(ρ) refers to clarity;
[0081] S312. Based on the scoring results of the two models, perform weighted average processing on each instruction to obtain a comprehensive scoring result. The calculation formula for the comprehensive scoring result is as follows:
[0082]
[0083] Among them, is the scoring result of model M 1 , is the scoring result of model M 2 . α and β are the weight coefficients of model M 1 and model M 2 respectively, and are usually set to be equal;
[0084] S313. Sort the final scoring results, retain the instruction data with both top rankings, remove the instructions with too low scores or inconsistent rankings, and finally form a preliminary instruction set.
[0085] Preferably, step S1 specifically includes the following sub-steps:
[0086] S11. Define the tasks of obtaining aviation scientific research literature data and the instruction set. The instruction set tasks include aviation knowledge Q&A, literature abstract generation, keyword generation, literature catalog generation, paragraph generation, paragraph optimization, and paragraph continuation;
[0087] S12. Generate the input statements and output statements of the aviation scientific research literature instruction set: Generate the input statements and output statements respectively for the instruction set tasks.
[0088] Preferably, on the other hand, the present invention also provides a multi-dimensional extended aviation domain large model instruction set generation system, which includes an instruction set input statement and output statement generation unit, an aviation domain instruction set generation unit, and a screening unit. The instruction set input statement and output statement generation unit is used to generate the instruction set input statements and output statements based on aviation scientific research literature; the aviation domain instruction set generation unit is used to generate a multi-dimensional extended aviation domain instruction set; the screening unit is used to screen high-quality aviation domain instruction sets and output the final instruction set.
[0089] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0090] (1) By using the large model to automatically construct the input and output in different task instruction sets in the aviation field, the present invention reduces the need for manual intervention and quickly responds to diverse task requirements, thus solving the problem of low efficiency of traditional instruction data generation methods.
[0091] (2) The present invention generates a rich and diverse instruction set by automatically constructing instructions in the instruction data using a large model and expanding the instructions in multiple dimensions, meeting the requirements of various application scenarios, solving the problems of insufficient instruction diversity and insufficient dimensions of the instruction set, and improving the practicality and adaptability of the instruction data.
[0092] (3) The present invention constructs a set of high-quality automatic screening methods for instruction data, effectively filtering and evaluating the quality, coverage, etc. of the generated instructions, ensuring the accuracy and practicality of the final instruction set, and solving the problem that it is difficult to control the quality of the automatically generated instruction set. Brief Description of the Drawings
[0093] Figure 1 is a schematic flowchart of the present invention;
[0094] Figure 2 is a schematic overall flowchart of the present invention;
[0095] Figure 3 is a schematic block diagram of the system structure of the present invention. Detailed Embodiments
[0096] Hereinafter, the embodiments of the present invention will be described with reference to the drawings.
[0097] Specifically, the present invention provides a method for generating a multi-dimensionally extended large model instruction set in the aviation field, as Figure 1 and Figure 2 shown, which includes the following steps:
[0098] S1. Generate instruction set input statements and output statements based on aviation research literature, specifically including the following sub-steps:
[0099] S11. Define the acquisition of aviation research literature data and instruction set tasks. The instruction set tasks include aviation knowledge Q&A, literature abstract generation, keyword generation, literature table of contents generation, paragraph generation, paragraph optimization, and paragraph continuation. The specific task definitions are introduced as follows:
[0100] Aviation knowledge Q&A: By analyzing the content of aviation research literature, form Q&A instruction pairs. The question types can cover basic knowledge, cutting-edge technologies, experimental data, etc.
[0101] Literature abstract generation: Generate a concise and information-rich abstract based on the given literature title to help researchers quickly obtain the core content of the literature.
[0102] Keyword generation: Generate keywords highly relevant to the literature content based on the given literature title and literature abstract (the literature abstract is not required) for literature retrieval and classification.
[0103] Bibliography generation: Based on the given literature title and literature abstract (the literature abstract is not required), generate the bibliography of the literature to help researchers quickly browse the organization and main content of the literature.
[0104] Paragraph generation: Generate high-quality paragraphs according to the given paragraph titles to help researchers with preliminary writing or content expansion on specific topics.
[0105] Paragraph optimization: Optimize the existing paragraphs in scientific research literature to make their expressions clearer and logic more rigorous, improving the overall text quality.
[0106] Paragraph continuation: Generate subsequent content based on the beginning paragraph of the given aviation scientific research literature to help researchers expand the writing of the literature.
[0107] S12. Generate input statements and output statements for the instruction set of aviation scientific research literature: Generate input statements and output statements for the instruction set tasks respectively. Specifically, for the aviation knowledge Q&A task, each aviation scientific research literature is segmented according to a fixed number of characters and sent to a large model such as GPT-4 in turn, allowing the large model to automatically extract 0 or 1 knowledge Q&A pairs from it, and the question forms should be diverse, including imperative sentences, interrogative sentences, etc. If extraction fails, it can be not extracted. Example of the extracted instruction input and output:
[0108] Input: Please explain the influence of supersonic airflow on the aircraft structure.
[0109] Output: The influence of supersonic airflow on the aircraft structure is mainly manifested as high heat and high pressure effects, which may lead to material fatigue and structural damage.
[0110] For the literature abstract generation task, directly extract the title and abstract of each scientific research literature. The instruction input is the title, and the instruction output is the corresponding abstract.
[0111] For the keyword generation task, directly extract the title and abstract of each scientific research literature. The instruction input is the title and abstract, or only the title, and the instruction output is the corresponding keywords.
[0112] For the bibliography generation task, directly extract the title and abstract of each scientific research literature. The instruction input is the title and abstract, or only the title, and the instruction output is the corresponding bibliography.
[0113] For the paragraph generation task, directly extract the titles and corresponding contents of some paragraphs in some scientific research literatures. The instruction input is the title, and the instruction output is the corresponding content.
[0114] The paragraph optimization task is to directly extract the titles and corresponding contents of some paragraphs in a part of scientific research literature. The corresponding contents are used as the optimized paragraphs. For the paragraphs before optimization, a relatively poor version is generated by a large model such as GPT-4. Therefore, the instruction input is the paragraph title plus the poor version generated by the large model, and the instruction output is the original paragraph in the literature.
[0115] The task of paragraph continuation is to extract the title and corresponding content of a part of the paragraphs in a part of scientific research literature. The corresponding content is used as the continued paragraph. For the paragraph before continuation, it is obtained by truncating the beginning of the original paragraph. Therefore, the instruction input is the paragraph after the beginning is truncated, and the instruction output is the original paragraph in the literature.
[0116] S2. Generate a multi-dimensionally extended aviation domain instruction set, specifically including the following sub-steps:
[0117] S21. Generate command input and output based on aviation scientific research literature for each task, and obtain multiple sets of few-sample command input and output pairs, specifically:
[0118] S211, extract and filter key features of the input statement and output statement of each instruction:
[0119] Use the pre-trained model BERT to semantically represent the input and output:
[0120]
[0121] Among them, X i For command input, Y i is the expected instruction output, is the semantic representation of the instruction input, The semantic representation of the instruction output.
[0122] S212, vector representation of the input statement and output statement of each instruction in the aviation knowledge question answering task (QX i ,QY i ), calculate the semantic similarity S(QX i ,QY i ):
[0123]
[0124] S213, setting a similarity threshold θ, screening out high-correlation input sentences and output sentences that meet the conditions,
[0125] Z filtered ={(QX i ,QY i )|S(QX i ,QYi ) > θ, i = 1, 2,.., N}
[0126] Where N is the number of extracted aviation knowledge Q&A task instructions.
[0127] S214. For each task, introduce a feature grouping strategy based on the K-Means clustering algorithm; concatenate the context representations of the input and output (X i , Y i ) to obtain a joint representation Perform K-Means clustering on all vectors V : i C
[0128] = {v k | Cluster(v i ) = k, i = 1, 2,.., N} i
[0129] Where C k represents the k-th clustering cluster. Through K-Means clustering, the input-output pairs are divided into different topic clusters.
[0130] S215. Finally, each task generates a large number of few-shot instruction input and output pair sets.
[0131] S22. Generate instruction statements based on multi-level optimization, specifically including the following sub-steps:
[0132] S221. Task description: Without distinguishing tasks, mix the above-mentioned multiple few-shot instruction input and output pair sets to obtain:
[0133]
[0134] Where T is the number of topic / domain clusters, R is a certain task, R ∈ {aviation knowledge Q&A, literature abstract generation, keyword generation, literature bibliography generation, paragraph generation, paragraph optimization, paragraph continuation}, and the number of all few-shot instruction input and output pair sets is T i represents the number of topic / domain clusters in the i-th task, is the number of input and output pairs under the j-th topic / domain cluster of the n-th task.
[0135] For one of the few-shot input and output pair sets The number of input and output pairs in the set is γ, where X i is the instruction input, Y i is the expected instruction output, and the missing instruction ρ is generated according to the instruction input and the instruction output, so that the instruction ρ can be obtained given the input X i Under the condition of i The output that best matches.
[0136] This step generates a candidate instruction sentence set {ρ 1 ,ρ 2 ,...,ρ n}, and perform a two-step optimization on the generation problem to ensure that the generated instruction statements match the input-output pairs, while improving their usability and adaptability in actual tasks.
[0137] S222. Minimize the semantic deviation of instruction statements.
[0138] By maximizing the similarity S(ρ,X i ,Y i ) based on the instruction ρ for the input X i Generate Y i The similarity score quantifies the similarity between the generated output and the expected output:
[0139]
[0140] in, The instruction ρ is the input X i A vector representation of the generated output.
[0141] The optimization objective function L1(ρ) for maximizing similarity is specifically:
[0142]
[0143] S223. Introduce domain correlation constraint optimization objective function: The optimization objective function L2(ρ) is:
[0144]
[0145] Among them, w i is the weight of the field priority, which is used to reflect the relevance or importance of different fields. The weight w i Dynamically adjust according to the actual needs of the field.
[0146] After completing the first step of optimization, in order to enhance the actual application effect of the instruction statements, a domain priority weight mechanism is introduced in the optimization process. The domain priority is scored by weighted similarity, so that the generated instruction statements not only meet the semantic matching in theory, but also have higher practicality in actual task scenarios.
[0147] S224. Determine the optimal instruction statement and the suboptimal instruction statement.
[0148] By performing the above steps, each set of input and output of the few-sample instruction can be calculated according to S(ρ,Xi , Y i ) Sort the candidate instruction statements in descending order. After the descending order sorting, take the instruction statement ranked first as the optimal instruction statement, that is, the one that satisfies the following formula
[0149]
[0150] After that, introduce the diversity score Select the sub-optimal instruction statement
[0151]
[0152] Determine based on the Euclidean distance. The larger the distance, the greater the difference in instructions:
[0153]
[0154] At the same time, in order to balance the similarity and diversity of instruction statements, introduce the adjustment parameter λ. This parameter determines the diversity and the weight relationship of the similarity S(ρ, X i , Y i ) when selecting the sub-optimal instruction statement. When λ is larger, it is more inclined to select instruction statements that match the input and output; when λ is smaller, more attention is paid to the diversity of instruction statements.
[0155] S225. After obtaining the optimal instruction statement and the sub-optimal instruction statement, for each instruction input and output pair in the few-shot instruction set, randomly select either the optimal instruction statement or the sub-optimal instruction statement as its instruction statement, so as to increase the diversity of instruction data. For all few-shot instruction sets, continuously repeat the above steps until each instruction contains the complete three parts of {instruction statement, input, output}.
[0156] Through the above steps, not only can it ensure a high matching degree between the generated instructions and the input-output pairs, but also it can generate diverse instruction statements to cover a variety of different application scenarios.
[0157] S226. Seed instruction extraction and multi-dimensional prompt word setting for instructions
[0158] For the seven major tasks, extract one instruction statement ρ 0 and one input and output instance (X 0 , Y 0 ), as the seed instruction, which is used to guide the expansion of subsequent instructions.
[0159] In this stage, first sample two instruction statements ρ 1 and ρ 2 from the seed instruction set., and respectively specify input and output examples (X 1 , Y 1 ) and (X 2 , Y 2 ). Initially, both instruction statements come from the seed instruction set of the seven major tasks, while in subsequent iterations, one of the two instruction statements comes from the generation of the large model. The generation process of the instruction is described by the following formula, where Model represents the large model. In the embodiments of the present invention, the large model generally selects GPT-4 to generate new instruction statements:
[0160] {ρ new , X new , Y new} = Model({ρ i , X i , Y i}), i = 1, 2
[0161] When generating new instructions, there are the following prompt requirements:
[0162] Diversification of statement types: The generated instruction statements should include various statement types, such as open imperative sentences, interrogative sentences, etc. At the same time, the generated input and output statements must also be diverse. For some special tasks, only output without input is also supported.
[0163] Instruction extension dimension: The instruction is extended in two directions: depth extension and breadth extension.
[0164] Depth extension: Depth extension is to complicate and refine the instruction statements, input statements, and output statements within the existing task framework. For example, on the basis of the existing instructions, expand the constraint conditions or introduce multi-step reasoning. A simple example is:
[0165] Before depth extension:
[0166] Instruction input: How to check the engine performance of an aircraft?;
[0167] Instruction output: Conduct an engine performance test and check the thrust, fuel consumption, and temperature data.
[0168] After depth extension:
[0169] Instruction input: How to check the engine performance of an aircraft under different flight phases and environmental conditions?;
[0170] Instruction output:
[0171] During the takeoff phase: Monitor the thrust, throttle setting, and temperature to ensure that the thrust reaches the standard.
[0172] During the cruise phase: Check fuel consumption, engine speed and temperature, and record various data to evaluate efficiency.
[0173] During the landing phase: Monitor reverse thrust and temperature changes to ensure the engine safely decelerates...
[0174] Breadth expansion: Breadth expansion emphasizes that the generated instructions should belong to the aviation field but not fall into the above seven major tasks or problem forms. For example, generate instruction statements in the form of multiple-choice questions:
[0175] Instruction input: Which of the following is not part of the standard procedure during the pre-flight inspection?
[0176] A. Check the fuel level of the aircraft
[0177] B. Verify the accuracy of the flight plan
[0178] C. Conduct a mental health assessment of the pilot
[0179] D. Confirm the normal function of the on-board equipment
[0180] Instruction output: C. Conduct a mental health assessment of the pilot.
[0181] S23. Generate instructions with multiple models and cross-validate, which specifically includes the following sub-steps:
[0182] S231. Generate new instructions: Generate new instruction statements and their corresponding input and output instances {ρ new1 ,X new1 ,Y new1} and {ρ new2 ,X new2 ,Y new2}, and calculate the vector representations after concatenating the instruction statements, input statements, and output statements respectively Calculate the similarity between each generated instruction and the seed instruction:
[0183]
[0184] In this embodiment, in step S231, two large models, GPT-4 and Qwen, are respectively used to generate new instruction statements and their corresponding input and output statements.
[0185] S232. Cross-validate with multiple models: Cross-validate the instruction statements, inputs, and outputs generated by different models, eliminate duplicate content, and retain the unique instruction statements, inputs, and output instances that do not appear in different models.
[0186] S233. Set dynamic stop and generation termination conditions, and terminate generating new instructions after reaching the dynamic stop and generation termination conditions.
[0187] Specifically, step S233 is as follows: To avoid generating too many similar instructions, a generation termination condition is set. After each round of generation, the overall similarity S between the newly generated instruction statements, inputs, and output instances and the existing set is calculated. overall The calculation formula is as follows:
[0188]
[0189] where D is the existing set, and D new is the newly generated set. When S overall remains high continuously for multiple times, the generation stops. The expanded instructions are mixed with all the original instructions to obtain the final instruction set for the aviation field.
[0190] S3. Screen the high-quality instruction set for the aviation field, which specifically includes the following sub-steps:
[0191] S31. Preliminary screening of the high-quality instruction set for the aviation field. Each instruction in the instruction set for the aviation field is scored, and the comprehensive results of the scoring are sorted. The instruction data with both top rankings are retained, and the instructions with too low scores or inconsistent rankings are removed to finally form a preliminary instruction set.
[0192] Through similarity calculation, instructions, inputs, and output instances with too high similarities are screened out, and combinations with similarities lower than a certain threshold are retained to ensure the high diversity of the generated content.
[0193] Preferably, step S31 specifically includes the following sub-steps:
[0194] S311. Each instruction is scored in multiple dimensions such as accuracy, integrity, and clarity through model M 1 and model M 2 respectively. The scoring formula is:
[0195] Score(ρ) = min(35, A(ρ)) + min(35, C(ρ)) + min(30, Q(ρ))
[0196] where A(ρ) refers to accuracy, C(ρ) refers to integrity, and Q(ρ) refers to clarity.
[0197] Accuracy A(ρ) refers to evaluating whether the generated instruction conforms to the input and output task requirements to ensure the correctness of the generated instruction.
[0198] Integrity C(ρ) refers to ensuring that the instruction output (i.e., the answer part) covers all the key points of the question to guarantee the integrity of the generated result.
[0199] Clarity Q(ρ) refers to evaluating whether the instruction statement structure is clear and the logic is smooth to ensure the readability and coherence of the instruction.
[0200] S312. Based on the scoring results of the two models, perform weighted average processing on each instruction to obtain a comprehensive scoring result. The calculation formula for the comprehensive scoring result is as follows:
[0201] Score(ρ) = α × Score M1 (ρ) + β × Score M2 (ρ)
[0202] where Score M1 (ρ) is the scoring result of model M 1 , Score M2 (ρ) is the scoring result of model M 2 , α and β are the weight coefficients of model M 1 and model M 2 respectively, and are usually set to be equal.
[0203] S313. Sort the final scoring results, retain the instruction data with both top rankings, remove the instructions with too low scores or inconsistent rankings, and finally form a preliminary instruction set.
[0204] S32. Fine-screen the instruction set in the high-quality aviation field. Screen the data of the preliminary instruction set through two indicators: quality assessment and knowledge coverage to obtain the final high-quality aviation field instruction set.
[0205] Preferably, step S32 specifically includes the following sub-steps:
[0206] S321. Quality assessment based on the reward model: First, splice the instruction statement, input instance, and output instance into a complete instruction expression, and then send the spliced instruction set into the reward model for scoring. When the score exceeds the scoring threshold, it is considered that the data quality meets the standard; in this embodiment, the reward model is the reward-model-debertav3-large-v2 reward model.
[0207] S322. Screening based on the coverage range, specifically including the following sub-steps:
[0208] S3221. Define the semantic embedding representation V (ρ,X,Y) of the instruction dataset, and each instruction triple (ρ, X, Y) is represented as its embedding vector.
[0209] S3222. Use the K-Center-Greedy algorithm to gradually select the instruction triples with the greatest diversity from the dataset that has passed the quality assessment, and maximize the coverage range of the selected data through the distance function d(v a , v b ):
[0210] d(va , v b ) = ||v a -v b || 2
[0211] where v a and v b are the embedding vectors of instructions a and b, and |||| 2 represents the Euclidean distance.
[0212] In this embodiment, step S3222 uses the K-Center-Greedy algorithm to gradually select the instruction triple with the greatest diversity from the dataset that has passed the quality assessment, which specifically includes the following sub-steps:
[0213] First, randomly select a point from the dataset as the center v c .
[0214] Second, for each unselected data point v i , calculate its minimum distance from the set of selected center points:
[0215] d min (v i ) = min(v i , v c ) · u(v i , v c )
[0216] where u(v i , v c ) is a penalty factor adaptively adjusted according to the data density, which is used to increase the selection probability of sparse regions and help the K-Center-Greedy algorithm better cover the entire data space, especially those regions with lower density. Specifically, the average distance between v i and its k nearest neighbors is used to adjust the value of u(v i , v c ). If the average distance between v i and other points is large, indicating that v i is in a sparse region, then the value of u(v i , v c ) should be large. The formula is as follows:
[0217]
[0218] where C is a normalization constant used to approximate the maximum nearest neighbor distance.
[0219] Finally, select the point with the maximum distance from the set of center points:
[0220]
[0221] Repeat this process until the preset coverage is met or the scale of the dataset is minimized.
[0222] On the other hand, the present invention also provides a multi-dimensional extended large model instruction set generation system for the aviation field, as Figure 3 shown, which includes an instruction set input statement and output statement generation unit 1, an aviation field instruction set generation unit 2, and a screening unit 3. The instruction set input statement and output statement generation unit 1 is used to generate instruction set input statements and output statements based on aviation research literature; the aviation field instruction set generation unit 2 is used to generate a multi-dimensional extended aviation field instruction set; the screening unit 3 is used to screen high-quality aviation field instruction sets and output the final instruction set.
[0223] In summary, the present invention automatically constructs the input and output in different task instruction sets in the aviation field by using a common large model, reduces the need for manual intervention, can quickly respond to diverse task requirements, thereby solving the problem of low efficiency of traditional instruction data generation methods, and can ensure the accuracy of instruction generation.
[0224] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for generating a large model instruction set in the aviation field with multi-dimensional expansion, characterized by: It includes the following steps: S1. Generate input statements and output statements of instructions based on aviation scientific research literature; S2. Generate a multi-dimensionally extended aviation domain instruction set, specifically including the following sub-steps: S21, generating an input statement and an output statement of an instruction for each task, and obtaining a plurality of sets of input and output pairs of a few samples of instructions; S22, the instruction statement generated based on the multi-level optimization specifically includes the following sub-steps: S221, mixing the above-mentioned plurality of few-sample instruction input and output pair sets, for one of the few-sample input and output pair sets The number of input and output pairs in the set is γ. The missing instruction statement ρ is generated by the large model based on the input statement and the output statement, so that the instruction statement ρ can be found in the given input X i Under the condition of i The most matching output; S222, minimizing the semantic deviation of the instruction statement ρ; S223, introducing a domain correlation constraint optimization objective function to optimize the instruction statement ρ; S224, determine the optimal instruction statement and the suboptimal instruction statement: Execute the above steps, and sort the candidate instruction statements in descending order for each set of few-sample instruction input and output pairs. After descending order, the instruction statement ranked first is taken as the optimal instruction statement, that is, the optimal instruction statement that satisfies the following formula Later, the diversity score was introduced Select suboptimal instruction statement Among them, S(ρ,X i ,Y i ) is the similarity of the instruction statements, is the diversity of instruction statements, and λ is a tuning parameter used to adjust the weight between similarity and diversity; S225, after obtaining the optimal instruction statement and the suboptimal instruction statement, for each set of few-sample instruction input and output pairs, randomly select the optimal instruction statement or the suboptimal instruction statement as the instruction statement; for all sets of few-sample instruction input and output pairs, continuously repeat the above steps until each instruction contains the complete three parts of {instruction statement, input statement, output statement}; S226, randomly extracting an instruction from each task as a seed instruction {ρ0, X0, Y0} and setting a large model for extending the prompt words of the instruction; S23, multi-model instruction expansion and cross-validation, specifically includes the following sub-steps: S231, expand new instructions: use two large models to expand the original task instructions in depth and breadth, and use seed instructions as initial instructions to iteratively generate new instructions in sequence {ρ new1 ,X new1 ,Y new1 } and {ρ new2 ,X new2 ,Y new2 }, calculate the similarity between each generated instruction statement and the seed instruction statement, and based on the similarity value, select multiple instructions with low similarity as backup instructions. new ,ρ0) is calculated as: In the formula, is the vector representation of the newly generated instruction, is the vector representation of the seed instruction, {ρ new1 ,X new1 ,Y new1 } is the instruction generated for the first large model, {ρ new2 ,X new2 ,Y new2 }Generated instructions for the second model; S232, multi-model cross-validation: cross-validate the backup instructions generated by different models and eliminate duplicate content; S233, setting dynamic stop and generation termination conditions, terminating the generation of new instructions after the dynamic stop and generation termination conditions are met, and mixing the expanded instructions with all original instructions to obtain an aviation field instruction set; S3. Screening high-quality aviation instruction sets, including the following sub-steps: S31. Preliminary screening of high-quality aviation instruction sets: Score each instruction in the expanded aviation instruction set, sort the comprehensive scoring results, retain the instructions with the highest rankings, remove the instructions with low scores or inconsistent rankings, and finally form a preliminary instruction set; S32. Detailed screening of high-quality aviation field instruction sets. Data screening of preliminary instruction sets is performed through two indicators: quality assessment and knowledge coverage, to obtain the final high-quality aviation field instruction set.
2. The method for generating a large model instruction set for a multi-dimensionally expanded aviation field according to claim 1, characterized in that: Step S32 specifically includes the following sub-steps: S321, quality assessment based on reward model: first, the instruction statement, input statement and output statement of each instruction are spliced to form a complete instruction expression, and then the spliced instruction set is sent to the reward model for scoring. When the score exceeds the scoring threshold, the data quality is considered to meet the standard; S322, screening based on coverage, specifically includes the following sub-steps: S3221. Define the semantic embedding representation V of the instruction dataset (ρ,X,Y) , each instruction triple (ρ, X, Y) is represented as its embedding vector; S3222. Use the K-Center-Greedy algorithm to gradually select the instruction triplets with the greatest diversity from the data set that has passed the quality assessment until the preset coverage is met or the size of the data set is minimized, complete the data screening, and obtain the final high-quality aviation field instruction set.
3. The method for generating a large model instruction set for a multi-dimensionally expanded aviation field according to claim 1, characterized in that: Step S3222 specifically includes the following sub-steps: First, randomly select a point from the data set as the center v c ; Second, for each unselected data point v i , calculate the data point v i Minimum distance to a set of selected center points: d min (v i )=min(v i ,v c )·u(v i ,v c ) Among them, u(v i ,v c ) is the penalty factor, and the formula is as follows: Among them, C is a normalization constant used to represent the maximum nearest neighbor distance; Finally, select the point with the largest distance from the center point set As the instruction triple with the greatest diversity: This process is repeated until the preset coverage is met or the size of the dataset is minimized.
4. The method for generating a large model instruction set for a multi-dimensionally expanded aviation field according to claim 1, characterized in that: Step S21 specifically includes the following sub-steps: S211, extract and filter key features of the input statement and output statement of each instruction: Use the pre-trained model BERT to semantically represent the input and output: Among them, X i For command input, Y i is the expected instruction output, is the semantic representation of the instruction input, The semantic representation of the instruction output; S212, vector representation QX of the input statement and output statement of each instruction in the aviation knowledge question answering task i ,QY i , calculate the semantic similarity S(QX i ,QY i ): S213, set a similarity threshold θ, and filter out high-correlation input sentences and output sentences that meet the conditions. The formula is as follows: Z filtered ={(QX i ,QY i )|S(QX i ,QY i )>θ,i=1,2,..,N} Among them, N is the number of aviation knowledge question-answering task instructions extracted, and θ is the similarity threshold; S214, introduce a feature grouping strategy based on the K-Means clustering algorithm for each task; the semantic representation of the input and output pairs Perform vector concatenation to obtain a joint vector representation For all vectors V i Perform K-Means clustering: C k ={v i |Cluster(v i )=k,i=1,2,..,N} Among them, C k represents the kth cluster, and through K-Means clustering, the input and output pairs are divided into different topic clusters; S215. Each task generates a large number of clusters, each of which contains instruction input and output pairs, and ultimately obtains a large number of sets of few-sample instruction input and output pairs.
5. The method for generating a large model instruction set for a multi-dimensionally expanded aviation field according to claim 1, characterized in that: After mixing in step S221, the following is obtained: Where T is the number of subject / domain clusters, R is a task, R∈{aviation knowledge question answering, document abstract generation, keyword generation, document catalog generation, paragraph generation, paragraph optimization, paragraph continuation}, and the number of all sets of few-sample instruction input and output pairs is T i represents the number of subject / domain clusters in the i-th task, is the number of input and output pairs under the jth topic / domain cluster of the nth task; In step S222, the similarity S(ρ, X i ,Y i ) for the input X based on the instruction statement ρ i Generate Y i The similarity score quantifies the similarity between the generated output and the expected output: in, is the instruction statement ρ for the input statement X i A vector representation of the generated output; The optimization objective function L2(ρ) in step S223 is: Among them, w i is the weight of the task priority, which is used to reflect the relevance or importance of different tasks. The weight w i Dynamically adjust according to the actual needs of the task.
6. The method for generating a large model instruction set for a multi-dimensionally expanded aviation field according to claim 1, characterized in that: In step S11, scientific research papers related to the aviation field are first retrieved from the open paper title query interface as input data for constructing the aviation field instruction set, and then the aviation field instruction set tasks are defined based on the acquired aviation scientific research literature data and common tasks constructed by the instruction set.
7. The method for generating a large model instruction set for a multi-dimensionally expanded aviation field according to claim 1, characterized in that: Step S233 is as follows: To avoid generating too many similar instructions, a generation termination condition is set. After each round of generation, the overall similarity S between the newly generated instructions and the existing set is calculated. overall , the calculation formula is as follows: Where D is the existing set, D new is a newly generated set, when S overall When it remains high for many consecutive times, the generation is stopped, and the expanded instructions are mixed with all the original instructions to obtain the aviation field instruction set.
8. The method for generating a large model instruction set for a multi-dimensionally expanded aviation field according to claim 1, characterized in that: Step S31 specifically includes the following sub-steps: S311. Use model M1 and model M2 to score each instruction in terms of accuracy, completeness, and clarity. The scoring formula is: Score(ρ)=min(35,A(ρ))+min(35,C(ρ))+min(30,Q(ρ)) Among them, A(ρ) refers to accuracy, C(ρ) refers to completeness, and Q(ρ) refers to clarity; S312. Based on the scoring results of the two models, a weighted average is performed on each instruction to obtain a comprehensive scoring result. The calculation formula for the comprehensive scoring result is as follows: in, is the scoring result of model M1, is the scoring result of model M2, α and β are the weight coefficients of model M1 and model M2 respectively and are set equal; S313, sorting the final scoring results, retaining the instruction data with the highest rankings, removing the instructions with too low scores or inconsistent rankings, and finally forming a preliminary instruction set.
9. The method for generating a large model instruction set for a multi-dimensionally expanded aviation field according to claim 1, characterized in that: Step S1 specifically includes the following sub-steps: S11. Define the tasks of aviation scientific research literature data acquisition and instruction set. The instruction set tasks include aviation knowledge question answering, literature abstract generation, keyword generation, literature catalog generation, paragraph generation, paragraph optimization, and paragraph continuation. S12. Generate input statements and output statements for the aviation scientific research literature instruction set: generate input statements and output statements for the instruction set tasks respectively.
10. An instruction set generation system for the multi-dimensionally expanded aviation field large model instruction set generation method according to claim 1, characterized in that: It includes an instruction set input statement and output statement generating unit, an aviation field instruction set generating unit and a screening unit. The instruction set input statement and output statement generating unit is used to generate instruction set input statements and output statements based on aviation scientific research literature; the aviation field instruction set generating unit is used to generate a multi-dimensionally extended aviation field instruction set; the screening unit is used to screen high-quality aviation field instruction sets and output a final instruction set.
Citation Information
Patent Citations
Session recommendation system fusing sparse graph and multi-hop attention
CN114817508A
Aviation literature keyword similarity judgment method fused with multi-modal semantic association map
CN116362221A