Method and device for generating evaluation questions in field of aerospace science and technology, and medium

By filtering through a subject classification table and a categorized corpus in the field of aerospace technology, a balanced number of evaluation questions were generated, which solved the problem of uneven distribution of the number of evaluation questions, improved the comparability and representativeness of the evaluation conclusions, and enhanced the coverage verification of the model's capabilities.

CN122045406APending Publication Date: 2026-05-15ZHEJIANG LAB
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610518662.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-20
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing technologies, the distribution of evaluation questions in the aerospace technology field across different disciplines is uneven, which can lead to biased evaluation conclusions and fail to fully reflect the model's true performance in key application scenarios.

Method used

By using a subject classification table in the field of aerospace technology, the corpus is screened and classified to ensure that the amount of corpus in each subject is balanced. A preset amount of corpus is extracted from each subject to generate evaluation questions that meet the condition of quantity balance. Combined with the optimization conditions of semantic repetition, quality and difficulty, a high-quality and high-difficulty evaluation question set is formed.

Benefits of technology

It has achieved a balance in the number of evaluation questions in the field of aerospace technology across different disciplines, improved the comparability and representativeness of evaluation conclusions, alleviated the shortcomings of evaluation conclusions that are prone to bias, and enhanced the coverage verification effect of model capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122045406A_ABST
    Figure CN122045406A_ABST
Patent Text Reader

Abstract

The invention relates to a method and device for generating evaluation questions in the field of spaceflight science and technology and a medium, and the method comprises the steps: judging whether corpora of a first corpus meet correlation conditions of the field of spaceflight science and technology or not according to a subject classification table of the field of spaceflight science and technology; removing the corpora which are judged to be satisfied from the first corpus, and putting the corpora into another corpus to obtain a second corpus; according to the subject classification table, classifying the corpora of the second corpus according to subjects; extracting a first preset number of corpora from the corpora of each subject; according to the first preset number of corpora extracted from all the subjects, aerospace science and technology field evaluation questions meeting the number balance condition are generated. According to the method and the device, the defect that the evaluation conclusion is easy to deviate due to unbalanced distribution of the number of evaluation questions in the field of spaceflight science and technology in related technologies is relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of aerospace technology, and in particular to a method, device, and medium for generating evaluation questions. Background Technology

[0002] As large language models extend from general scenarios to the aerospace technology field, model evaluation needs to rely on assessment questions within that field. Because the actual application of the model often involves multiple sub-disciplines within aerospace technology, the distribution of assessment questions across these sub-disciplines needs to be reasonably balanced. This ensures that the assessment questions comprehensively test the model's capabilities, making the evaluation conclusions comparable and representative. Conversely, if the distribution of assessment questions across the sub-disciplines is uneven, concentrated in a few, the evaluation conclusions are easily dominated by the sub-disciplines with the highest proportion of questions, making it difficult to fully reflect the model's true performance in key application scenarios.

[0003] In related technologies, the first step is to select corpora related to the aerospace technology field from general public corpora; then, through a large language model, evaluation questions in the aerospace technology field are generated according to instructions based on the selected corpora; or, evaluation questions in the aerospace technology field are compiled manually based on the selected corpora.

[0004] It is evident that both large language models and manual methods rely on the selection of corpora to generate or compile evaluation questions in the aerospace technology field. The coverage and content density of general public corpora vary across different sub-disciplines. For example, the first sub-discipline has a smaller proportion of corpora, while the second sub-discipline has a larger proportion. When corpora from the entire corpus are sampled with equal random probability, the sampling results will naturally skew towards the proportion of corpora. Because the second sub-discipline has a larger corpus, it is sampled more frequently, resulting in a more concentrated sampling result across the second sub-discipline. Consequently, more evaluation questions in the aerospace technology field are generated or compiled in the second sub-discipline, leading to an uneven distribution of evaluation questions across the disciplines and biased evaluation conclusions. Summary of the Invention

[0005] This application provides a solution to address some or all of the shortcomings in the related technologies.

[0006] According to a first aspect of the embodiments of this application, a method for generating evaluation questions in the aerospace technology field for large language models is provided, the method comprising: Based on the subject classification table of aerospace technology, determine whether the corpus of the first corpus meets the relevance condition to the aerospace technology field; Remove the corpus whose judgment result is satisfied from the first corpus and put it into another corpus to obtain the second corpus; According to the subject classification table, the corpus of the second corpus is classified by subject. Extract a first predetermined amount of data from the corpus of each subject. Based on the first preset amount of corpus extracted from all disciplines, generate evaluation questions in the field of aerospace technology that meet the quantity balance condition.

[0007] Optionally, before extracting a first preset number of corpora from the corpus of each subject, the generation method further includes: For subjects where the quantity of corresponding corpus is less than the first preset quantity, supplement the corresponding corpus until the quantity of corresponding corpus reaches the first preset quantity.

[0008] Optionally, after generating the aerospace technology evaluation questions that satisfy the quantity balance condition, the generation method further includes: Among the evaluation questions in the field of aerospace technology, those that meet the optimization criteria in terms of semantic redundancy, quality, and difficulty are selected as the preferred evaluation questions in the field of aerospace technology. If the evaluation questions in the preferred aerospace technology field do not meet the quantity balance condition, supplementary corpus is obtained for subjects whose proportion of the corresponding evaluation questions in the preferred aerospace technology field is lower than a preset proportion threshold. Based on the supplementary corpus, supplementary aerospace technology evaluation questions are obtained that satisfy the preferred conditions in terms of semantic redundancy, quality, and difficulty. The supplementary aerospace technology evaluation questions are merged with the preferred aerospace technology evaluation questions to obtain new preferred aerospace technology evaluation questions. Then, when the preferred aerospace technology evaluation questions do not meet the quantity balance condition, supplementary corpus is obtained for the disciplines whose proportion of the corresponding preferred aerospace technology evaluation questions is lower than a preset proportion threshold.

[0009] Optionally, among the evaluation questions in the aerospace technology field, those that meet the preferred criteria in terms of semantic redundancy, quality, and difficulty are selected as preferred aerospace technology field evaluation questions, including: Semantic deduplication is performed on the evaluation questions in the aerospace technology field. Based on the deduplicated aerospace technology evaluation questions, high-quality aerospace technology evaluation questions that meet the preset quality conditions are obtained. The high-quality aerospace technology assessment questions are divided into high-difficulty aerospace technology assessment questions and low-difficulty aerospace technology assessment questions according to their difficulty. The preferred aerospace technology evaluation questions are those with high difficulty and a sampling accuracy that meet the preset accuracy conditions.

[0010] Optionally, the step of generating aerospace technology evaluation questions that satisfy the quantity balance condition based on the first preset amount of corpus extracted from all disciplines includes: The first preset number of corpora and prompt words extracted from each subject are input into a large language model to generate a second preset number of aerospace technology evaluation questions. The large language model is used to generate the same preset number of aerospace technology evaluation questions based on each corpus, and to summarize all aerospace technology evaluation questions for each subject to obtain the second preset number of aerospace technology evaluation questions.

[0011] Optionally, determining whether the corpus in the first corpus meets the relevance condition to the aerospace technology field based on the subject classification table includes: Based on the subject classification table and keyword table, a TF-IDF (Term Frequency-Inverse Document Frequency) vector for the keyword table is obtained; the keyword table includes keywords. Determine the TF-IDF values ​​of the keyword in the candidate corpus of the first corpus; According to the order of the keywords in the keyword list, the TF-IDF values ​​of all candidate corpora are sorted to obtain the TF-IDF vector of the candidate corpora; Determine the similarity between the TF-IDF vectors of the candidate corpus and the TF-IDF vectors of the keyword list to obtain a first similarity value; Determine the similarity between the TF-IDF vectors of each candidate corpus to obtain the second similarity value; Based on the first similarity value and the second similarity value, it is determined whether the corpus of the first corpus meets the relevance condition to the aerospace technology field.

[0012] Optionally, determining whether the corpus of the first corpus meets the relevance condition to the aerospace technology field based on the first similarity value and the second similarity value includes: When determining for the first time whether the corpus in the first corpus meets the relevance condition, it is determined that the corpus with the largest first similarity value in the first corpus meets the relevance condition.

[0013] Optionally, determining whether the corpus of the first corpus meets the relevance condition to the aerospace technology field based on the first similarity value and the second similarity value further includes: When it is not the first time to determine whether the corpus in the first corpus meets the relevance condition, for each corpus in the first corpus that is not determined to meet the relevance condition, determine the maximum value of the second similarity value between it and the corpus in the first corpus that has been determined to meet the relevance condition. The corpus with the largest difference between the first similarity value and the maximum value in the first corpus is determined to satisfy the relevance condition; Returning to each corpus in the first corpus that is not determined to meet the relevance condition, the maximum value of the second similarity value between it and the corpus in the first corpus that has been determined to meet the relevance condition is determined, until the number of corpora in the first corpus that have been determined to meet the relevance condition reaches a preset first quantity threshold.

[0014] Optionally, determining whether the corpus of the first corpus meets the relevance condition to the aerospace technology field based on the subject classification table of the aerospace technology field further includes: When the number of corpora in the first corpus that have been determined to meet the relevance conditions reaches the first threshold, the corpora in the first corpus that have been determined to meet the relevance conditions are input into a pre-trained first classification model to obtain the relevance confidence; the first classification model is used to output the relevance classification result between the corpora in the first corpus that have been determined to meet the relevance conditions and the aerospace technology field; The corpora in the first corpus that were previously determined to meet the relevance condition but whose relevance confidence score is lower than a preset confidence score threshold are re-determined to not meet the relevance condition.

[0015] Optionally, the keyword list is obtained in the following way: For each subject in the subject classification table, the same preset number of reference corpora are selected; Determine the TF-IDF values ​​of the candidate words for each subject in the subject classification table in the corresponding reference corpus; Based on the TF-IDF values ​​of the candidate words, the same preset number of keywords are determined from the candidate words of each subject in the subject classification table; All keywords are summarized to obtain the keyword table.

[0016] Optionally, obtaining the TF-IDF vector of the keyword table based on the subject classification table and the keyword table includes: All reference corpora are merged to obtain the total reference corpus; Determine the keyword TF-IDF value for each keyword in the total reference corpus; According to the order of the keywords in the keyword table, sort all the keyword TF-IDF values ​​to obtain the keyword table TF-IDF vector.

[0017] Optionally, classifying the corpus of the second corpus by subject according to the subject classification table includes: The corpus data of the second corpus is input into a pre-trained second classification model to obtain subject classification results; the second classification model is used to determine the specific subject in the subject classification table to which the corpus data of the second corpus belongs.

[0018] According to a second aspect of the embodiments of this application, an electronic device is provided, including one or more processors, for implementing the aforementioned method for generating evaluation questions in the aerospace technology field for large language models.

[0019] According to a third aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a program is stored, which, when executed by a processor, implements the aforementioned method for generating evaluation questions in the aerospace technology field for large language models.

[0020] The technical solutions provided by the embodiments of this application may include the following beneficial effects: The method for generating evaluation questions in the aerospace technology field for large language models in this application is based on the subject classification table of the aerospace technology field. It performs relevance determination on the corpus of the first corpus and transfers the corpus whose determination results meet the relevance conditions of the aerospace technology field from the first corpus to the second corpus, so as to avoid the interference of corpus unrelated to the aerospace technology field with the subsequent subject classification.

[0021] Furthermore, based on the subject classification table, the corpora in the second corpus are classified by subject, so that different subjects correspond to independent corpus sets. On this basis, a first preset number of corpora are extracted from the corpora of each subject, and based on the first preset number of corpora extracted from all subjects, evaluation questions in the field of aerospace technology that meet the quantity balance condition are generated. This ensures that the number of evaluation questions in the field of aerospace technology corresponding to each subject is constrained by the quantity balance condition, and is no longer dominated by the differences in the coverage scale and content density of general public corpora in different subjects in the field of aerospace technology, thus avoiding an uneven distribution of the number of evaluation questions in the field of aerospace technology across the subject dimensions.

[0022] Based on the above settings, the evaluation conclusions are less likely to be dominated by disciplines with a high proportion of evaluation questions in the aerospace technology field. This improves the coverage of the model's capabilities in the aerospace technology field, making the evaluation conclusions more comparable and representative in the aerospace technology field. This alleviates the defect that the uneven distribution of the number of evaluation questions in the aerospace technology field in related technologies can lead to biased evaluation conclusions.

[0023] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart illustrating a method for generating evaluation questions in the field of aerospace technology, according to an embodiment of this application. Figure 2 This is a flowchart illustrating another method for generating evaluation questions in the field of aerospace technology, according to an embodiment of this application. Figure 3 yes Figure 1 The flowchart of step S106 in the method for generating evaluation questions in the field of aerospace technology is shown. Figure 4 This is a schematic diagram of a high-difficulty aerospace technology evaluation set A according to an embodiment of this application; Figure 5 This is a schematic diagram of a low-difficulty aerospace technology evaluation set B according to an embodiment of this application; Figure 6 yes Figure 1 The flowchart shown is a step S101 of the method for generating evaluation questions in the field of aerospace technology. Figure 7 yes Figure 1 Another flowchart illustrating step S101 in the method for generating evaluation questions in the aerospace technology field; Figure 8 This is the loss (error) curve of the training process of the first classification model; Figure 9 This is a confusion matrix diagram showing the classification performance of the first classification model after training on the validation set. Figure 10This is a TruncatedSVD Visualization (truncated singular value decomposition and data visualization) graph showing the classification performance of the first classification model after training on the validation set. Figure 11 yes Figure 1 The flowchart illustrates the method for determining the keyword list in step S1011 of the method for generating evaluation questions in the aerospace technology field. Figure 12 yes Figure 4 The flowchart shown is a step S1011 of the method for generating evaluation questions in the field of aerospace technology. Figure 13 This is a schematic diagram of a module for generating evaluation questions in the field of aerospace technology, according to an embodiment of this application; Figure 14 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0026] The technical solutions in the embodiments of this application will be clearly and completely described herein with reference to the accompanying drawings. In the following description, when referring to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements.

[0027] The terms "first" and "second" used in the embodiments of this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0028] This application provides a star catalog data management method and a star catalog data management system. The star catalog data management method and system of this application will be described in detail below with reference to the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0029] Figure 1 This application illustrates a method for generating evaluation questions in the aerospace technology field for large language models, according to an exemplary embodiment. For example... Figure 1 As shown, the generation method includes steps S101 to S105: S101. Based on the subject classification table of aerospace technology, determine whether the corpus of the first corpus meets the relevance condition to the aerospace technology field.

[0030] Specifically, the first corpus is typically constructed using publicly available corpora, including sources such as Wikipedia and general question-and-answer datasets. The corpus in the first corpus may contain: books, papers, patents, technical reports, standard documents, etc. A relevance condition is used to determine whether the corpus in the first corpus is related to any discipline in the aerospace technology subject classification table. Corpora that meet this condition can be used to generate subsequent aerospace technology evaluation questions.

[0031] First corpora related to aerospace technology, obtained through various means, often contain a large amount of pseudo-domain-related data. For example, some aerospace-themed literary texts or general science articles that only mention names like space stations and artificial satellites involve only conceptual descriptions and non-engineering content, having no substantial connection to core aerospace technology areas such as launch vehicle development, spacecraft guidance and control, and aerospace material preparation. However, in the initial corpus construction stage, to ensure corpus size, screening criteria are usually relaxed, making such texts easily mixed into the corpus.

[0032] Therefore, determining whether the data in the first corpus meets the relevance criteria to the aerospace technology field, and conducting aerospace technology relevance assessment and screening on the data in the first corpus, can effectively eliminate pseudo-relevant noise data and provide a reliable data foundation for the subsequent generation of aerospace technology evaluation questions.

[0033] The subject classification table is a list of subjects within the aerospace technology field, broken down into sub-disciplines. Specifically, the subject classification table for aerospace technology can be obtained through the following steps: Step 1: Extracting basic disciplines Based on the "GB / T13745-2009 Subject Classification and Codes", experts in the field of aerospace technology sorted out the secondary and tertiary subjects, removed basic subjects that are not related to aerospace technology, formed a list of basic subjects applicable to the field of aerospace technology, and used this to build the core framework of the subject classification table.

[0034] Step 2: Domain Relevance Filtering Based on the core business scenarios and technical needs in the aerospace technology field, the correlation strength of each basic discipline was assessed one by one. Strongly related disciplines were retained, weakly related disciplines were marked and secondary verification was carried out to ensure that the screening results fit the actual application scope of the field.

[0035] Step 3: Optimize and validate the category table Experts in the field, taking into account the interdisciplinary nature of the disciplines, broke down, merged and adjusted the selected basic disciplines, added specific sub-directions for the aerospace technology field, standardized the discipline names and classification logic, and formed an initial version of the discipline classification table.

[0036] Experts from fields who were not involved in the initial work conducted cross-validation and filled in gaps in the initial classification table, ultimately forming a complete subject classification table that is suitable for the aerospace technology field.

[0037] The following is a subject classification table in the field of aerospace technology, adapted based on the above steps: Table 1: Subject Classification Table in the Field of Aerospace Science and Technology

[0038] S102. Remove the corpus whose judgment result is satisfied from the first corpus and put it into another corpus to obtain the second corpus.

[0039] S103. According to the subject classification table, classify the corpus of the second corpus by subject.

[0040] S104. Extract a first preset amount of data from the corpus of each subject.

[0041] The first preset quantity is the number of corpora extracted uniformly from the corpus set corresponding to each discipline after the corpus of the second corpus has been classified by discipline.

[0042] S105. Based on the first preset number of corpora extracted from all disciplines, generate evaluation questions in the field of aerospace technology that meet the condition of quantity balance.

[0043] The quantity balance condition is that the number of evaluation questions in the aerospace technology field corresponding to each discipline is evenly distributed. This can be either completely consistent or similar within a preset ratio range, without the need for strict equality.

[0044] The above steps effectively constrain the distribution of assessment questions across disciplines, improving the comparability and representativeness of the evaluation results. The method categorizes the second corpus by discipline according to a subject classification table, ensuring each discipline corresponds to an independent corpus set. Then, a first predetermined number of questions are extracted from the corpus of each discipline. Based on these extracted questions from all disciplines, assessment questions in the aerospace technology field that meet the quantity balance condition are generated. This ensures that the number of assessment questions in the aerospace technology field corresponding to each discipline is constrained by the quantity balance condition, thereby avoiding an uneven distribution of assessment questions across disciplines.

[0045] The method for generating evaluation questions in the aerospace technology field for large language models in this embodiment is based on the subject classification table of the aerospace technology field. It performs relevance determination on the corpus of the first corpus and transfers the corpus whose determination results meet the relevance conditions of the aerospace technology field from the first corpus to the second corpus, so as to avoid corpus that is not related to the aerospace technology field from interfering with the subsequent subject classification.

[0046] Furthermore, based on the subject classification table, the corpora in the second corpus are classified by subject, so that different subjects correspond to independent corpus sets. On this basis, a first preset number of corpora are extracted from the corpora of each subject, and based on the first preset number of corpora extracted from all subjects, evaluation questions in the field of aerospace technology that meet the quantity balance condition are generated. This ensures that the number of evaluation questions in the field of aerospace technology corresponding to each subject is constrained by the quantity balance condition, and is no longer dominated by the differences in the coverage scale and content density of general public corpora in different subjects in the field of aerospace technology, thus avoiding an uneven distribution of the number of evaluation questions in the field of aerospace technology across the subject dimensions.

[0047] Based on the above settings, the evaluation conclusions are less likely to be dominated by disciplines with a high proportion of evaluation questions in the aerospace technology field. This improves the coverage of the model's capabilities in the aerospace technology field, making the evaluation conclusions more comparable and representative in the aerospace technology field. This alleviates the defect that the uneven distribution of the number of evaluation questions in the aerospace technology field in related technologies can lead to biased evaluation conclusions.

[0048] In an optional embodiment, before step S104, the generation method further includes: for disciplines where the number of corresponding corpora is less than a first preset number, supplementing the corresponding corpora until the number of corresponding corpora reaches the first preset number. Specifically, the corresponding corpora are corpora that have been preliminarily determined to be related to the discipline. After supplementing the corresponding corpora in the first corpus, the filtering and classification processes in steps S101 to S103 can be used to determine whether the corpora actually belong to the discipline; corpora belonging to the discipline are retained, and corpora not belonging to the discipline are removed, until the total number of corpora for the discipline reaches the first preset number. This step can ensure that each discipline in the first corpus has sufficient corpora, avoiding an imbalance in the number of assessment questions generated due to insufficient corpora in some disciplines, thereby achieving a balanced distribution of assessment questions across different discipline dimensions.

[0049] The method for producing evaluation questions in the aerospace technology field for large language models in this embodiment supplements the corresponding corpus for disciplines with a corpus quantity less than a first preset quantity until the corpus quantity reaches the first preset quantity, thus providing a guarantee for each discipline to generate evaluation questions in the aerospace technology field using quantitative corpus.

[0050] In one alternative embodiment, such as Figure 2 As shown, after step S105, the generation method further includes steps S106 to S109: S106. Among the evaluation questions in the field of aerospace technology, select those that meet the selection criteria in terms of semantic redundancy, quality and difficulty as the preferred aerospace technology evaluation questions.

[0051] The preferred criteria are used to screen assessment questions in the aerospace technology field. These criteria characterize whether the assessment questions in the aerospace technology field meet the selection requirements in terms of semantic redundancy, quality, and difficulty. The preferred aerospace technology assessment questions are those selected from the existing aerospace technology assessment questions that meet the preferred criteria in terms of semantic redundancy, quality, and difficulty.

[0052] S107. If the number of evaluation questions in the field of aerospace technology does not meet the condition of quantity balance, supplementary data shall be obtained for subjects whose proportion of the number of evaluation questions in the corresponding field of aerospace technology is lower than the preset proportion threshold.

[0053] Specifically, the preset percentage threshold is a pre-defined threshold used to compare with the percentage of evaluation questions in the corresponding preferred aerospace technology field. When the percentage of evaluation questions in the corresponding preferred aerospace technology field is lower than this threshold, the subject is determined to be the subject for which supplementary corpus is obtained. The supplementary corpus is the corpus obtained for subjects where the percentage of evaluation questions in the corresponding preferred aerospace technology field is lower than the preset percentage threshold, when the number of evaluation questions in the preferred aerospace technology field does not meet the quantity balance condition. This supplementary corpus is preliminarily determined to be relevant to the subject.

[0054] S108. Based on the supplementary corpus, obtain supplementary evaluation questions in the field of aerospace technology that meet the optimal conditions in terms of semantic repetition, quality, and difficulty.

[0055] Specifically, the supplementary aerospace technology evaluation questions are derived from the supplementary corpus, and these questions meet the optimal criteria in terms of semantic redundancy, quality, and difficulty. The supplementary corpus is sequentially processed through relevance filtering and classification in steps S101-S103 to determine whether it actually belongs to the corresponding discipline. Corpus belonging to the relevant discipline is retained, while those not belonging to the relevant discipline are discarded. Subsequently, new aerospace technology evaluation questions are generated based on the selected supplementary corpus. These new evaluation questions, which meet the optimal criteria in terms of semantic redundancy, quality, and difficulty, are selected as the supplementary aerospace technology evaluation questions.

[0056] S109. Merge the supplementary aerospace technology evaluation questions with the selected aerospace technology evaluation questions to obtain new selected aerospace technology evaluation questions. Then, if the selected aerospace technology evaluation questions do not meet the quantity balance condition, obtain supplementary corpus for the disciplines whose corresponding selected aerospace technology evaluation questions have a quantity ratio lower than the preset ratio threshold.

[0057] Steps S106 to S109 involve filtering the obtained aerospace technology evaluation questions, retaining those that meet the selection criteria in terms of semantic redundancy, quality, and difficulty. If the selected aerospace technology evaluation questions do not meet the quantity balance condition, new selected aerospace technology evaluation questions are generated. Finally, a selection of aerospace technology evaluation questions with a quantity balance across disciplines is obtained. The aerospace technology evaluation question generation method for large language models described in this embodiment, through the iterative completion mechanism of supplementary corpus and supplementary evaluation questions, can maintain a long-term, stable balance in the quantity distribution of aerospace technology evaluation questions across disciplines, avoiding imbalances in the number of questions between disciplines. It possesses the ability to repeatedly adjust and maintain balance, ultimately achieving iterative completion and balanced distribution by discipline.

[0058] In one alternative embodiment, such as Figure 3 As shown, S106 includes steps S1061 to S1064: S1061. Perform semantic deduplication on evaluation questions in the field of aerospace technology.

[0059] Specifically, the embedding model can be used to sequentially deduplicate the aerospace technology evaluation questions within each discipline, and only one aerospace technology evaluation question with a semantic similarity greater than 0.9 can be retained.

[0060] S1062. Based on the deduplicated evaluation questions in the aerospace technology field, high-quality evaluation questions in the aerospace technology field that meet the preset quality conditions are obtained.

[0061] Specifically, the quality criteria may include at least one of the following: whether the question stem is complete and clear, whether the knowledge point is clearly defined, whether the answer is unique and correct, whether there is any ambiguity or missing conditions, whether the difficulty of the question matches the corresponding subject classification and grade level, and whether the question is presented in a standardized manner. High-quality assessment questions are those in the aerospace technology field that meet these quality criteria.

[0062] Large language models can be used to quality-label the deduplicated aerospace technology assessment questions. Suitable models include the GPT series, Claude series, Gemini series, Llama series, and Qwen series. Based on the quality labeling results, assessment questions under each subject category are divided into high-quality and low-quality categories, with high-quality questions being retained directly. Low-quality assessment questions are rewritten and then quality-labeled again: if the rewritten questions meet the quality criteria, they are retained and combined with the original high-quality questions to form the final high-quality aerospace technology assessment questions; if the rewritten questions are still low-quality, they are discarded.

[0063] S1063. High-quality aerospace technology evaluation questions are divided into high-difficulty aerospace technology evaluation questions and low-difficulty aerospace technology evaluation questions according to their difficulty.

[0064] High-difficulty aerospace technology assessment questions are those classified as high-difficulty aerospace technology assessment questions based on difficulty labeling results, while low-difficulty aerospace technology assessment questions are those classified as low-difficulty aerospace technology assessment questions based on difficulty labeling results, among the high-difficulty aerospace technology assessment questions.

[0065] Specifically, stable large language models such as GPT-5, Gemini series, Qwen3-235B-Think, DeepSeek-R1, and Kimi-K2-Instruct can be used to label the difficulty of high-quality aerospace technology evaluation questions. The labeling results are uniformly set to "high difficulty" or "low difficulty" with "difficulty" as the key field, so as to divide high-quality aerospace technology evaluation questions into two categories: high difficulty and low difficulty.

[0066] S1064. Low-difficulty aerospace technology evaluation questions that meet the preset accuracy conditions in both high-difficulty aerospace technology evaluation questions and sampling surveys are selected as preferred aerospace technology evaluation questions.

[0067] The preset accuracy rate conditions are pre-defined criteria used to determine whether the accuracy rate of a sample survey of low-difficulty aerospace technology assessment questions meets the retention requirements. Correct high-difficulty aerospace technology assessment questions refer to those questions that, after being categorized by subject and verified by experts in their respective fields, are accurately worded, have a high degree of matching between the question stem and the corresponding answer, and have no incorrect answers or mismatches between the question stem and the answer. Low-difficulty aerospace technology assessment questions whose sampling survey accuracy rate meets the preset accuracy rate conditions refer to those questions whose accuracy is verified through manual random sampling, and where all assessment questions within the sampling range meet the requirements of matching the question and answer, with no answer deviation, and whose sampling survey accuracy rate meets the preset conditions. Specifically, low-difficulty aerospace technology assessment questions are verified using manual random sampling; high-difficulty aerospace technology assessment questions are assigned to experts in their respective fields for verification, categorized by subject.

[0068] Specifically, since the evaluation questions in the field of aerospace technology are divided into two categories based on difficulty: high difficulty and low difficulty, in step S107, when determining whether the evaluation questions in the field of aerospace technology meet the quantity balance condition, it is necessary to judge the high difficulty aerospace technology evaluation questions and the low difficulty aerospace technology evaluation questions separately, rather than merging the two types of evaluation questions and judging them uniformly. Ultimately, a set of high difficulty aerospace technology evaluation questions A and a set of low difficulty aerospace technology evaluation questions B can be obtained, and the quantity distribution of the evaluation questions in the field of aerospace technology in both sets remains uniform across all disciplines. Figure 4 This showcases a high-difficulty evaluation set A in the field of aerospace technology. Figure 5 It showcases a set of evaluations in the field of low-difficulty aerospace technology, titled B.

[0069] The method for producing aerospace technology evaluation questions for large language models in this embodiment establishes a controllable mechanism for semantic redundancy, quality, and difficulty, thereby improving the credibility of the evaluation results. After generating aerospace technology evaluation questions, evaluation questions that meet the optimization criteria in terms of semantic redundancy, quality, and difficulty are selected as preferred aerospace technology evaluation questions. Specifically, the evaluation questions are first semantically deduplicated to obtain a deduplicated set of evaluation questions. Then, based on this set, high-quality aerospace technology evaluation questions that meet the quality criteria are selected. Subsequently, the high-quality aerospace technology evaluation questions are divided into two categories according to difficulty: high difficulty and low difficulty. Correct high-difficulty aerospace technology evaluation questions and low-difficulty aerospace technology evaluation questions whose accuracy rate meets the accuracy rate criteria in the sampling survey are selected as preferred aerospace technology evaluation questions, thus forming a high-difficulty aerospace technology evaluation set A and a low-difficulty aerospace technology evaluation set B. This effectively improves the stability of the evaluation question quality and difficulty classification and enhances the credibility of the evaluation results.

[0070] In an optional embodiment, step S105 includes: inputting a first preset number of corpora and prompt words extracted from each subject into a large language model to generate a second preset number of aerospace technology evaluation questions; the large language model is used to generate the same preset number of aerospace technology evaluation questions based on each corpus, and to summarize all aerospace technology evaluation questions for each subject to obtain the second preset number of aerospace technology evaluation questions.

[0071] The second preset quantity is the product of the number of disciplines and the first preset quantity. This value represents the total number of assessment questions in the aerospace technology field after summarizing all disciplines. Specifically, according to discipline classification, each discipline extracts the first preset quantity of corpus data from the second corpus based on the assessment question requirements. Each corpus can correspond to one document. Stable large language models such as GPT5, Gemini, qwen3-235b-think, deepseek-R1, and kimi-K2-instruct are used, explicitly specifying the corresponding discipline in the prompt words, and assessment question distillation is performed on each corpus. The following shows the prompt words for the disciplines of spacecraft control and navigation technology in the aerospace technology field: For example, your role and task could be: You are a rigorous interdisciplinary expert and an experienced test creator, proficient in generating three multiple-choice questions based on the following document content on aircraft control and navigation technologies.

[0072] The document content and question requirements are as follows: 1. Generate 3 multiple-choice questions, each containing a stem, 4 options, the correct answer, and an explanation; 2. The prompt should be 60-120 words long and based on document content related to aircraft control and navigation technology; 3. The options are concise and clear, less than 40 words; 4. There is only one correct answer.

[0073] Following the above method, the same preset number of aerospace technology evaluation questions are generated for each corpus entry in each subject, in batches according to discipline. This preset number can be selected as 3. After obtaining the evaluation questions for each subject based on the subject classification, all evaluation questions are summarized.

[0074] In one alternative embodiment, such as Figure 6 As shown, S101 includes steps S1011 to S1016: S1011. Based on the subject classification table and keyword table, obtain the TF-IDF vector of the keyword table.

[0075] The keyword list includes keywords.

[0076] The keyword list is a collection of keywords determined and summarized based on the disciplines in the subject classification table. The keyword list TF-IDF vector is a vector formed by sequentially arranging the TF-IDF values ​​of each keyword in the aerospace technology reference corpus according to the fixed order of the keyword list. Each element in the vector corresponds one-to-one with a keyword in the keyword list, and the order remains consistent. Specifically, the reference corpus can be selected from core books, with textbooks serving as the selection criteria for such core books.

[0077] S1012. Determine the TF-IDF values ​​of candidate corpora for keywords in the first corpus.

[0078] S1013. Sort the TF-IDF values ​​of all candidate corpora according to the order of the keywords in the keyword table to obtain the TF-IDF vector of the candidate corpora.

[0079] Specifically, the number of keywords is not limited; this embodiment selects 600. A single corpus in the first corpus can be a single document. Each document can be represented as a TF-IDF space vector, i.e., the TF-IDF value of the candidate corpus. Each element of this vector represents the TF-IDF value of the candidate corpus corresponding to the 600 keywords in that document. For example, if the keywords in the keyword list for the aerospace technology field are in the order of [aerospace, aircraft, system, ...], then document A, which is about the research content of composite materials for the space shuttle fuselage, might have a candidate corpus TF-IDF vector of [0.8, 0.9, 0.1, ...]; document B, which is about the research content of key technologies for aerospace measurement, control, and command systems, might have a candidate corpus TF-IDF vector of [0.8, 0.1, 0.9, ...].

[0080] S1014. Determine the similarity between the TF-IDF vectors of the candidate corpus and the TF-IDF vectors of the keyword list to obtain the first similarity value.

[0081] This similarity score represents the degree of similarity between the TF-IDF vectors of the candidate corpus and the TF-IDF vectors of the keyword list. Cosine similarity can be used. The first similarity value is a numerical value used to characterize the degree of similarity between the TF-IDF vectors of the candidate corpus and the TF-IDF vectors of the keyword list.

[0082] S1015. Determine the similarity between the TF-IDF vectors of each candidate corpus to obtain the second similarity value.

[0083] The first similarity is the degree of similarity between the TF-IDF vectors of each candidate corpus, and this similarity can be expressed as cosine similarity. The second similarity value is a numerical value obtained by determining the similarity between the TF-IDF vectors of each candidate corpus, and is used to characterize the degree of similarity between the TF-IDF vectors of each candidate corpus.

[0084] S1016. Based on the first similarity value and the second similarity value, determine whether the corpus of the first corpus meets the relevance condition to the field of aerospace technology.

[0085] In an optional embodiment, the MMR (Maximal Marginal Relevance) algorithm can be used to determine whether the corpus in the first corpus meets the relevance condition to the aerospace technology field based on a first similarity value and a second similarity value. During the MMR initialization phase, step S1016 includes: when first determining whether the corpus in the first corpus meets the relevance condition, determining that the corpus with the highest first similarity value in the first corpus meets the relevance condition. Therefore, the first corpus in the second corpus is the corpus in the first corpus whose corresponding candidate TF-IDF vector is most similar to the TF-IDF vector of the keyword list.

[0086] In the MMR iterative selection phase, step S1016 further includes: when it is not the first time to determine whether the corpus in the first corpus meets the relevance condition, for each corpus in the first corpus that is not determined to meet the relevance condition, determine the maximum value of the second similarity value between it and the corpus in the first corpus that has been determined to meet the relevance condition; determine that the corpus in the first corpus with the largest difference between the first similarity value and the maximum value meets the relevance condition; return to determining the maximum value of the second similarity value between each corpus in the first corpus that is not determined to meet the relevance condition and the corpus in the first corpus that has been determined to meet the relevance condition, until the number of corpus in the first corpus that has been determined to meet the relevance condition reaches a preset first quantity threshold.

[0087] For all documents in the first corpus that are not deemed not to meet the relevance criteria, their MMR scores are determined based on a first similarity value and a second similarity value. The document with the highest MMR score is then considered to meet the relevance criteria and added to the second corpus. Through iterative MMR scoring, documents that meet both relevance and diversity criteria and reach a first quantity threshold can be selected from the first corpus according to actual needs. Specifically, the MMR score is determined using the following formula:

[0088]

[0089] in, This represents the TF-IDF vector of the candidate corpus corresponding to the documents in the first corpus; This represents the TF-IDF vector of the candidate corpus corresponding to the documents in the second corpus; This represents the TF-IDF vector of the keyword list; This indicates the collection of documents that have been selected for the second corpus. Represents the cosine similarity function; It is a trade-off parameter between 0 and 1.

[0090] For documents in the first corpus, the first part of the formula The second part of the formula is used to measure the relevance of documents to the keyword list. This is used to measure the maximum similarity between a document and documents in the second corpus. Therefore, each new document selected through iterative MMR formula added to the second corpus maintains a high relevance to the keyword list while exhibiting low similarity to existing documents in the second corpus. (Balanced parameters) When the value approaches 1, it emphasizes correlation; when it approaches 0, it emphasizes diversity.

[0091] The method for generating evaluation questions in the aerospace technology field provided in this embodiment can improve the relevance determination and screening effect, thereby enhancing the data foundation reliability of the second corpus. This method performs relevance determination on the corpus in the first corpus based on a subject classification table. Corpus whose determination results meet the relevance conditions of the aerospace technology field is transferred to the second corpus, effectively avoiding interference from corpus unrelated to the aerospace technology field in subsequent subject classification. Furthermore, a first similarity value is obtained by calculating the similarity between the TF-IDF vector of the keyword table and the TF-IDF vector of the candidate corpus. Simultaneously, a second similarity value is obtained based on the similarity between the candidate corpus. These two similarity values ​​are combined to determine whether the corpus meets the relevance conditions. Corpus meeting the relevance conditions is iteratively screened based on the MMR algorithm. Once a first quantity threshold is reached, the corpus determined to meet the relevance conditions is input into a pre-trained first classification model to obtain relevance confidence. Corpus with relevance confidence below the confidence threshold is re-determined as not meeting the relevance conditions, thus ensuring the relevance of the corpus in the second corpus.

[0092] In one alternative embodiment, such as Figure 7 As shown, S101 also includes S1017~S1018: S1017. When the number of corpora in the first corpus that have been determined to meet the relevance conditions reaches the first quantity threshold, the corpora in the first corpus that have been determined to meet the relevance conditions are input into the pre-trained first classification model to obtain the relevance confidence.

[0093] The first classification model is used to output the relevance classification results between the corpus in the first corpus that has been determined to meet the relevance conditions and the field of aerospace technology.

[0094] S1018. Re-determine whether the corpus in the first corpus that has been judged to meet the relevance condition has a relevance confidence level lower than the preset confidence level threshold as not meeting the relevance condition.

[0095] The confidence threshold is a pre-set threshold used to compare the relevance confidence with the output of the first classification model. If the relevance confidence of a corpus that has been determined to meet the relevance condition in the first corpus is lower than the confidence threshold, then the corpus is reclassified as not meeting the relevance condition.

[0096] In the relevance filtering in step S101, keyword filtering and classifier filtering are used in combination. The two methods complement each other and can maximize the relevance of the second corpus to the field of aviation technology.

[0097] Specifically, after completing keyword filtering in S1011~S1016, further classifier filtering can be carried out: First, a 0 / 1 classifier is trained based on the BERT (Bidirectional Encoder Representations from Transformers) model under the Transformer framework, and then the trained 0 / 1 classifier is used to classify and filter the corpus in the second corpus.

[0098] The training process for the 0 / 1 classifier is as follows: First, a certain amount of data is selected from the second corpus as needed, and its relevance and subject are manually cross-labeled. Simultaneously, the reference data used in generating the TF-IDF vector of the keyword table in S1011 is labeled as relevant, and its corresponding subject is also labeled. A portion of all the labeled data is used as the training set, and the other portion as the validation set. Then, the classifier is trained and validated through programming. Specifically, an appropriate BERT model is selected, and the 0 / 1 classifier is trained using the training set. The trained classifier is then used to perform 0 / 1 classification on the validation set data, determining whether the data is "relevant" or "irrelevant." The classification effect is then statistically analyzed, specifically by recording the training process using a training loss curve. The classification effect is quantitatively measured using four metrics: accuracy, precision, recall, and F1 score. A confusion matrix diagram and a visualization of the clustering results based on Truncated Singular Value Decomposition (TSVD) are also plotted to visually demonstrate the final classification effect. If the loss value and all metrics reach the preset target values, the classifier training is complete; otherwise, iterative training continues. Specifically, the F1 score can be determined by the following formula:

[0099] in, Indicates accuracy. This indicates the recall rate.

[0100] Figure 8The loss curve of the training process is shown. Figure 9 The diagram shows the confusion matrix of the trained 0 / 1 classifier on the validation set. Figure 10 The TruncatedSVD Visualization plot shows the classifier's classification performance on the validation set. Figures 8-10 As shown, the trained 0 / 1 classifier achieved an accuracy of 0.8313, a precision of 0.8341, a recall of 0.7623, and an F1 score of 0.7966 on the validation set.

[0101] In one alternative embodiment, such as Figure 11 As shown, the keyword list in S1011 is obtained in the following way: S001. For each subject in the subject classification table, select the same preset number of reference corpora.

[0102] Specifically, domain experts can select 25 core books as reference corpora for each subject in the subject classification table.

[0103] S002. Determine the TF-IDF values ​​of the candidate words for each subject in the subject classification table in the corresponding reference corpus.

[0104] S003. Based on the TF-IDF values ​​of the candidate words, determine the same preset number of keywords for each subject in the subject classification table.

[0105] S004. Summarize all the keywords to obtain a keyword table.

[0106] Candidate words are words from the reference corpus corresponding to each discipline in the subject classification table. They are used to filter keywords based on their TF-IDF irrelevance values. Specifically, the Bag of Words model can be used to determine the candidate words for each discipline in the subject classification table and their corresponding TF-IDF values ​​in the reference corpus. A predetermined number of candidate words with the highest TF-IDF values ​​are selected as keywords for that discipline; this predetermined number can be set to 50. Table 1 contains a total of 12 disciplines, and based on the above method, a keyword table with a capacity of 600 is finally obtained.

[0107] In one alternative embodiment, such as Figure 12 As shown, S1011 includes steps S10111 to S10113: S10111. Merge all reference corpora to obtain the total reference corpus.

[0108] S10112. Determine the keyword TF-IDF value of each keyword in the total reference corpus.

[0109] S10113. Sort all the TF-IDF values ​​of the keywords according to the order of the keywords in the keyword table to obtain the TF-IDF vector of the keyword table.

[0110] The total reference corpus is the corpus obtained by merging all reference corpora. Specifically, 25 core books corresponding to each subject are merged to form the total reference corpus; then, the TF-IDF value of each keyword in the total reference corpus is calculated, and the resulting TF-IDF values ​​are sorted according to the order of the keywords in the keyword table, finally yielding a 600-dimensional keyword table TF-IDF vector. The TF-IDF value can be determined by the following formula:

[0111]

[0112]

[0113] in, Indicates the first The first in the corpus The term frequency-inverse document frequency value of each keyword; Indicates the first The first in the corpus The word frequency of each keyword; Indicates the first Inverse document frequency of each keyword; Indicates the first The keyword in the first The number of times it appears in the corpus; Indicates the first Total number of words in the corpus; This indicates the total number of words in the total reference corpus; Indicates containing the first The number of keywords in the corpus. When the total number of core books in the total reference corpus is 300, .

[0114] The following are the top 10 dimensions of the TF-IDF vector of the keyword table in the field of aerospace technology, adapted according to the above steps. The table below indicates the keywords corresponding to the TF-IDF values ​​and the disciplines corresponding to the keywords.

[0115] Table 2: Keyword List in Aerospace Technology Field - Top 10 Dimensions of TF-IDF Vector

[0116] The TF-IDF bag-of-words model differs significantly from the TF bag-of-words model. The TF bag-of-words model only counts the frequency of words in the corpus and cannot distinguish between general terms and domain-specific terms. The TF-IDF bag-of-words model, however, considers both term frequency and inverse document frequency, effectively distinguishing the domain specificity of words. Keywords with high TF-IDF values ​​indicate that the keyword appears frequently in the current corpus and is relatively rare in the overall reference corpus. Compared to the TF bag-of-words model, the TF-IDF bag-of-words model can directly filter out highly frequent non-domain-related general terms, such as "you," "I," "is," and "not."

[0117] In an optional embodiment, step S103 includes: inputting the corpus of the second corpus into a pre-trained second classification model to obtain a subject classification result; the second classification model is used to determine the specific subject in the subject classification table to which the corpus of the second corpus belongs.

[0118] The subject classification result is the specific subject to which the corpus belongs after the second classification model identifies the subject of the corpus. Specifically, the number of classifications in the classification model is consistent with the number of subjects in the aerospace technology subject classification table. Since Table 1 contains 12 subjects, a suitable BERT model can be selected to train a twelve-subject classifier, and the trained twelve-subject classifier can be used to complete the subject classification of documents in the second corpus.

[0119] Based on the same inventive concept as the aforementioned method for generating assessment questions in the field of aerospace technology, this application also provides an apparatus for generating assessment questions in the field of aerospace technology, such as... Figure 13 As shown, the aerospace technology evaluation question generation device 2 includes: a first judgment module 21, a second corpus generation module 22, a classification module 23, a corpus extraction module 24, and an aerospace technology evaluation question generation module 25.

[0120] The first judgment module 21 is used to determine whether the corpus in the first corpus meets the relevance condition to the aerospace technology field based on the subject classification table. The second corpus generation module 22 is used to remove the corpus whose judgment result meets the condition from the first corpus and put it into another corpus to obtain the second corpus. The classification module 23 is used to classify the corpus in the second corpus according to the subject classification table. The corpus extraction module 24 is used to extract a first preset number of corpus from the corpus of each subject. The aerospace technology field assessment question generation module 25 is used to generate aerospace technology field assessment questions that meet the quantity balance condition based on the first preset number of corpus extracted from all subjects.

[0121] Each module of the aforementioned aerospace technology evaluation question generation device corresponds to the steps of the aforementioned aerospace technology evaluation question generation method. For details on the implementation process of the functions and roles of each module in the aforementioned aerospace technology evaluation question generation device, please refer to the implementation process of the corresponding steps in the aforementioned aerospace technology evaluation question generation method. The same technical effect can be achieved, and will not be repeated here.

[0122] Figure 14 The diagram shown is a structural schematic of the electronic device 30 provided in an embodiment of this application.

[0123] like Figure 14 As shown, the electronic device 30 includes one or more processors 31 for implementing the above-described method for generating evaluation questions in the field of aerospace technology.

[0124] In some embodiments, the electronic device 30 may include a storage medium 39. For example, a computer-readable storage medium may store a program that can be invoked by a processor 31, and may include a non-volatile storage medium. In some embodiments, the electronic device 30 may include memory 38 and an interface 37. In some embodiments, the electronic device 30 may also include other hardware depending on the specific application.

[0125] The computer-readable storage medium of this application embodiment stores a program that, when executed by processor 31, is used to implement the method for generating evaluation questions in the field of aerospace technology as described above.

[0126] This application provides a computer program product, including a computer program / instructions, which, when executed by a processor, implements a method for generating evaluation questions in the aerospace technology field as described above.

[0127] This application also provides a computer program stored in a computer-readable storage medium, for example... Figure 14 The storage medium 39, and when the processor executes the computer program, it causes the processor 31 to execute the method for generating evaluation questions in the field of aerospace technology described above.

[0128] This application may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing program code. Computer-readable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented using any method or technology. Information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media include, but are not limited to: phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0129] The aforementioned electronic device can execute the method for generating assessment questions in the aerospace technology field provided in the embodiments herein. The aforementioned electronic device may include the aforementioned apparatus for generating assessment questions in the aerospace technology field, such as one or more of a processor, a controller, and a PC (Personal Computer) terminal device. The server terminal device and the PC terminal device may include, but are not limited to, a server, a desktop computer, a tablet computer, or a laptop computer.

[0130] It should be noted that the technical solutions or features described in the above embodiments can be combined or supplemented with each other without conflict. The scope of protection of this application is not limited to the precise structures described in the above embodiments and shown in the accompanying drawings; all modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for generating assessment questions in the aerospace technology field for large language models, characterized in that, The generation method includes: Based on the subject classification table of aerospace technology, determine whether the corpus of the first corpus meets the relevance condition to the aerospace technology field. Remove the corpus whose judgment result is satisfied from the first corpus and put it into another corpus to obtain the second corpus; According to the subject classification table, the corpus of the second corpus is classified by subject. Extract a first predetermined amount of data from the corpus of each subject. Based on the first preset amount of corpus extracted from all disciplines, generate evaluation questions in the field of aerospace technology that meet the quantity balance condition.

2. The method for generating evaluation questions in the aerospace technology field for large language models as described in claim 1, characterized in that, Before extracting a first preset number of corpora from the corpus of each subject, the generation method further includes: For subjects where the quantity of corresponding corpus is less than the first preset quantity, supplement the corresponding corpus until the quantity of corresponding corpus reaches the first preset quantity.

3. The method for generating assessment questions in the aerospace technology field for large language models as described in claim 1, characterized in that, After generating the aerospace technology evaluation questions that satisfy the quantity balance condition, the generation method further includes: Among the evaluation questions in the field of aerospace technology, those that meet the optimization criteria in terms of semantic redundancy, quality, and difficulty are selected as the preferred evaluation questions in the field of aerospace technology. If the evaluation questions in the preferred aerospace technology field do not meet the quantity balance condition, supplementary corpus is obtained for subjects whose proportion of the corresponding evaluation questions in the preferred aerospace technology field is lower than a preset proportion threshold. Based on the supplementary corpus, supplementary aerospace technology evaluation questions are obtained that satisfy the preferred conditions in terms of semantic redundancy, quality, and difficulty. The supplementary aerospace technology evaluation questions are merged with the preferred aerospace technology evaluation questions to obtain new preferred aerospace technology evaluation questions. Then, when the preferred aerospace technology evaluation questions do not meet the quantity balance condition, supplementary corpus is obtained for the disciplines whose proportion of the corresponding preferred aerospace technology evaluation questions is lower than a preset proportion threshold.

4. The method for generating assessment questions in the aerospace technology field for large language models as described in claim 3, characterized in that, Among the evaluation questions in the aerospace technology field, those that meet the preferred criteria in terms of semantic redundancy, quality, and difficulty are selected as preferred aerospace technology evaluation questions, including: Semantic deduplication is performed on the evaluation questions in the aerospace technology field. Based on the deduplicated aerospace technology evaluation questions, high-quality aerospace technology evaluation questions that meet the preset quality conditions are obtained. The high-quality aerospace technology assessment questions are divided into high-difficulty aerospace technology assessment questions and low-difficulty aerospace technology assessment questions according to their difficulty. The preferred aerospace technology evaluation questions are those with high difficulty and a sampling accuracy that meet the preset accuracy conditions.

5. The method for generating evaluation questions in the aerospace technology field for large language models as described in claim 1, characterized in that, The process of generating assessment questions in the aerospace technology field that satisfy the quantity balance condition based on the first preset amount of corpus extracted from all disciplines includes: The first preset number of corpora and prompt words extracted from each subject are input into a large language model to generate a second preset number of aerospace technology evaluation questions. The large language model is used to generate the same preset number of aerospace technology evaluation questions based on each corpus, and to summarize all aerospace technology evaluation questions for each subject to obtain the second preset number of aerospace technology evaluation questions.

6. The method for generating assessment questions in the aerospace technology field for large language models as described in claim 1, characterized in that, The step of determining whether the corpus in the first corpus meets the relevance condition to the aerospace technology field based on the subject classification table of the aerospace technology field includes: Based on the subject classification table and the keyword table, the TF-IDF vector of the keyword table is obtained; the keyword table includes keywords; Determine the TF-IDF values ​​of the keyword in the candidate corpus of the first corpus; According to the order of the keywords in the keyword list, the TF-IDF values ​​of all candidate corpora are sorted to obtain the TF-IDF vector of the candidate corpora; Determine the similarity between the TF-IDF vectors of the candidate corpus and the TF-IDF vectors of the keyword list to obtain a first similarity value; Determine the similarity between the TF-IDF vectors of each candidate corpus to obtain the second similarity value; Based on the first similarity value and the second similarity value, it is determined whether the corpus of the first corpus meets the relevance condition to the aerospace technology field.

7. The method for generating assessment questions in the aerospace technology field for large language models as described in claim 6, characterized in that, The step of determining whether the corpus of the first corpus meets the relevance condition to the aerospace technology field based on the first similarity value and the second similarity value includes: When determining for the first time whether the corpus in the first corpus meets the relevance condition, it is determined that the corpus with the largest first similarity value in the first corpus meets the relevance condition.

8. The method for generating evaluation questions in the aerospace technology field for large language models as described in claim 7, characterized in that, The step of determining whether the corpus of the first corpus meets the relevance condition to the aerospace technology field based on the first similarity value and the second similarity value further includes: When it is not the first time to determine whether the corpus in the first corpus meets the relevance condition, for each corpus in the first corpus that is not determined to meet the relevance condition, determine the maximum value of the second similarity value between it and the corpus in the first corpus that has been determined to meet the relevance condition. The corpus with the largest difference between the first similarity value and the maximum value in the first corpus is determined to satisfy the relevance condition; Returning to each corpus in the first corpus that is not determined to meet the relevance condition, the maximum value of the second similarity value between it and the corpus in the first corpus that has been determined to meet the relevance condition is determined, until the number of corpora in the first corpus that have been determined to meet the relevance condition reaches a preset first quantity threshold.

9. The method for generating assessment questions in the aerospace technology field for large language models as described in claim 8, characterized in that, The step of determining whether the corpus in the first corpus meets the relevance condition to the aerospace technology field based on the subject classification table of aerospace technology further includes: When the number of corpora in the first corpus that have been determined to meet the relevance conditions reaches the first threshold, the corpora in the first corpus that have been determined to meet the relevance conditions are input into a pre-trained first classification model to obtain the relevance confidence; the first classification model is used to output the relevance classification result between the corpora in the first corpus that have been determined to meet the relevance conditions and the aerospace technology field; The corpora in the first corpus that were previously determined to meet the relevance condition but whose relevance confidence score is lower than a preset confidence score threshold are re-determined to not meet the relevance condition.

10. The method for generating evaluation questions in the aerospace technology field for large language models as described in claim 6, characterized in that, The keyword list was obtained in the following way: For each subject in the subject classification table, the same preset number of reference corpora are selected; Determine the TF-IDF values ​​of the candidate words for each subject in the subject classification table in the corresponding reference corpus; Based on the TF-IDF values ​​of the candidate words, the same preset number of keywords are determined from the candidate words of each subject in the subject classification table; All keywords are summarized to obtain the keyword table.

11. The method for generating evaluation questions in the aerospace technology field for large language models as described in claim 10, characterized in that, The step of obtaining the TF-IDF vector of the keyword table based on the subject classification table and the keyword table includes: All reference corpora are merged to obtain the total reference corpus; Determine the keyword TF-IDF value for each keyword in the total reference corpus; According to the order of the keywords in the keyword table, sort all the keyword TF-IDF values ​​to obtain the keyword table TF-IDF vector.

12. The method for generating assessment questions in the aerospace technology field for large language models as described in claim 1, characterized in that, The step of classifying the corpus of the second corpus by subject according to the subject classification table includes: The corpus data of the second corpus is input into a pre-trained second classification model to obtain subject classification results; the second classification model is used to determine the specific subject in the subject classification table to which the corpus data of the second corpus belongs.

13. An electronic device, characterized in that, It includes one or more processors for implementing the method for generating evaluation questions in the aerospace technology field for large language models as described in any one of claims 1 to 12.

14. A computer-readable storage medium, characterized in that, It stores a program that, when executed by a processor, implements the method for generating evaluation questions in the aerospace technology field for large language models as described in any one of claims 1 to 12.