Test question generation method and system based on retrieval enhancement generation
By constructing a dual-track parallel hybrid knowledge base based on retrieval enhancement and combining a large language model with multi-dimensional quality assessment, the problems of low efficiency and insufficient accuracy in the psychological quality assessment of civil aviation flight attendants have been solved, achieving efficient and professional automated generation and assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies for assessing the psychological qualities of civil aviation flight attendants suffer from inefficiency, strong subjectivity, insufficient professional accuracy, and difficulty in generating content that meets actual needs. In particular, they are prone to generating "illusions" in situational judgment tests, making it impossible to comprehensively and objectively assess the competence of high-risk positions.
We employ a retrieval-enhanced generation (RAG) approach to construct a dual-track parallel hybrid knowledge base. This base combines a structured database with a vector knowledge base and automatically extracts structured competency features and key events using a large language model to generate contextual judgment test questions. Through multi-dimensional quality assessment and optimization, we ensure the professionalism and accuracy of the generated content.
It has achieved the automation, standardization, and efficient generation of psychological quality assessments for civil aviation flight attendants. The generated questions are closely related to the actual job requirements, which improves the efficiency and quality of the assessments and ensures the professional accuracy and credibility of the generated content.
Smart Images

Figure CN121833705A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of cross-fusion of artificial intelligence and psychological assessment, and in particular to large language model application technology based on retrieval augmented generation (RAG), text structured extraction and controllable text generation technology in natural language processing, and psychological competency model construction and situational judgment test (SJT) automation preparation technology, which is suitable for psychological quality assessment, talent selection, training effect diagnosis and human resource decision support of high-risk and high-service requirement posts such as civil aviation cabin attendants. BACKGROUND
[0002] At present, the psychological quality and working ability (such as emotion management, team cooperation, emergency decision-making, etc.) of cabin attendants selected or evaluated by airlines mainly rely on two ways: traditional "manual experience driven" process and emerging but imperfect "basic AI assisted" method.
[0003] Traditional manual development process and its problems: The linear segmented development process with manual as the core is the current mainstream way. The process is as follows: Manual competency modeling and interview text coding: a team of psychologists manually defines the competency model through limited literature review and interviews. Then, the interview texts of a small number of cabin attendants are manually coded, that is, the coders read the texts word by word and sentence by sentence, subjectively judge and mark the competency behavior indicators embodied in them according to the coding manual. This process has the problems of difficult to unify the coding standards, time-consuming and labor-intensive, long cycle, and serious dependence on the personal experience and subjective judgment of the coders, which easily leads to inconsistent coding results, low reliability, and difficulty in covering all key information in a large amount of interview data.
[0004] Manual SJT item development: Based on the limited situational materials, the item stem, options, scoring criteria and reasons are manually written by the assessment experts. The core problems faced in this stage are: (1) The subjective limitations of experts are prominent: the quality of the items is highly dependent on the personal knowledge and experience of the item writers, and it is difficult to ensure that the content of the items can fully and objectively reflect the diversity of real work situations. (2) Low efficiency of item generation: from situational collection and screening to item writing and review, all rely on manual work, resulting in high development cost and long cycle of individual items, which may take several months or even longer. (3) Insufficient coverage and precision of items: due to limited manpower and cost, the size of the item bank is limited, which may lead to repeated use of items in selection, reducing safety; at the same time, the complexity and authenticity of the item situation are also difficult to fully guarantee. (4) Long iteration and update cycle: when industry standards or work situations change, the review and update of the entire item bank also requires a lot of manual work, which may lead to a lag between the assessment content and the actual demand. (5) Leakage of items and loss of discrimination: a limited number of manually prepared items may easily lead to content leakage after administration, affecting the discrimination of the items and the fairness of the test.
[0005] Limitations and needs of emerging AI-assisted methods: To overcome the drawbacks of the above-mentioned manual process, the industry has begun to explore the use of AI large model technology for improvement. However, the current method of relying solely on basic large language models has exposed a series of new problems when applied to civil aviation assessment, which is highly professional and has high stakes: (1) "Illusion" problem difficult to solve: basic large models are prone to generate seemingly reasonable but actually incorrect or inaccurate information (i.e. "illusion") when generating professional content. In civil aviation competency assessment, any deviation in professional knowledge may lead to ineffective items, even misleading selection decisions, with high risk. (2) Insufficient application in professional recruitment selection: existing AI-based assessment generation research focuses mainly on general knowledge or cognitive ability tests, with little research on automated item development for specific industries such as civil aviation and non-cognitive psychological abilities (such as communication and cooperation, decision-making under stress, etc.). (3) Measurement content is one-sided: even with AI, existing technology mainly focuses on generating items that test cognitive abilities, while there is a lack of effective means for automated and precise measurement of deep psychological traits such as emotional attitudes, values, interpersonal skills, which are crucial for civil aviation cabin crew.
[0006] In summary, the existing technologies have the following core problems: (1) Traditional manual methods are professional but inefficient, subjective and difficult to scale; (2) AI methods driven by basic large language models are efficient but lack professional accuracy, are prone to "illusions" and cannot ensure that the generated content conforms to the real norms and practices in the civil aviation field; (3) There is a lack of technical mechanisms to effectively integrate professional domain knowledge into the AI generation process, which makes it difficult for the generated assessment questions to meet the actual needs of high-stakes job selection in terms of contextual authenticity, content validity and measurement depth. Summary of the Invention
[0007] The following provides a brief overview of one or more aspects to offer a basic understanding of them. This overview is not an exhaustive summary of all conceived aspects, nor is it intended to identify key or decisive elements of all aspects, nor to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form to prepare for the more detailed descriptions that follow.
[0008] The purpose of this invention is to solve the above-mentioned problems by providing a test question generation method and system based on retrieval-enhanced generation. This method combines the precision of professional fields with the efficiency of AI technology. It can not only use AI to achieve automatic and standardized coding of interview texts and efficient and batch generation of situational judgment test questions, but also ensure that the generated content is rooted in the professional field through retrieval-enhanced generation. This fundamentally solves the "illusion" problem and enables a comprehensive, objective, efficient and reliable automated assessment of the test subjects' abilities.
[0009] The technical solution of this invention is as follows: This invention discloses a test question generation method based on retrieval enhancement, the method comprising: Step 1: Perform data parsing and vectorization processing on the raw data to build a vector knowledge base and a structured database; Step 2: Using the vector knowledge base and structured database built in Step 1, a large language model is driven by retrieval enhancement generation technology to automatically extract structured competency features and key events from the practical cases to be processed, and store them in the competency feature and key event database. Step 3: Based on retrieval enhancement generation, use vector knowledge base, structured database, and competency feature and key event database to automatically generate situational judgment test questions.
[0010] According to one embodiment of the test question generation method based on retrieval enhancement according to the present invention, the original data is obtained from multiple sources including theoretical literature, practical case database and historical question database.
[0011] According to an embodiment of the test question generation method based on retrieval enhancement according to the present invention, step one further includes: Step 1-1: Obtain the raw data; Steps 1-2: Based on the characteristics of different data sources, differentiated data processing and data entry are performed to generate a hybrid knowledge base consisting of a structured database and a vector knowledge base. The differentiation includes: extracting key information from unstructured interview texts in the practical case library and storing the extracted key events and competency features in the structured database; performing structured parsing on the historical question bank and storing the parsed standardized situational judgment test questions in the structured database; and vectorizing theoretical literature and storing the generated semantic vectors in the vector knowledge base.
[0012] According to an embodiment of the test question generation method based on retrieval enhancement according to the present invention, step two further includes: Step 2-1: For the practical cases to be processed, retrieve relevant psychological theories and practical cases from the vector knowledge base through retrieval enhancement, and initially screen and identify the original text fragments containing specific behaviors and situations; Step 2-2: The large language model takes the original text fragments corresponding to the key events retrieved in Step 2-1 and the predefined competency rating criteria as input to the model. It guides the large language model to perform semantic analysis and reasoning through preset analysis prompts to obtain competency features and their levels. Steps 2-3: Quantitative evaluation based on coding reliability coefficients, transforming expert experience into evaluation indicators, and generating a reliable, structured competency feature and key event library.
[0013] According to an embodiment of the test question generation method based on retrieval enhancement according to the present invention, step three further includes: Step 3-1: The Big Language Application Development Platform takes the target competency features as input, performs a search in a multi-source database to obtain reference material fragments related to the target competency features, and uses the reference material fragments as contextual information to construct prompt words for generating test questions; Step 3-2: Based on the input content sent in Step 3-1, the large language model performs tasks according to prompt words to generate test questions. Multiple test questions form a candidate situation judgment test question set.
[0014] According to an embodiment of the test question generation method based on retrieval enhancement according to the present invention, the method further includes: Step 4: Evaluate and optimize the quality of the situational judgment test questions generated in Step 3.
[0015] This invention also discloses a test question generation system based on retrieval enhancement, the system comprising: The multi-source knowledge base construction module performs data parsing and vectorization processing on the raw data, thereby constructing a vector knowledge base and a structured database; The competency feature intelligent extraction module utilizes a vector knowledge base and structured database built by a multi-source knowledge base construction module. It employs retrieval enhancement generation technology to drive a large language model, automatically extracting structured competency features and key events from the practical cases to be processed, and storing them in a competency feature and key event database. The automatic question generation module, based on retrieval-enhanced generation, utilizes a vector knowledge base, a structured database, and a competency feature and key event database to automatically generate situational judgment test questions.
[0016] According to an embodiment of the test question generation system based on retrieval enhancement according to the present invention, the system further includes: The question quality assessment and closed-loop optimization module assesses and optimizes the quality of the situational judgment test questions generated by the automatic question generation module.
[0017] The present invention also discloses an electronic device, the electronic device including a controller, the controller including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When the processor executes the program stored in memory, it implements the steps of the question generation method based on retrieval enhancement as described above.
[0018] The present invention also discloses a computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the steps of the test question generation method based on retrieval enhancement as described above.
[0019] Compared with the prior art, the present invention has the following advantages: First, this invention constructs a "dual-track parallel hybrid knowledge base" and applies it to the field of psychological testing. This knowledge base consists of two parts: a structured database storing key events and competency features automatically extracted from flight attendant interview texts using a large language model, as well as standardized Historical Contextual Judgment Test (SJT) questions; and a vectorized knowledge base storing vector representations of relevant documents such as psychology and civil aviation regulations after semantic segmentation and embedding, supporting deep semantic retrieval. By transforming unstructured interview content into structured data and using it in conjunction with vectorized theoretical knowledge, this dual-track parallel hybrid knowledge base achieves an organic combination of precise querying and semantic understanding. This technical feature, to a certain extent, solves the problem of scattered professional knowledge and difficulty in systematic utilization in traditional methods, providing a unified, reliable, and high-fidelity knowledge source for subsequent automated coding and question generation, significantly improving the professional accuracy and contextual relevance of the generated content.
[0020] Secondly, the test item generation method of this invention is based on retrieval-enhanced generation (RAG) technology. RAG organically links item generation, multi-dimensional quality assessment (including automatic rule filtering, expert review, and psychometric analysis), and dynamic optimization of preceding generation stages, forming a feedback-driven closed-loop system. Large language models are prone to creating "illusions," fabricating seemingly reasonable but actually erroneous or inaccurate professional content. Existing mainstream technologies mainly rely on knowledge stored within the model, whose accuracy and timeliness cannot be guaranteed. RAG technology, by retrieving information in real time from external authoritative knowledge bases (such as civil aviation regulations, operation manuals, and competency model documents) and providing references for the model, enhances the professionalism and credibility of the generated content, representing the cutting-edge direction of technological development in this field. The method does not merely modify individual unqualified items locally, but rather transforms expert review opinions and psychometric indicators (such as discrimination, criterion-related validity, etc.) into optimization instructions for feature extraction prompts or item generation prompts, thus influencing the workflow in reverse. This technical feature enables the entire assessment system to have adaptive evolution capabilities, continuously iterating to improve the quality of the question bank, thereby significantly enhancing the long-term effectiveness, stability, and maintainability of the assessment tool.
[0021] Third, this invention establishes a data-driven empirical construction and optimization process for competency models. This process begins with the construction of a theoretical model (12 dimensions) based on literature and interviews, followed by the development of an initial question bank (10 dimensions / 152 questions) incorporating expert opinions. Then, through exploratory factor analysis, reliability and validity testing, and question selection using large-sample testing data from multiple (e.g., 818) active flight attendants, a more stable and accurate 5-dimensional psychological competency model and corresponding high-quality SJT questions (e.g., 59 questions) are established, and its predictive validity for actual job performance is verified. This technical feature transforms the establishment of competency models from relying on expert subjective experience to a scientific verification process based on large-scale empirical data, overcoming the subjectivity and instability of traditional modeling methods. It ensures that the constructed assessment system possesses excellent psychometric characteristics (such as high internal consistency reliability and strong criterion-related validity) and effective predictive ability for actual job performance.
[0022] In summary, this invention achieves efficient, accurate, and automated generation of interview texts for psychological competence assessment of test subjects, such as flight attendants, through fusion retrieval-enhanced generation (RAG), structured information extraction from large language models, construction of hybrid knowledge bases, and closed-loop optimization mechanisms. This enables the rapid generation of diverse and highly realistic test questions tailored to the specific job requirements. Simultaneously, it constructs a standardized and precise automatic assessment system, thereby improving both the efficiency of assessing job psychological competence and the quality of job selection.
[0023] This invention is not only applicable to the selection, training, and assessment of cabin crew in China's civil aviation transportation, but its technical framework also has good versatility and can be transferred to other fields with high requirements for job competency, such as the selection of personnel in safety-critical positions in transportation industries such as high-speed rail and public transportation, the assessment of soft skills of medical and health personnel in the healthcare industry, and the assessment of leadership potential for enterprise managers, providing data-driven and intelligent talent assessment solutions for various industries. Attached Figure Description
[0024] The above-described features and advantages of the present invention will be better understood after reading the following detailed description of embodiments of the present disclosure in conjunction with the accompanying drawings. In the drawings, components are not necessarily drawn to scale, and components having similar related characteristics or features may have the same or similar reference numerals.
[0025] Figure 1 The flowchart of an embodiment of the test question generation method based on retrieval enhancement of the present invention is shown.
[0026] Figure 2 It shows Figure 1 A detailed flowchart of step one in the method shown.
[0027] Figure 3 It shows Figure 1 The flowchart shown is for the application scenario of step one in the method, which targets the psychological competence characteristics of civil aviation flight attendants.
[0028] Figure 4 It shows Figure 1 The detailed flowchart of step two in the method shown.
[0029] Figure 5 It shows Figure 1 A detailed flowchart of step three in the method shown.
[0030] Figure 6 It shows Figure 1 A detailed flowchart of step four in the method shown.
[0031] Figure 7 A schematic diagram of an embodiment of the test question generation system based on retrieval enhancement of the present invention is shown.
[0032] Figure 8 A schematic diagram of the structure of an electronic device according to the present invention is shown. Detailed Implementation
[0033] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. It should be noted that the aspects described below with reference to the accompanying drawings and specific embodiments are merely exemplary and should not be construed as limiting the scope of protection of the present invention in any way.
[0034] Figure 1 The overall flow of an embodiment of the test question generation method based on retrieval enhancement of the present invention is shown. Please refer to [link to relevant documentation]. Figure 1 The implementation steps of the method in this embodiment are detailed below. In the following detailed description of the implementation steps, some examples will be used to illustrate the testing of civil aviation flight attendants. However, this scenario is only for illustrative purposes and is not intended to limit the scope of the invention. The present invention can be extended to other job competency testing and selection scenarios.
[0035] Step 1: Perform data parsing and vectorization on the raw data to build a vector knowledge base and a structured database.
[0036] The raw data is obtained from three categories: theoretical literature, practical case studies, and historical question banks. The theoretical knowledge base includes literature and monographs, the practical case studies include transcripts of interviews with frontline personnel, and the historical question bank includes historical SJT (Situation Judgment Test) questions. The vector knowledge base supports semantic retrieval, while the structured database is used for precise queries.
[0037] refer to Figure 2 As shown, step one further includes the following processing steps.
[0038] Step 1-1: Obtain the raw data.
[0039] The raw data is divided into three categories: theoretical literature, practical case database, and historical question database. The specific data sources are as follows.
[0040] (1) Case studies based on interview texts: For example, one-on-one in-depth interviews with multiple industry professionals (e.g., 30 front-line flight attendants), each interview is a verbatim transcript of 10,000 to 20,000 words (e.g., Word or TXT format), rich in key events and behaviors in real work situations.
[0041] (2) Historical question bank: Situational Judgment Test (SJT) questions that have been verified through public or internal experiments are usually stored in a structured Excel spreadsheet format, which includes the question stem, options, answers and explanations.
[0042] (3) Relevant theoretical literature: professional works and academic papers covering fields such as psychology and organizational behavior (mainly in PDF format) to provide theoretical support and professional knowledge background.
[0043] Steps 1-2: Based on the characteristics of different data sources, differentiated data processing and data entry are performed to generate a hybrid knowledge base consisting of a structured database and a vector knowledge base. Differentiation includes: extracting key information from unstructured interview texts in the practical case study database and storing the extracted key events and competency features in the structured database; performing structured parsing on the historical question bank and storing the parsed standardized situational judgment test questions in the structured database; and vectorizing theoretical literature and storing the generated semantic vectors in the vector knowledge base.
[0044] (1) For interview texts (unstructured data processing): An automated workflow is built using a Large Language Model (LLM) application development platform such as Dify. This workflow leverages the powerful natural language understanding and information extraction capabilities of LLM to automatically process the interview transcripts and extract structured key events and related competency features. For example, the LLM identifies the specific process by which flight attendants handle sudden passenger conflicts (key events) and labels them with "adaptability" or "customer orientation" (competency features). The extracted structured data is directly inserted into a relational database (such as PostgreSQL). In this way, massive amounts of unstructured interview content are directly transformed into structured knowledge that can be accurately queried, statistically analyzed, and processed.
[0045] (2) For existing question banks (structured data processing): Construct a workflow to automatically import questions in Excel format into the database. This workflow is responsible for parsing the Excel file and mapping each row of data to the corresponding fields in a relational database (such as PostgreSQL), such as: question ID, question stem text, option A, option B, option C, option D, correct answer, scoring reason, etc. The entire process is automated, ensuring that the question bank data can be quickly and accurately imported into a relational database such as PostgreSQL in batches, forming a standardized and easy-to-manage structured question bank.
[0046] (3) For relevant theoretical literature (vectorization processing): A dedicated vector knowledge base is constructed using a Retrieval-Augmented Generation (RAG) development platform (such as the RAGFlow platform) that supports semantic chunking and vectorization processing. RAGFlow first performs semantic-based intelligent chunking on the uploaded literature (PDF, Word, etc.) to obtain semantic-based text fragments. This method avoids the semantic fragmentation that may be caused by simple word count cutting and ensures the integrity of the fragments. Subsequently, a text vectorization embedding model (such as BGE, M3E, etc.) is called to convert each text fragment into a high-dimensional floating-point vector. These vectors and their corresponding original text fragments are uniformly stored in a vector database (such as Milvus, Chroma), ultimately forming a vector knowledge base that focuses on academic literature and supports deep semantic retrieval.
[0047] The structured relational database (such as PostgreSQL) formed by the above (1) and (2) processes centrally stores key events and competency characteristics extracted from interviews, as well as standardized situational judgment test items, supporting precise SQL queries, correlation analysis and statistics.
[0048] The vector knowledge base formed by the above (3) process specifically stores the semantic vectors of academic documents, supports fuzzy query and related knowledge discovery through natural language, and provides a theoretical basis for understanding problems and generating content.
[0049] This dual-track parallel (structured database and vector knowledge base) hybrid knowledge base system combines the structured extraction of real work experience with the semantic understanding of academic theories, providing solid, multimodal data support for subsequent intelligent applications such as question generation and context analysis.
[0050] Taking the psychological competence characteristics of civil aviation flight attendants as an example, Figure 3 This illustrates the preparatory stage of step one. (Reference) Figure 3The initial preparation stage is divided into four stages, which are detailed below.
[0051] Phase 1: Preliminary Research and Interviews We systematically retrieved and reviewed domestic and international academic literature, industry standards, and policy documents on topics such as "civil aviation flight attendant competence" and "cabin crew competence." Based on the literature review and combined with organizational behavior and psychology theories, we predefined a set of multiple (e.g., 45) specific competency characteristics. For each competency characteristic, we developed clear and observable behavioral rating standards (usually divided into 1-3 levels, representing "incompetent," "basically competent," and "excellent competence," respectively).
[0052] Using the critical incident method, we conducted in-depth structured or semi-structured interviews with multiple (e.g., 30) frontline flight attendants (covering different job levels) to collect typical and key work situations and behavioral events they encountered in their actual work, obtaining first-hand, authentic work materials. These interview transcripts are the core raw materials for generating situational judgment test questions that closely resemble reality.
[0053] The interview transcripts were carefully analyzed by psychology experts or trained coders. First, independent and complete "critical incidents" were identified. Then, based on a pre-defined set of 45 competency rating criteria, the flight attendants' behaviors in the incidents were coded, marking the specific competency characteristics they demonstrated and their corresponding levels.
[0054] Phase Two: Assessment Question Bank Development The numerous specific competency features extracted in the first phase were summarized, clustered, and theoretically analyzed to form a higher-order, multi-dimensional (e.g., 12 dimensions) psychological competency model (such as "leadership," "emotional intelligence," and "risk decision-making") for the development of psychological assessment questions. The assessment expert team selected the 10 dimensions most suitable for situational judgment tests from the 12-dimensional model. Then, based on the "critical events" and "competency feature" database extracted in the first phase, the stems, options, scoring criteria, and rationale for the situational judgment questions were manually written.
[0055] The prepared questions were tested on a small sample. Based on the feedback from the test (such as participants' comprehension of the questions and completion time) and the preliminary results of psychometric analysis, the questions underwent multiple revisions, including streamlining the wording and correcting ambiguous options. Questions with obvious defects were removed, and the quality of the question bank was optimized, ultimately forming a relatively mature initial question bank containing 152 questions across 10 dimensions, preparing for large-scale testing.
[0056] Phase Three: Testing and Data Analysis An initial question bank of multiple items (e.g., 152 items), along with two other mature scales (General Cognitive Ability Test and Mental Health Scale), was administered online to a large number of active flight attendants (e.g., 818). Exploratory factor analysis was performed on the responses to the 152 items. This is a data dimensionality reduction technique aimed at discovering the underlying actual structure behind the items (i.e., which items actually measure the same psychological trait). Items were selected based on statistical indicators (such as factor loadings, communality, etc.). After factor analysis and multiple iterative selections of items, a more stable 5-dimensional psychological competence model was finally determined. Simultaneously, reliability and validity indicators were incorporated, and 59 SJT items with excellent psychometric indicators were ultimately retained. The data analysis results include reliability and criterion-related validity. Reliability refers to the calculation of internal consistency reliability (such as the alpha coefficient) to assess the stability of the item measurement. Criterion-related validity refers to the analysis of the correlation between SJT scores and actual work performance scores to verify the predictive ability of the assessment on work performance.
[0057] Phase Four: Establishment of the Indicator System Experts from civil aviation human resources, psychology, and senior flight attendants were convened to assess the relative importance of the seven dimensions through multiple rounds of anonymous, back-to-back scoring (the Delphi method). The subjective weights determined by the experts were then combined with the criterion-related validity (objective weights) calculated in Phase 3, weighted according to a certain ratio (e.g., 53% expert weight and 47% validity weight), to calculate the final composite weight for each dimension. Based on the final weights of each dimension and measurement norms, a method was developed to synthesize an individual's standardized scores across all dimensions into a weighted total score.
[0058] Step 2: Using the vector knowledge base and structured database built in Step 1, the large language model is driven by the combined retrieval-enhanced generation (RAG) technology to automatically extract structured competency features and key events from the practical cases to be processed and store them in the competency feature and key event database.
[0059] refer to Figure 4 Step two further includes the following processing flow.
[0060] Step 2-1: For the practical cases to be processed, retrieve relevant psychological theories and practical cases from the vector knowledge base, and initially screen and identify the original text fragments containing specific behaviors and situations.
[0061] The case study to be processed is a complete transcript of a flight attendant interview.
[0062] For this transcript of the flight attendant interview, the Dify platform's workflow invokes a configured large language model (LLM, such as DeepSeek), using a specific prompt to guide the LLM in processing the transcript input. The prompt design is as follows: "You are a psychology assistant. Your task is to extract events that flight attendants have personally experienced from interview transcripts so that psychology experts can analyze the flight attendants' competency characteristics."
[0063] [The interview transcript is as follows]
[0064] Please analyze the interview transcript and perform the following tasks.
[0065] 1. Extract the events that the interviewees personally experienced, ensuring that each event is unique.
[0066] 2. Write a brief overview for each event, which should include the time, place, people involved, and core information about the event.
[0067] 3. The output format should be a single event summary per line, without any other content. Note: Ensure that each event is unique, and only output the event summary, not the original text.
[0068] Finally, based on the prompts, the large language model reads through the verbatim transcript of the flight attendant interviews and outputs one or more structured descriptions of key events. For example: {"Event 1": "Handling a passenger who suddenly falls ill during the flight", "Event 2": "Resolving a dispute between two passengers over overhead bin space", ...}. These outputs lay the foundation for further precise analysis.
[0069] Step 2-2: The large language model takes the original text fragments corresponding to the key events retrieved in Step 2-1 and the predefined competency rating criteria as input to the model. It guides the large language model to perform semantic analysis and reasoning through preset analysis prompts to obtain competency features and their levels.
[0070] In step 2-2, the input received by the large language model includes: [Original Interview Transcript] The original text fragments corresponding to the key events extracted in the previous RAG retrieval step; The [Competency Feature Rating Criteria] is a predefined, detailed document containing multiple (e.g., 45) competency features, with each competency feature having a clear level 1, 2, or 3 behavioral description. This is the sole basis for the large language model to make judgments.
[0071] The Dify platform's workflow combines the two inputs (the original interview transcript and the competency rating criteria) into a rich context and invokes the large language model using expert-level prompts. These expert-level prompts aim to transform the large language model's role from "generator" to "evaluation matching engine."
[0072] For example, expert-level prompts are: "You are an experienced organizational behavior expert and AI assessment assistant. Your core task is to act as a precise matching and assessment engine. You will receive a detailed [Competency Rating Criteria] containing 45 competencies, each with a clear level 1, 2, and 3 behavioral description. Your job is not to infer based on general knowledge, but to rigorously match the specific behavioral descriptions in the [Original Interview Transcript] with the items in the [Competency Rating Criteria]. You should only rate a level when the behavior in the original text is highly consistent with the description of a certain level." Under the guidance of expert-level prompts, the large language model analyzes each key event and outputs structured analysis results, directly mapping them to specific competency features and their levels. For example: {"Key Event": "Passengers argue over luggage placement", "Mapped Competency Feature": [{"Feature Name": "Rule Application Ability", "Rating Level": 3}, {"Feature Name": "Leadership", "Rating Level": 2}]}. This structured data is stored in a structured database such as PostgreSQL.
[0073] Steps 2-3: Workflow reliability verification. Based on the coded reliability coefficient, quantitative evaluation is carried out to transform expert experience into evaluation indicators and generate a reliable and structured competency feature and key event library.
[0074] First, for the same interview transcript, both manual and AI coding were performed to generate separate competency lists. Manual coding involved one or more human experts analyzing the interview transcript and extracting what they considered to be the correct competency list. AI coding used the Dify platform's workflow to automatically extract and process the same interview transcript, resulting in the competency list.
[0075] Then, consistency is calculated based on the coding reliability coefficient, which is, for example, the classification agreement coefficient (CA). This involves comparing the competency feature lists generated by human coding and AI coding respectively, and calculating the classification agreement coefficient (CA). The formula for calculating the classification consistency coefficient is: CA = (2 S) / (T1 + T2), where CA is the classification consistency coefficient, T1 is the number of competency features extracted by manual coding, T2 is the number of competency features extracted by AI coding, and S is the number of identical competency features extracted by manual coding and AI coding.
[0076] For example, regarding interview transcripts: The expert coder extracted 16 features (T1 = 16), and the AI automatically extracted 17 features (T2 = 17). Six features were identical in both extraction results (S = 11). Substituting these into the formula for calculation: CA = (2 S) / (T1 + T2) = (2 11) / (16 + 17) = 22 / 33 = 0.67 Finally, the results of the consistency calculation are evaluated: when the CA reliability coefficient is greater than the preset value (e.g., 0.6), it is evaluated that the extraction results of the AI workflow have acceptable consistency with human experts.
[0077] Step 3: Based on retrieval enhancement generation, use vector knowledge base, structured database, and competency feature and key event database to automatically generate situational judgment test questions.
[0078] refer to Figure 5 Step 3-1: The Dify platform takes the target competency features as input, performs a search in a multi-source database to obtain reference material fragments related to the target competency features, uses the reference material fragments as contextual information, and constructs prompt words for generating test questions.
[0079] The Dify workflow injects all the search results obtained from multi-source retrieval—that is, the information most relevant to the competency features of the target (such as definition, behavioral indicators, key events, and question examples)—as context into the prompt word generated for the question, and sends them together to the large language model.
[0080] The search targets for multi-source retrieval based on search enhancement generation (RAG) include: Retrieve the definition and behavioral indicators of the competency features for the target from the structured database; retrieve all key events marked as competency features for the target from the competency feature and key event database generated in step two; retrieve historical high-quality question examples from the structured database as a reference for generating the format and style.
[0081] Step 3-2: Based on the input content sent in Step 3-1, the large language model performs tasks according to prompt words to generate test questions with standardized format and professional content. These test questions form a candidate SJT question set with ability mapping.
[0082] For example, a test question includes a stem, options, a scoring reference, and a reasoning.
[0083] The prompt word instructions in step 3-2 of the large language model are highly structured.
[0084] Taking a scenario targeting airline flight attendants as an example, here is an example of the prompt words: "You are a civil aviation scenario SJT question design expert. You combine historical question styles and event analysis to generate scenarios and questions for the emotion management ability dimension."
[0085] The meanings of the competency dimensions being assessed are as follows: (retrieved from the [Competency Feature Rating Criteria])
[0086] Refer to the history question template: (searched from the [History Question Bank])
[0087] Refined event background: (extracted from interview transcripts and existing questions)
[0088] Ability and behavior levels (fixed): Level 1 = Ability deficiency, Level 2 = Basic standard met, Level 3 = Excellent standard met.
[0089] Question structure requirements (strictly adhere to): Scenario Description: The description must include "Identity + Specific Work Situation + Core Conflict + Constraints." The scenario should closely resemble the actual work of the target position. The conflict should focus on the key assessment dimensions (e.g., "Conflict between rules and flexible service," "Conflict between limited resources and high customer demands"). The constraints should be clearly defined (e.g., "Limited crew resources," "Tight timeframe"). 1. Structure: Fully matches the historical question template, including "Context (based on a concise background of about 100 words, supplemented with details to make the scene realistic) + Question (aligned with the key points of the dimension, such as "How would you handle this?" or "What would you do?" For the dimension of rule application ability, the question can also be "What is the most compliant action at this time?" etc.) + 4 options (including 1 correct option = level 3, 1 distractor option = level 2, 2 incorrect options = level 1, with scores of 3, 1, and 0 respectively)".
[0090] 2. Question design: mainly refer to the template of history questions, which is roughly the main body + scenario + question. The references should be strongly related to the core logic of the question and avoid using the same reference repeatedly. 3. Option Design: All operations must conform to logic.
[0091] - Correct item: Meets the criteria for Level 3 behavior in this dimension.
[0092] - Disruptive items: Meets Level 2 behavior, exhibiting "partial compliance but defects" (e.g., "delayed processing, incomplete rule execution"). - Error: Meets Level 1 behavior, with issues such as "violation of regulations, shirking responsibility, and exceeding authority"; 4. Language style: Consistent with the supplementary exam questions, formal, professional, and relevant to work scenarios, without colloquial expressions.
[0093] 5. If the question stem mentions "according to xxx rule or manual", then the following must be added to the end of the question stem: (Note: the manual or standard mentioned in the question are fictitious). 6. In the scoring reference and reasoning, replace behavioral levels 1, 2, and 3 with incompetent, basically competent, and competent. Output format (strictly fixed and cannot be modified) {Question number (e.g., 1.)}. {Question content} A. {Option A} B. {Option B} C. {Option C} D. {Option D} Scoring criteria and reasons: A---{Corresponding Score} Score: {Reasons, based on the reasonableness of the behavior, consequences, and basis, clearly indicating competence / basic competence / incompetence} B---{Corresponding Score}: {Reasons, written according to the above requirements} C---{Corresponding Score}: {Reasons, written according to the above requirements} D---{Corresponding Score}: {Reasons, written according to the above requirements} Requirements: Options should be randomly shuffled; the 3-point correct answer is not fixed; the reasoning for each point should correspond to one of the options; the language should be concise and conform to the SJT specification. Step 4: Evaluate and optimize the quality of the situational judgment test questions generated in Step 3.
[0094] Step four is optional. The results of the quality assessment are used to determine whether the quality of the questions meets the standards. If they do, the questions are included in the final situational judgment test question bank. If they do not meet the standards, optimization instructions are generated. These optimization instructions guide and optimize the generation process of steps two and three through a feedback loop.
[0095] refer to Figure 6 Step four further includes the following processing steps.
[0096] First, the candidate SJT questions generated in step three are subjected to automated pre-evaluation in multiple dimensions (including structure, language and relevance). The basic rules for automatic filtering set in the pre-evaluation include checking the question length, the number of options, and whether there are sensitive words.
[0097] The generated candidate SJT items are then submitted to civil aviation and psychology experts for manual review. The experts assess: face validity (e.g., are the scenarios realistic? Is the language fluent? Does it reflect the actual work of flight attendants?) and content validity (e.g., do the items accurately test the target competency characteristics? Are the options reasonably designed?)
[0098] Next, a decision is made based on the evaluation results of automated assessment and manual review. If the standard is met, the question is included in the final question bank; if the standard is not met, it enters the optimization process to form a closed loop.
[0099] This step also includes experts providing specific suggestions for improvement (e.g., "Option C is not misleading enough and should be changed to a more common incorrect response"). These suggestions are not directly used to modify the questions themselves, but are recorded, analyzed, and transformed into optimization instructions for previous steps. For example, if multiple experts point out that "the generated options are not discriminatory," this feedback will be passed to step three, suggesting the need to optimize the prompts used for question generation, such as emphasizing "generating highly misleading distractors" in the instructions. If inaccurate feature extraction is found, feedback will be passed to step two, suggesting the need to supplement the theoretical knowledge base or adjust the feature extraction prompts.
[0100] Finally, the revised questions, based on feedback, will be used for testing, and the collected data will be analyzed using psychometrics. The analytical indicators include: difficulty (pass rate / scoring rate), discrimination (comparing the pass rate difference between high and low groups on this question), confirmatory factor analysis (measuring whether actual ability performance matches the dimensions of the question), and criterion-related validity analysis (analyzing the correlation between question scores and employees' actual work performance).
[0101] also, Figure 7 This illustrates the principle of an embodiment of the test question generation system based on retrieval enhancement according to the present invention. See also Figure 7 The system in this embodiment includes: a multi-source knowledge base construction module, a competency feature intelligent extraction module, a test question automatic generation module, and a test question quality assessment and closed-loop optimization module.
[0102] The multi-source knowledge base construction module performs data parsing and vectorization processing on the raw data, thereby constructing a vector knowledge base and a structured database. The internal processing of the multi-source knowledge base construction module and... Figure 1 The process in step one is the same as shown, so it will not be repeated here.
[0103] The competency feature intelligent extraction module utilizes a vector knowledge base and structured database built by a multi-source knowledge base construction module. It employs retrieval-enhanced generative techniques to drive a large language model, automatically extracting structured competency features and key events from the practical cases to be processed, and storing them in a competency feature and key event database. The internal processing of the competency feature intelligent extraction module and... Figure 1 The process in step two is the same as shown, so it will not be repeated here.
[0104] The automatic question generation module, based on retrieval-enhanced generation, utilizes a vector knowledge base, a structured database, and a competency feature and key event database to automatically generate situational judgment test questions. The internal processing of the automatic question generation module and... Figure 1 The process in step three is the same as shown, so it will not be repeated here.
[0105] The question quality assessment and closed-loop optimization module is optional. It assesses and optimizes the quality of situational judgment test questions generated by the automatic question generation module. The specific internal processing of the question quality assessment and closed-loop optimization module is as follows: Figure 1 The process in step four is the same as shown, so it will not be repeated here.
[0106] See Figure 8 The electronic device of the present invention includes a controller, which comprises a processor, a communication interface, a memory, and a communication bus. The processor, communication interface, and memory communicate with each other via the communication bus. The memory stores computer programs; the processor, when executing the programs stored in the memory, implements, as follows: Figure 1 The steps of the test question generation method based on retrieval enhancement are described in the embodiment.
[0107] Furthermore, the present invention also discloses a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements as follows: Figure 1 The steps of the test question generation method based on retrieval enhancement are described in the embodiment.
[0108] Although the methods described above are illustrated and depicted as a series of actions for the sake of simplicity, it should be understood and appreciated that these methods are not limited by the order of the actions, as some actions may occur in a different order and / or concurrently with other actions from the illustrations and descriptions herein or not illustrated and described herein but which may be understood by those skilled in the art, according to one or more embodiments.
[0109] Those skilled in the art will further appreciate that the various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in a generalized manner in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of the invention.
[0110] The various illustrative logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein can be implemented or performed using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor may be a microprocessor, but in alternatives, it may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration.
[0111] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor such that the processor can read and write information to / from the storage medium. In an alternative, the storage medium may be integrated into the processor. The processor and storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor and storage medium may reside as discrete components in the user terminal.
[0112] In one or more exemplary embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software as a computer program product, the functionality may be stored or transmitted as one or more instructions or code on or through a computer-readable medium. A computer-readable medium includes both computer storage media and communication media, encompassing any medium that facilitates the transfer of a computer program from one location to another. A storage medium may be any available medium accessible to a computer. By way of example and not limitation, such a computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and is accessible to a computer. Any connection is also legitimately referred to as a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of a medium. As used in this article, disk and disc include compact discs (CDs), laser discs, optical discs, digital multi-purpose discs (DVDs), floppy disks, and Blu-ray discs. Disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of these should also be included within the scope of computer-readable media.
[0113] The prior description of this disclosure is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the examples and designs described herein, but should be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A test question generation method based on retrieval enhancement, characterized in that, The methods include: Step 1: Perform data parsing and vectorization processing on the raw data to build a vector knowledge base and a structured database; Step 2: Using the vector knowledge base and structured database built in Step 1, and employing a large language model that combines retrieval enhancement and generation techniques, structured competency features and key events are automatically extracted from the practical cases to be processed and stored in the competency feature and key event database. Step 3: Based on retrieval enhancement generation, use vector knowledge base, structured database, and competency feature and key event database to automatically generate situational judgment test questions.
2. The test question generation method based on retrieval enhancement according to claim 1, characterized in that, The raw data was obtained from multiple sources, including theoretical literature, practical case studies, and historical question banks.
3. The test question generation method based on retrieval enhancement according to claim 2, characterized in that, Step one further includes: Step 1-1: Obtain the raw data; Steps 1-2: Based on the characteristics of different data sources, differentiated data processing and data entry are performed to generate a hybrid knowledge base consisting of a structured database and a vector knowledge base. The differentiation includes: extracting key information from unstructured interview texts in the practical case library and storing the extracted key events and competency features in the structured database; performing structured parsing on the historical question bank and storing the parsed standardized situational judgment test questions in the structured database; and vectorizing theoretical literature and storing the generated semantic vectors in the vector knowledge base.
4. The test question generation method based on retrieval enhancement according to claim 1, characterized in that, Step two further includes: Step 2-1: For the practical cases to be processed, retrieve relevant psychological theories and practical cases from the vector knowledge base through retrieval enhancement, and initially screen and identify the original text fragments containing specific behaviors and situations; Step 2-2: The large language model takes the original text fragments corresponding to the key events retrieved in Step 2-1 and the predefined competency rating criteria as input to the model. It guides the large language model to perform semantic analysis and reasoning through preset analysis prompts to obtain competency features and their levels. Steps 2-3: Quantitative evaluation based on coding reliability coefficients, transforming expert experience into evaluation indicators, and generating a reliable, structured competency feature and key event library.
5. The test question generation method based on retrieval enhancement according to claim 1, characterized in that, Step three further includes: Step 3-1: The Big Language Application Development Platform takes the target competency features as input, performs a search in a multi-source database to obtain reference material fragments related to the target competency features; and uses the reference material fragments as contextual information to construct prompt words for generating test questions. Step 3-2: Based on the input content sent in Step 3-1, the large language model performs tasks according to prompt words to generate test questions. Multiple test questions form a candidate situation judgment test question set.
6. The test question generation method based on retrieval enhancement according to claim 1, characterized in that, The method also includes: Step 4: Evaluate and optimize the quality of the situational judgment test questions generated in Step 3.
7. A test question generation system based on retrieval enhancement, characterized in that the system... include: The multi-source knowledge base construction module performs data parsing and vectorization processing on the raw data, thereby constructing a vector knowledge base and a structured database; The competency feature intelligent extraction module utilizes a vector knowledge base and structured database built by a multi-source knowledge base construction module. It employs retrieval enhancement generation technology to drive a large language model, automatically extracting structured competency features and key events from the practical cases to be processed, and storing them in a competency feature and key event database. The automatic question generation module, based on retrieval-enhanced generation, utilizes a vector knowledge base, a structured database, and a competency feature and key event database to automatically generate situational judgment test questions.
8. The test question generation system based on retrieval enhancement according to claim 7, characterized in that, The system also includes: The question quality assessment and closed-loop optimization module assesses and optimizes the quality of the situational judgment test questions generated by the automatic question generation module.
9. An electronic device, characterized in that, The electronic device includes a controller, which includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus. Memory, used to store computer programs; When a processor executes a program stored in memory, it implements the steps of the question generation method based on retrieval enhancement as described in any one of claims 1-6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the question generation method based on retrieval enhancement as described in any one of claims 1-6.