A railway document training question generation method and device, electronic equipment and storage medium

CN122674643APending Publication Date: 2026-09-01BEIJING TRAFFIC & TRANSPORT TECH CORP LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611075960.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-20
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

传统技术通常采用人工出题的方式,需要行业专家深刻理解文档内容和重点,并提取出关键字和重点、难点内容,再通过人工的方式进行题目生成,时间成本较高,且随着文档修订、变化,需要持续投入人力进行相关题库的编写、修正工作

Benefits of technology

[0008]本发明的有益效果是:通过将铁路行业文档自动划分为多个文档片段并与各岗位培训内容大纲进行语义关联,实现了文档内容到知识点的精准映射,避免了传统规则匹配方式的扩展性与泛化性不足问题,能够适应文档频繁更新的需求;进而利用基于历史题库训练得到的生成式大语言模型进行出题,使模型通过学习历史题目与文档片段的对应关系掌握铁路专业知识的出题规范,有效避免了生成抓不住重点或包含页码条款编号等无效信息的低质量题目,在无需多模型级联和复杂逻辑判定的前提下显著提升了出题效率、题库质量与系统可维护性,解决了人工出题成本高、更新滞后以及传统自动化出题质量不稳定的技术问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122674643A_ABST
    Figure CN122674643A_ABST
Patent Text Reader

Abstract

This invention relates to a method, apparatus, electronic device, and storage medium for generating training questions for railway documents. The method includes: S10, acquiring railway industry documents and training content outlines for various positions, and dividing the railway industry documents into multiple document fragments; S20, performing semantic association on all document fragments and all training content outlines to obtain at least one document fragment matching each training content outline; S30, based on the at least one document fragment matching each training content outline, obtaining a target question bank corresponding to the railway industry documents through a pre-trained question generation model. This invention achieves automatic question generation for railway documents based on large language model fine-tuning and semantic association technology, improving question generation efficiency and quality while solving the problems of high manual costs, delayed updates, and insufficient generalization of traditional automated question generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence and natural language processing technology. Specifically, this invention relates to a method, apparatus, electronic device, and storage medium for generating questions for railway document training. Background Technology

[0002] The railway industry has a large number of non-public documents and knowledge bases. Many front-line employees need to learn and understand the content of these documents through training systems. However, the documents themselves are voluminous and do not highlight the key points. Furthermore, the documents are revised and updated frequently, resulting in high learning and understanding costs for front-line employees, making it difficult for them to quickly grasp the key points and difficulties of the documents.

[0003] Creating targeted questions based on documents and conducting exams and training based on these questions is an effective way to help employees understand key points and document content. Traditional methods usually involve manually creating questions, which requires industry experts to have a deep understanding of the document content and key points, extract keywords and key and difficult points, and then manually generate questions. This is time-consuming, and as the document is revised and changed, it requires continuous investment of manpower to compile and revise the relevant question bank.

[0004] There are also some traditional technical means to automate question generation, but the question generation method is usually rule matching and hard coding of logic, which results in insufficient scalability and generalization of the generated questions, making it impossible to generate a sufficient question bank and thus failing to play an effective role in vocational education and training.

[0005] These problems and shortcomings ultimately affect the quality and quantity of the question bank, which in turn affects frontline employees' mastery of industry knowledge, hindering daily work and improving office efficiency. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a method, apparatus, electronic device and storage medium for generating questions for railway document training, which aims to solve at least one of the above-mentioned technical problems.

[0007] Firstly, the technical solution of the present invention to solve the above-mentioned technical problems is as follows: a method for generating questions for railway document training, the method comprising: S10: Obtain railway industry documents and training content outlines for various positions, and divide the railway industry documents into multiple document segments; S20, Semantically correlate all document fragments and all training content outlines to obtain at least one document fragment that matches each training content outline; S30: Based on at least one document fragment matched to each training content outline, a target question bank corresponding to railway industry documents is obtained through a pre-trained question generation model. The question generation model is a generative large language model trained based on a historical question bank.

[0008] The beneficial effects of this invention are as follows: By automatically dividing railway industry documents into multiple document fragments and semantically associating them with the training content outlines for each position, a precise mapping from document content to knowledge points is achieved, avoiding the insufficient scalability and generalization of traditional rule matching methods, and adapting to the needs of frequent document updates; furthermore, by using a generative large language model trained based on a historical question bank to generate questions, the model learns the correspondence between historical questions and document fragments to master the question-generating norms of railway professional knowledge, effectively avoiding the generation of low-quality questions that fail to grasp the key points or contain invalid information such as page numbers and clause numbers. Without the need for multi-model cascading and complex logical judgments, it significantly improves the efficiency of question generation, the quality of the question bank, and the maintainability of the system, solving the technical problems of high cost, delayed updates, and unstable quality of traditional automated question generation.

[0009] Based on the above technical solution, the present invention can be further improved as follows.

[0010] Furthermore, the semantic association of all document fragments and all training content outlines yields at least one document fragment matching each training content outline, including: For each training content outline, perform semantic relevance analysis between the training content outline and each document fragment to obtain at least one document fragment that matches the training content outline. or, Extract the key field identifiers for each document fragment. The key field identifiers are used to identify the core content of the corresponding document fragment. For each training content outline, perform semantic relevance analysis between the training content outline and each key field identifier to obtain at least one document fragment that matches the training content outline. or, Generative large models are used to associate each document fragment with each training content outline, resulting in at least one document fragment that matches each training content outline.

[0011] Furthermore, the above question-generating model was trained using the following method: Obtain the historical question bank and the corresponding historical question documents. Each historical question in the historical question bank is marked with a relevant document fragment from the historical question document. Extract the key fields of each question in the historical question bank. The key fields include at least the question type, question stem, options, correct answer, related document fragments and explanations. The related document fragments are used to indicate the specific clauses or paragraphs in the historical question-setting documents corresponding to the question. The key fields of each question are associated with relevant document fragments in the corresponding historical question documents to construct a training sample set. Each training sample in the training sample set includes a document fragment as input and the key fields of the question as output. The generative large language model is trained in a supervised manner using the training sample set to obtain the question generation model.

[0012] Furthermore, each training content outline includes at least one knowledge point; based on at least one document fragment matched to each training content outline, a target question bank corresponding to railway industry documents is obtained through a pre-trained question generation model, including: S301, For each training content outline, based on at least one document fragment matched with the training content outline, the question set corresponding to each knowledge point in the training content outline under different question types is obtained through the question generation model; S302 performs semantic-level deduplication on the question bank sets generated for each knowledge point under each question type to obtain a target question bank that does not contain questions with substantially the same or highly similar content.

[0013] Furthermore, the question bank sets generated for each knowledge point under each question type are semantically deduplicated to obtain a target question bank containing no substantially identical or highly similar questions, including: Construct a vector knowledge base, which is used to store the semantic feature vectors of questions in the target question bank after deduplication and retention, question by question; For each question in each question bank set, perform the following processing to obtain the deduplicated target question bank: Calculate the semantic feature vector of the problem; The semantic feature vector of the question is compared with the semantic feature vectors of each question already stored in the vector knowledge base. If the similarity between the semantic feature vector of a question and the semantic feature vector of any question in the vector knowledge base exceeds a preset threshold, and the question type of the question is the same as that of any question, then the question is deleted; otherwise, the question is retained in the target question bank, and the semantic feature vector of the question is stored in the vector knowledge base.

[0014] Furthermore, after obtaining the target question bank, the method also includes: Based on the mapping relationship between each question in the target question bank and the training content outline and document fragments, check whether each question type corresponding to each training content outline meets the preset question quantity requirements; If the number of questions for any question type under any training content outline does not meet the preset question quantity requirement, then for any training content outline, repeat steps S301 to S302 to supplement and generate questions of the corresponding question type to the target question bank.

[0015] Secondly, in order to solve the above-mentioned technical problems, the present invention also provides a railway document training question-generating device, the device comprising: The acquisition module is used to acquire railway industry documents and training content outlines for various positions, and divides the railway industry documents into multiple document fragments; The semantic association module is used to perform semantic association on all document fragments and all training content outlines to obtain at least one document fragment that matches each training content outline. The question generation module is used to generate a target question bank corresponding to railway industry documents by matching at least one document fragment with each training content outline and using a pre-trained question generation model. The question generation model is a generative large language model trained based on the historical question bank.

[0016] Thirdly, in order to solve the above-mentioned technical problems, the present invention also provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the railway document training question generation method of the present application.

[0017] Fourthly, in order to solve the above-mentioned technical problems, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the railway document training question generation method of the present application.

[0018] Additional aspects and advantages of this application will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of this application. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments of the present invention will be briefly introduced below.

[0020] Figure 1 A flowchart illustrating a railway document training question generation method according to an embodiment of the present invention; Figure 2 A flowchart illustrating another method for generating questions for railway document training, provided as an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a railway document training question generation device according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present invention. Detailed Implementation

[0021] The principles and features of the present invention are described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0022] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.

[0023] The data acquisition process involved in this invention follows the principles of legality, legitimacy, and necessity. Based on obtaining the explicit authorization and consent of the user, only the minimum necessary information required to achieve the purpose is collected, and data security protection obligations are fulfilled in accordance with the law.

[0024] The solution provided in this invention can be applied to any application scenario that requires generating questions for railway industry documents. The solution provided in this invention can be executed by any electronic device, such as a user's terminal device, including at least one of the following: smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, smart TV, or smart in-vehicle device.

[0025] This invention provides a possible implementation, such as... Figure 1 As shown, a flowchart of a railway document training question generation method is provided. This method can be executed by any electronic device, such as a terminal device, or by a terminal device and a server. For ease of description, the method provided in this embodiment will be described below using a terminal device as the execution subject as an example. Figure 1 The flowchart shown indicates that the method may include the following steps: S10: Obtain railway industry documents and training content outlines for various positions, and divide the railway industry documents into multiple document segments; S20, Semantically correlate all document fragments and all training content outlines to obtain at least one document fragment that matches each training content outline; S30: Based on at least one document fragment matched to each training content outline, a target question bank corresponding to railway industry documents is obtained through a pre-trained question generation model. The question generation model is a generative large language model trained based on a historical question bank.

[0026] By automatically dividing railway industry documents into multiple document fragments and semantically associating them with the training content outlines for each position, a precise mapping from document content to knowledge points is achieved. This avoids the limitations of traditional rule-based matching methods in terms of scalability and generalization, and can adapt to the needs of frequent document updates. Furthermore, a generative large language model trained on a historical question bank is used to generate questions. The model learns the correspondence between historical questions and document fragments to master the question-generating standards of railway professional knowledge, effectively avoiding the generation of low-quality questions that fail to grasp the key points or contain invalid information such as page numbers and clause numbers. Without the need for multiple model cascading and complex logical judgments, this significantly improves question generation efficiency, question bank quality, and system maintainability, solving the technical problems of high cost, delayed updates, and unstable quality of traditional automated question generation.

[0027] The following specific embodiments further illustrate the solution of the present invention. The purpose of the present invention is to address the current situation where railway industry issues, documents, and training content outlines for various positions are updated frequently, and manual question creation and evaluation are inefficient. This invention combines a generative large language model to extract and understand knowledge points and perform semantic association, thereby creating a multi-dimensional and automated question creation and update system. This system aims to improve the work efficiency of users, enhance the efficiency and effectiveness of employee training, and reduce the burden of question creation and review on managers (such as vocational education staff).

[0028] Based on this, combined Figure 2 This embodiment provides a method for generating questions for railway document training, which may include the following steps: S10: Obtain railway industry documents and training content outlines for various positions, and divide the railway industry documents into multiple document segments; The railway industry documents include technical specifications, emergency plans, equipment manuals, etc. Before dividing the railway industry documents into multiple document fragments, the method also includes: parsing the railway industry documents through a multi-source parser to ensure the consistency of the parsing structure and parsing field dimensions. Specifically, for railway industry documents in doc format, the parsing tool converts them into docx format documents; for railway industry documents in docx format, an open-source parser is used for parsing; for railway industry documents in text-based pdf format, layout and content recognition are performed using an API; for scanned pdf format railway industry documents, OCR models and text recognition software are used for parsing; and for railway industry documents in html and txt formats, the text content is extracted and html tags are removed.

[0029] Specifically, railway industry documents can be divided into multiple document segments according to chapters, clauses, or natural paragraphs. A document segment can be a chapter or a paragraph.

[0030] The training outlines for the aforementioned positions refer to the systematic training requirements and knowledge structure developed for each specific position in the railway industry (such as station duty officers, train drivers, dispatchers, and train attendants). These outlines are typically organized in a hierarchical or tree-like structure, clearly defining the theoretical knowledge, practical skills, safety regulations, and other key training points that personnel in that position must master. As an example, the training outlines for each position may include: Theoretical Knowledge 1.1 Basic Theory 1.1.1 Railway Lines and Stations.

[0031] S20, Semantically correlate all document fragments and all training content outlines to obtain at least one document fragment that matches each training content outline; Optionally, semantic association is performed on all document fragments and all training content outlines to obtain at least one document fragment matching each training content outline, including: For each training content outline, perform semantic relevance analysis between the training content outline and each document fragment to obtain at least one document fragment that matches the training content outline. Specifically, for each document fragment in set A formed by all document fragments, a semantic relevance analysis is performed directly between the document fragment and the training content outline. This analysis combines the semantic features of BM2.5 literal features and the embedding model, and uses a multi-classification task to determine the relevance.

[0032] or, Extract the key field identifiers for each document fragment. The key field identifiers are used to identify the core content of the corresponding document fragment. For each training content outline, perform semantic relevance analysis between the training content outline and each key field identifier to obtain at least one document fragment that matches the training content outline. Optionally, before extracting the key field identifiers of each document fragment, the set A formed by all document fragments can be preprocessed to extract the tag outline of set A, that is, to remove irrelevant content in set A, and then the key field identifiers can be extracted.

[0033] or, Generative large models are used to associate each document fragment with each training content outline, resulting in at least one document fragment that matches each training content outline.

[0034] S30: Based on at least one document fragment matched to each training content outline, a target question bank corresponding to railway industry documents is obtained through a pre-trained question generation model. The question generation model is a generative large language model trained based on a historical question bank.

[0035] Traditional question generation methods typically use context engineering (prompt projects) to directly generate questions. However, these models often lack a good understanding of the constraints, characteristics, and question types specific to the railway industry, resulting in questions that fail to capture the key points (e.g., specifying the page number of content xx). One-shot or few-shot question generation methods consume the model's context window, leading to truncation when the scope of questions is large, preventing some original input text from being included. This solution collects historical question documents and a historical question bank as a positive sample set, performs labeled training, and associates these with historical context to obtain a complete and scenario-based question bank. By learning from this data, the model develops the ability to generate professional, clear, and compliant questions.

[0036] Based on this, optionally, the above question generation model is trained in the following way: Obtain the historical question bank and the corresponding historical question documents. Each historical question in the historical question bank is marked with a relevant document fragment from the historical question document. Extract the key fields of each question in the historical question bank. The key fields include at least the question type, question stem, options, correct answer, related document fragments and explanations. The related document fragments are used to indicate the specific clauses or paragraphs in the historical question-setting documents corresponding to the question. The key fields of each question are associated with relevant document fragments in the corresponding historical question documents to construct a training sample set. Each training sample in the training sample set includes a document fragment as input and the key fields of the question as output. The generative large language model is trained in a supervised manner using the training sample set to obtain the question generation model.

[0037] Each question can be a single choice, multiple choice, true or false, fill in the blank, or short answer.

[0038] In this application, each relevant document fragment and explanation can be highlighted, and a generative large language model can be used to repeatedly generate questions and explanations for the same relevant document fragment with different content and types of annotations. This allows the generative large language model to understand the key knowledge point features of the questions, reduce irrelevant content (such as page numbers, clause numbers), and enhance the question generation effect. Simultaneously, to deepen the model's understanding of the questions, strong negative samples can be manually labeled to reinforce types and content that should not be generated as questions, and the reasons can be explained in the explanation field to further enhance the model's understanding.

[0039] By employing the aforementioned combination of data annotation strategies, high-quality fine-tuning data was obtained, and the question-generating model was fine-tuned (model selection: qwen3-14B). Since the fields of the annotated data are all from railway industry documents, the model fine-tuned after this stage of annotation incorporates railway knowledge representations, and its relevance results on railway professional documents outperform public models (trained using general public datasets), context engineering, or pure few-shot solutions.

[0040] Optionally, the above training sample set can be divided into a training set, a validation set, and a test set in a 7:1:2 ratio, and then the model can be trained in a supervised manner.

[0041] Optionally, each training content outline includes at least one knowledge point; based on at least one document fragment matched to each training content outline, a target question bank corresponding to railway industry documents is obtained through a pre-trained question generation model, including: S301, For each training content outline, based on at least one document fragment matched with the training content outline, the question set corresponding to each knowledge point in the training content outline under different question types is obtained through the question generation model; Furthermore, in this solution, the number of questions and the types of questions required in the actual use scenario can be combined to call up the question generation task. For each knowledge point in the training content outline, a certain number of question sets for each of the five question types can be generated, denoted as q1 q2...qi, where i represents the total number of questions in the question set.

[0042] S302 performs semantic-level deduplication on the question bank sets generated for each knowledge point under each question type to obtain a target question bank that does not contain questions with substantially the same or highly similar content.

[0043] Furthermore, the question bank sets generated for each knowledge point under each question type are semantically deduplicated to obtain a target question bank containing no substantially identical or highly similar questions, including: Construct a vector knowledge base, which is used to store the semantic feature vectors of questions in the target question bank after deduplication and retention, question by question; For each question in each question bank set, perform the following processing to obtain the deduplicated target question bank: Calculate the semantic feature vector of the problem; The semantic feature vector of the question is compared with the semantic feature vectors of each question already stored in the vector knowledge base. If the similarity between the semantic feature vector of a question and the semantic feature vector of any question in the vector knowledge base exceeds a preset threshold, and the question type of the question is the same as that of any question, then the question is deleted; otherwise, the question is retained in the target question bank, and the semantic feature vector of the question is stored in the vector knowledge base.

[0044] For each question bank set q1 q2…qi, similarity association can be performed through semantic vector matching to avoid the same / similar questions appearing for the same knowledge point or the same segment.

[0045] Specifically, a vector knowledge base is first constructed. For the nth question qn (n not greater than i), the vector knowledge base is first checked to see if it is empty. If it is empty, question qn and its semantic features (semantic vector) are stored in the vector knowledge base. If the vector knowledge base is not empty, the similarity value between question qn and the semantic features of a question in the vector knowledge base is checked to see if it exceeds a specific threshold (assumed to be 0.9). If it exceeds the specific threshold, it means that question qn has a high similarity to a question in the vector knowledge base. If question qn and a question in the vector knowledge base have the same question type, question qn is deleted; otherwise, question qn is retained in the target question bank. The above method can be used to process each question in the question bank set.

[0046] Optionally, after obtaining the target question bank, the method further includes: Based on the mapping relationship between each question in the target question bank and the training content outline and document fragments, check whether each question type corresponding to each training content outline meets the preset question quantity requirements; If the number of questions for any question type under any training content outline does not meet the preset question quantity requirement, then for any training content outline, repeat steps S301 to S302 to supplement and generate questions of the corresponding question type to the target question bank.

[0047] Specifically, since each question in the generated target question bank contains the relationship between the question and the training content outline (outline-segment-question), it is possible to check that each training content outline has a certain number of questions for different question types. If a question type is missing a question, repeat steps S301 to S302 to supplement the question.

[0048] Based on this, since railway industry documents and training content outlines for various positions may change, when the training content outlines for various positions change, repeat steps S10 to S30 to obtain the relationship between the new training content outline and the document fragments in the railway industry documents, and perform model reasoning for the new training content outline to obtain different types of question banks corresponding to the new training content outline.

[0049] If the training content outline remains unchanged, but the railway industry documents are updated, it is necessary to create new questions while deleting the questions that were previously associated with the old document fragments. Since the original questions recorded their relationship with the outline and document fragments, this step can be achieved directly through cascading deletion.

[0050] To better illustrate and understand the principle of the method provided by this invention, the following description uses an optional specific embodiment to illustrate the solution of this invention. It should be noted that the specific implementation of each step in this specific embodiment should not be construed as a limitation of the solution of this invention. Other implementations that can be conceived by those skilled in the art based on the principle of the solution provided by this invention should also be considered within the scope of protection of this invention.

[0051] In this embodiment, examples of training content outlines and railway industry documents are provided: The training content outline refers to the training requirements for each position recorded in the training manual for each position.

[0052] For example (taking the position of station duty officer as an example), the hierarchical structure of its training content outline is as follows: 1. Theoretical Knowledge - 1.1 Basic Theory; 1.1.1 Railway lines, stations, and turnouts; 1.1.2 Railway Signals; 1.1.3 Locomotives and rolling stock; The connection between the training content outline and training materials (railway industry documents) is as follows: Examples of correlation results (taking railway lines and stations as examples): 1. Theoretical Knowledge - 1.1 Basic Theory - 1.1.1 Railway Lines and Stations; Articles 4-6 of the "Station Operation Regulations" (hereinafter referred to as the "Station Regulations"); Articles 17-32 of the Railway Technical Management Regulations (hereinafter referred to as the "Regulations").

[0053] The solution of the present invention has the following beneficial effects: By fine-tuning the model strategy based on railway industry documents and building an automatic question generation system for multiple document types based on the fine-tuned model, the system can understand the characteristics of industry documents and knowledge points, automatically generate questions of multiple types, and has high usability. This fundamentally solves the problems of low efficiency in question generation due to manual document understanding, inconsistent question quality, and high document update and maintenance costs. It improves the usability of vocational education scenarios, reduces the cost of question generation, maintenance, and updates for managers, and improves the work efficiency of users. It better leverages the advantages of the railway big data model in multiple scenarios such as manual reference, knowledge point learning, vocational education training, and employee profiling.

[0054] Based on and Figure 1 Using the same principle as the method shown, this embodiment of the invention also provides a railway document training question-generating device 20, such as... Figure 3As shown, the railway document training question generation device 20 may include an acquisition module 210, a semantic association module 220, and a question generation module 230, wherein: The acquisition module 210 is used to acquire railway industry documents and training content outlines for various positions, and divides the railway industry documents into multiple document fragments; The semantic association module 220 is used to perform semantic association on all document fragments and all training content outlines to obtain at least one document fragment that matches each training content outline. The question generation module 230 is used to obtain the target question bank corresponding to railway industry documents by matching at least one document fragment with each training content outline through a pre-trained question generation model. The question generation model is a generative large language model trained based on the historical question bank.

[0055] Optionally, when the semantic association module 220 performs semantic association on all document fragments and all training content outlines to obtain at least one document fragment matching each training content outline, it is specifically used for: For each training content outline, perform semantic relevance analysis between the training content outline and each document fragment to obtain at least one document fragment that matches the training content outline. or, Extract the key field identifiers for each document fragment. The key field identifiers are used to identify the core content of the corresponding document fragment. For each training content outline, perform semantic relevance analysis between the training content outline and each key field identifier to obtain at least one document fragment that matches the training content outline. or, Generative large models are used to associate each document fragment with each training content outline, resulting in at least one document fragment that matches each training content outline.

[0056] Optionally, the above question generation model is trained in the following way: Obtain the historical question bank and the corresponding historical question documents. Each historical question in the historical question bank is marked with a relevant document fragment from the historical question document. Extract the key fields of each question in the historical question bank. The key fields include at least the question type, question stem, options, correct answer, related document fragments and explanations. The related document fragments are used to indicate the specific clauses or paragraphs in the historical question-setting documents corresponding to the question. The key fields of each question are associated with relevant document fragments in the corresponding historical question documents to construct a training sample set. Each training sample in the training sample set includes a document fragment as input and the key fields of the question as output. The generative large language model is trained in a supervised manner using the training sample set to obtain the question generation model.

[0057] Optionally, each training content outline includes at least one knowledge point; when the question generation module obtains the target question bank corresponding to railway industry documents based on at least one document fragment matched to each training content outline using a pre-trained question generation model, it specifically performs the following steps: S301, For each training content outline, based on at least one document fragment matched with the training content outline, the question set corresponding to each knowledge point in the training content outline under different question types is obtained through the question generation model; S302 performs semantic-level deduplication on the question bank sets generated for each knowledge point under each question type to obtain a target question bank that does not contain questions with substantially the same or highly similar content.

[0058] Optionally, when the question generation module performs semantic-level deduplication on the question bank sets generated for each knowledge point under each question type to obtain a target question bank that does not contain questions with substantially the same or highly similar content, it is specifically used for: Construct a vector knowledge base, which is used to store the semantic feature vectors of questions in the target question bank after deduplication and retention, question by question; For each question in each question bank set, perform the following processing to obtain the deduplicated target question bank: Calculate the semantic feature vector of the problem; The semantic feature vector of the question is compared with the semantic feature vectors of each question already stored in the vector knowledge base. If the similarity between the semantic feature vector of a question and the semantic feature vector of any question in the vector knowledge base exceeds a preset threshold, and the question type of the question is the same as that of any question, then the question is deleted; otherwise, the question is retained in the target question bank, and the semantic feature vector of the question is stored in the vector knowledge base.

[0059] Optionally, after obtaining the target question bank, the device further includes: The verification and supplementation module is used to verify whether each question type corresponding to each training content outline meets the preset question quantity requirement based on the mapping relationship between each question in the target question bank and the training content outline and document fragments. If the number of questions of any question type under any training content outline does not meet the preset question quantity requirement, then for any training content outline, S301 to S302 are executed repeatedly to supplement and generate questions of the corresponding question type to the target question bank.

[0060] The railway document training question-generating device of this invention can execute the railway document training question-generating method provided in this invention. The implementation principle is similar. The actions performed by each module and unit in the railway document training question-generating device in each embodiment of this invention correspond to the steps in the railway document training question-generating method in each embodiment of this invention. For detailed functional descriptions of each module of the railway document training question-generating device, please refer to the descriptions in the corresponding railway document training question-generating methods shown above, which will not be repeated here.

[0061] The aforementioned railway document training question-generating device can be a computer program (including program code) running on a computer device, such as an application software; the device can be used to execute the corresponding steps in the method provided in the embodiments of the present invention.

[0062] In some embodiments, the railway document training question generation device provided in this invention can be implemented using a combination of hardware and software. As an example, the railway document training question generation device provided in this invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the railway document training question generation method provided in this invention. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0063] In other embodiments, the railway document training question generation device provided in this invention can be implemented in software. Figure 3 A railway document training question generation device stored in a memory is shown. It can be software in the form of programs and plug-ins, and includes a series of modules, including an acquisition module 210, a semantic association module 220, and a question generation module 230, for implementing the railway document training question generation method provided in the embodiments of the present invention.

[0064] The modules described in the embodiments of the present invention can be implemented in software or hardware. The names of the modules are not, in some cases, limiting the scope of the module itself.

[0065] Based on the same principles as the methods shown in the embodiments of the present invention, the embodiments of the present invention also provide an electronic device, which may include, but is not limited to: a processor and a memory; the memory for storing computer programs; and the processor for executing the methods shown in any embodiment of the present invention by invoking the computer programs.

[0066] In one alternative embodiment, an electronic device is provided, such as Figure 4 As shown, Figure 4 The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present invention.

[0067] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this invention. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0068] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0069] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.

[0070] The memory 4003 stores application code (computer program) for executing the present invention, and its execution is controlled by the processor 4001. The processor 4001 executes the application code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.

[0071] Among these, electronic devices can also be terminal devices. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0072] This invention provides a computer-readable storage medium storing a computer program that, when run on a computer, enables the computer to execute the corresponding content in the aforementioned method embodiments.

[0073] According to another aspect of the present invention, a computer program product or computer program is also provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various embodiments described above.

[0074] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0075] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0076] The computer-readable storage medium provided in this invention can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0077] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.

[0078] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.

Claims

1. A method for generating training questions for railway documents, characterized in that, include: S10, Obtain railway industry documents and training content outlines for various positions, and divide the railway industry documents into multiple document fragments; S20, Semantically correlate all document fragments and all training content outlines to obtain at least one document fragment that matches each training content outline; S30. Based on at least one document fragment matched to each training content outline, a target question bank corresponding to the railway industry document is obtained through a pre-trained question generation model. The question generation model is a generative large language model trained based on a historical question bank.

2. The method according to claim 1, characterized in that, The semantic association of all document fragments and all training content outlines, resulting in at least one document fragment matching each training content outline, includes: For each training content outline, perform semantic relevance analysis between the training content outline and each document fragment to obtain at least one document fragment that matches the training content outline; or, Extract key field identifiers for each document fragment. The key field identifiers are used to identify the core content of the corresponding document fragment. For each training content outline, perform semantic relevance analysis between the training content outline and each key field identifier to obtain at least one document fragment that matches the training content outline. or, Generative large models are used to associate each document fragment with each training content outline, resulting in at least one document fragment that matches each training content outline.

3. The method according to claim 1, characterized in that, The question generation model was trained using the following method: Obtain a historical question bank and corresponding historical question documents, wherein each historical question in the historical question bank is marked with a relevant document fragment from the historical question document; Extract the key fields of each question in the historical question bank. The key fields include at least the question type, question stem, options, correct answer, related document fragments and explanations. The related document fragments are used to indicate the specific clauses or paragraphs in the historical question-setting documents corresponding to the question. The key fields of each question are associated with relevant document fragments in the corresponding historical question documents to construct a training sample set. Each training sample in the training sample set includes a document fragment as input and the key fields of the question as output. The generative large language model is trained in a supervised manner using the training sample set to obtain the question generation model.

4. The method according to claim 1, characterized in that, Each training content outline includes at least one knowledge point; the process of obtaining a target question bank corresponding to the railway industry documents by matching at least one document fragment with each training content outline and using a pre-trained question generation model includes: S301, for each training content outline, based on at least one document fragment matched with the training content outline, the question set corresponding to each knowledge point in the training content outline under different question types is obtained through the question generation model; S302 performs semantic-level deduplication on the question bank sets generated for each knowledge point under each question type to obtain a target question bank that does not contain questions with substantially the same or highly similar content.

5. The method according to claim 4, characterized in that, The process involves semantically deduplicating the question bank sets generated for each knowledge point under each question type to obtain a target question bank that does not contain questions with substantially the same or highly similar content, including: Construct a vector knowledge base, which is used to store the semantic feature vectors of questions in the target question bank after deduplication and retention, question by question; For each question in each of the aforementioned question bank sets, perform the following processing to obtain the deduplicated target question bank: Calculate the semantic feature vector of the question; The semantic feature vector of the question is compared with the semantic feature vectors of each question already stored in the vector knowledge base; If the similarity between the semantic feature vector of the question and the semantic feature vector of any question in the vector knowledge base exceeds a preset threshold, and the question type of the question is the same as that of any question, then the question is deleted; otherwise, the question is retained in the target question bank, and the semantic feature vector of the question is stored in the vector knowledge base.

6. The method according to claim 4, characterized in that, After obtaining the target question bank, the method further includes: Based on the mapping relationship between each question in the target question bank and the training content outline and document fragments, check whether each question type corresponding to each training content outline meets the preset question quantity requirements; If the number of questions of any question type under any training content outline does not meet the preset question quantity requirement, then for any training content outline, S301 to S302 are executed repeatedly to supplement the target question bank with questions of the corresponding question type.

7. A railway document training question-generating device, characterized in that, include: The acquisition module is used to acquire railway industry documents and training content outlines for various positions, and to divide the railway industry documents into multiple document fragments; The semantic association module is used to perform semantic association on all document fragments and all training content outlines to obtain at least one document fragment that matches each training content outline. The question generation module is used to obtain the target question bank corresponding to the railway industry document by using a pre-trained question generation model based on at least one document fragment matched to each training content outline. The question generation model is a generative large language model trained based on the historical question bank.

8. The apparatus according to claim 7, characterized in that, When the semantic association module performs semantic association on all document fragments and all training content outlines to obtain at least one document fragment matching each training content outline, it is specifically used for: For each training content outline, perform semantic relevance analysis between the training content outline and each document fragment to obtain at least one document fragment that matches the training content outline; or, Extract key field identifiers for each document fragment. The key field identifiers are used to identify the core content of the corresponding document fragment. For each training content outline, perform semantic relevance analysis between the training content outline and each key field identifier to obtain at least one document fragment that matches the training content outline. or, Generative large models are used to associate each document fragment with each training content outline, resulting in at least one document fragment that matches each training content outline.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method of any one of claims 1-6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1-6.