Railway industry question and answer method and system based on Deepseek reinforcement learning

By fine-tuning the large language model in multiple stages using Deepseek's reinforcement learning method and combining it with railway industry datasets and knowledge vector libraries, we solved the problem of insufficient accuracy of the railway industry's question-answering system in understanding complex contexts and providing precise answers, and achieved logically clear question-answering capabilities.

CN120596631AInactive Publication Date: 2025-09-05武汉铁路职业技术学院
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510737902.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing large-model-based question-answering assistants in the railway industry have difficulty effectively understanding complex contexts and providing precise answers, resulting in insufficient accuracy.

Method used

Using the Deepseek reinforcement learning method, through multi-stage fine-tuning of the large language model, we constructed a railway industry instruction dataset and knowledge vector library. We combined the BM25 algorithm and the Railway-Embedding model for vector retrieval and full-text retrieval to generate logically clear answers.

Benefits of technology

It has improved the accuracy of the railway industry's question-and-answer system in understanding complex contexts and providing precise answers, and enhanced its understanding of industry terminology and ability to respond to long inquiries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596631A_ABST
    Figure CN120596631A_ABST
Patent Text Reader

Abstract

The invention discloses a railway industry question answering method and system based on Deepseek reinforcement learning, and relates to the technical field of natural language processing. The method comprises the following steps: constructing a railway industry instruction data set and a railway question and answer model; the railway question and answer model is obtained by taking a large language model as a base model and performing multi-stage fine tuning on the base model through a railway industry instruction data set; collecting railway industry text data, integrating and vectorizing to form a knowledge vector library; a Railway-Embedding model is used for conducting vector retrieval on the user question text vector, meanwhile, a BM25 algorithm is used for conducting full-text retrieval, and two kinds of retrieval results are sorted to obtain a target retrieval result; and inputting the user question and the target retrieval result into the railway question and answer model, and outputting the thinking process and the answer content of the question. According to the method, the large model is finely adjusted by fully utilizing the Deepseek reinforcement learning technology, the understanding of industry terminologies and the coping ability of long-term inquiry are improved, and meanwhile, the model has the thinking ability, so that the answer logic of the model is clearer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing technology, and in particular to a railway industry question-answering method based on Deepseek reinforcement learning. Background Art

[0002] The railway industry encompasses a wide range of safety management policies, regulations, experience sharing, technical specifications, and training materials, creating an urgent need for professional Q&A assistants to support efficient query and learning for industry personnel. The rapid development of large-scale language model (LLM) technology has significantly lowered the barrier to entry into the field of artificial intelligence, enabling more individuals and companies to enter the field. Open-source large-scale models are particularly popular in China, attracting significant attention and anticipation. They not only reduce high technical costs but also promote the widespread adoption and application of the technology. Many companies hope to integrate LLM capabilities into their products, enabling the implementation of intelligent applications through deep integration with business scenarios, thereby enhancing their market competitiveness. Building an industry knowledge base and establishing a professional large-scale model assistant system have become the most direct and effective application approaches.

[0003] In existing technologies, with the development of technology, question-answering assistants based on large models have emerged. However, these large model assistants still have limitations in understanding complex contexts and providing accurate answers, resulting in accuracy challenges in practical applications.

[0004] Therefore, there is an urgent need for a railway industry question-answering assistant system to improve the accuracy of the question-answering system in understanding complex contexts and providing precise answers. Summary of the Invention

[0005] Based on this, it is necessary to provide a railway industry question-answering method and system based on Deepseek reinforcement learning to address the above technical issues.

[0006] The present invention adopts the following technical solutions: Constructing a railway industry instruction dataset and a railway question-answering model; the railway question-answering model is obtained by using a large language model as a base model and performing multi-stage fine-tuning on the base model using the railway industry instruction dataset, including: using Deepseek R1 to distill the railway industry instruction dataset to generate reasoning data with thought chains, and performing a first-stage fine-tuning training on the base model to obtain the Railway-R1-SFT1 model; generating reinforced reasoning data using the Railway-R1-SFT1 model, and performing a second-stage fine-tuning training on the base model to obtain the Railway-R1-RL1 model; based on the Railway-R1-RL1 model, performing data distillation on the instruction dataset to generate non-reasoning task data and reasoning task data with thought chains, and performing a third-stage fine-tuning training on the base model to obtain the railway question-answering model; Collect various text data from the railway industry, integrate them and quantify them to obtain a railway industry knowledge vector library; Based on the railway industry knowledge vector library, the trained Railway-Embedding model is used to perform vector retrieval on the text vectors corresponding to the user's questions. The BM25 algorithm is then used to perform full-text retrieval to obtain vector retrieval results and full-text retrieval results. The vector retrieval results and full-text retrieval results are then sorted to obtain the target retrieval results. The user questions and target retrieval results are input into the trained railway question-answering model to obtain the thinking process and answer content corresponding to the user questions.

[0007] Preferably, constructing a railway industry instruction dataset specifically includes: Based on knowledge fragments from multiple dimensions of the railway industry, Deepseek R1 is used to extract railway questions and generate multiple answers corresponding to railway questions; Each railway question and its corresponding answer are input into the Deepseek R1 model configured with vLLM to obtain the scores of multiple answers and rank them; All railway questions and their top-ranked answers are constructed as an instruction dataset.

[0008] Preferably, Deepseek R1 is used to distill the railway industry instruction dataset to generate reasoning data with thought chains, and the first stage of fine-tuning training is performed on the base model to obtain the Railway-R1-SFT1 model, which specifically includes: Input the railway industry instruction dataset into Deepseek R1, set the controlled sampling parameters, and output the thinking chain reasoning data; Specific railway terminology tokens are added to the thinking chain reasoning data and input into the base model for supervised fine-tuning training to obtain the Railway-R1-SFT1 model.

[0009] Preferably, the Railway-R1-SFT1 model is used to generate enhanced inference data, and the base model is fine-tuned for the second stage to obtain the Railway-R1-RL1 model, specifically including: Input the railway industry instruction dataset into the Railway-R1-SFT1 model, prompt engineering through its thought chain, and output enhanced reasoning data; the enhanced reasoning data includes natural language instructions and decision paths for railway scheduling scenarios; The reinforcement inference data is input into the base model, the dynamic penalty parameters and gradient clipping parameters are set, the base model is subjected to reinforcement learning for group relative strategy optimization, and the base model is subjected to the second stage of fine-tuning training to obtain the Railway-R1-RL1 model.

[0010] Preferably, the base model is fine-tuned in the third stage to obtain a railway question-answering model, specifically including: The railway industry instruction dataset is input into the Railway-R1-RL1 model for data distillation, which outputs non-reasoning task data and reasoning task data with thought chains. The non-reasoning task data with thought chain is input into the base model for supervised fine-tuning to obtain the Railway-R1-SFT2 model; The reasoning task data with thought chains was input into the Railway-R1-SFT2 model, and reinforcement learning fine-tuning training with group relative strategy optimization was performed to obtain the railway question-answering model.

[0011] Preferably, performing vector retrieval on the text vector corresponding to the user question specifically includes: Based on the railway industry knowledge vector library, calculate the knowledge vector in the railway industry knowledge vector library whose similarity with the user question is greater than a preset threshold; The corresponding text content is extracted based on the searched knowledge vectors to obtain multiple vector retrieval results with different relevance indicators.

[0012] Preferably, the training process of the Railway-Embedding model specifically includes: Deepseek R1 is used to extract user questions from the railway industry knowledge vector library and generate questions similar to the user questions and questions that are not similar to the user questions. The format is: {"query": str, "pos": List[str], "neg":List[str], "prompt": str, "type": str}; Among them, query is the user question to be queried, pos is the positive text whose similarity with the user question to be queried meets the similarity threshold, neg is the negative text whose similarity with the user question to be queried meets the similarity threshold, prompt is used for the prompt of query to cover the query retrieval instruction; Generate difficult negative samples through the FlagEmbedding framework; The Railway-Embedding model is trained with user questions, questions similar to user questions, questions dissimilar to user questions, and difficult negative samples to obtain a trained Railway-Embedding model.

[0013] Preferably, the vector search results and the full-text search results are sorted to obtain the target search results and the target full-text search results, specifically including: The multiple vector search results and full-text search results of different relevance indicators are sorted for the first time through inverted fusion sorting, and constructed into a search result set; The Railway-Rerank model is used to perform secondary sorting on the search results in the search result set to obtain the target search results.

[0014] Preferably, the method further comprises: Configure the large language model environment and deploy OpenAI-style API services based on the large language model framework; After deployment is complete, the user questions and the answers output by the railway question-answering model are output in the following format: content = "Question:" + query + "\nThink first, then answer the user's question in the following format: <think>Reasoning content< / think> content".

[0015] The present invention provides a railway industry question-answering system based on Deepseek reinforcement learning, comprising: Railway industry instruction dataset construction module, used to construct railway industry instruction dataset; A railway question-answering model construction module, which uses a large language model as a base model and performs multi-stage fine-tuning on the base model using a railway industry instruction dataset to obtain a railway question-answering model; Railway industry knowledge vector library construction module, used to collect various text data of the railway industry, integrate and quantify them, and obtain the railway industry knowledge vector library; The question retrieval module is used to perform vector retrieval and full-text retrieval on the text vectors corresponding to user questions based on the railway industry knowledge vector library and the trained Railway-Embedding model, obtaining vector retrieval results and full-text retrieval results; and sorting the vector retrieval results and full-text retrieval results; The answer generation module is used to input user questions, vector retrieval results, and full-text retrieval results into the trained railway question-answering model to obtain the thinking process and answer content corresponding to the user questions.

[0016] The present invention provides a computer-readable storage medium, wherein the storage medium stores a computer program. When the computer program is executed by a processor, the computer program implements the above-mentioned railway industry question-answering method based on Deepseek reinforcement learning.

[0017] The present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for answering questions in the railway industry based on Deepseek reinforcement learning is implemented.

[0018] At least one of the above technical solutions adopted by the present invention can achieve the following beneficial effects: In the railway industry question-answering method based on Deepseek reinforcement learning provided by the present invention, a large language model is used as the base model, and the base model is fine-tuned in multiple stages using a constructed railway industry instruction dataset. The resulting railway question-answering model fully utilizes Deepseek reinforcement learning technology to fine-tune the large model, improving the understanding of industry professional terms and the ability to respond to long inquiries. At the same time, it enables the model to have thinking ability, thereby making the model's answer logic clearer. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0020] Figure 1 A flowchart of a railway industry question-answering method based on Deepseek reinforcement learning provided by the present invention; Figure 2 A schematic diagram of a model fine-tuning method for the railway industry question-answering method based on Deepseek reinforcement learning provided by the present invention; Figure 3 A schematic diagram of a railway industry question-answering system based on Deepseek reinforcement learning provided by the present invention; Figure 4 A diagram of computer equipment for implementing a railway industry question-answering method based on Deepseek reinforcement learning, provided by the present invention. DETAILED DESCRIPTION

[0021] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of the present invention and corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in the specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0022] The technical solutions provided by the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0023] Figure 1 This is a flow chart of a railway industry question-answering method based on Deepseek reinforcement learning in the present invention, which specifically includes: S101: Construct a railway industry instruction dataset and a railway question-answering model; the railway question-answering model is based on a large language model and is obtained by multi-stage fine-tuning the base model using the railway industry instruction dataset.

[0024] Optionally, the base model is fine-tuned in multiple stages using the railway industry instruction dataset, including: using the DeepseekR1 model to distill the railway industry instruction dataset to generate reasoning data with thought chains, and performing the first stage of fine-tuning training on the base model to obtain the Railway-R1-SFT1 model; generating reinforced reasoning data through the Railway-R1-SFT1 model, and performing the second stage of fine-tuning training on the base model to obtain the Railway-R1-RL1 model; based on the Railway-R1-RL1 model, performing data distillation on the instruction dataset to generate non-reasoning task data and reasoning task data with thought chains, and performing the third stage of fine-tuning training on the base model to obtain the railway question-answering model.

[0025] Optionally, constructing instruction set data includes: extracting railway questions and generating multiple answers corresponding to the railway questions through Deepseek R1 based on knowledge fragments of multiple dimensions of the railway industry; inputting each railway question and its corresponding answer into a Deepseek R1 model configured with a vLLM, obtaining scores for the multiple answers and ranking them; and constructing all railway questions and their highest-ranked answers into an instruction data set.

[0026] Specifically, based on the compiled industry knowledge base, questions were extracted using Deepseek R1, organized into knowledge segments categorized by railway industry domain, knowledge type, and application scenario. For each knowledge segment, 50 questions were generated, covering common questions and key information points. These questions were then cleaned and expanded based on feedback from domain experts. Answers were then generated using Deepseek R1, following the recommended parameter settings in the DeepSeek technical report. The following instruction was added before the prompt: "Please think and reason step by step, and place your final answer within \boxed{}." To ensure efficient generation, the invention limited the number of tokens generated to 8k (testing found that 80% of questions could be solved within 8k tokens, while the remaining questions required 16k tokens). To provide greater flexibility in subsequent filtering and optimization, the invention generated two answers for each question (some questions even generated four answers). This ultimately replicated a method similar to DeepSeek R1's ability to reject samples, making the dataset suitable for preference optimization methods such as Deterministic Profiling (DPO). In addition, to ensure that the dataset contains only high-quality and correct reasoning results, the present invention conducts multiple rounds of review and optimization on the generated answers. Under the guidance of experts, accurate answers are screened out, and unclear or ambiguous content is re-edited. Based on the above requirements, a Railway-verify tool is built to review and optimize the generated answers. For those data rows containing multiple correct answers, the present invention attempts to use the reward model (RM) as the final filter to select the best answer. The specific operation is as follows: First, from each data row containing multiple correct answers, remove ( <think> …< / think> ), extract the final answer; second, input the question and the extracted answer into the Deepseek R1 model configured with vLLM to obtain the score of each answer; then, according to the model score, each data row containing multiple correct answers is ranked, and the highest-ranked answer is selected to be included in the instruction dataset.

[0027] Optionally, the railway industry instruction dataset is distilled using the Deepseek R1 model to generate reasoning data with COT thinking, and then the base model is subjected to the first stage of SFT fine-tuning training to obtain the Railway-R1-SFT1 model, specifically including: inputting the railway industry instruction dataset into the Deepseek R1 model, setting controlled sampling parameters, outputting the thinking chain reasoning data, adding specific railway terminology tokens to the thinking chain reasoning data, and inputting it into the base model for supervised fine-tuning training to obtain the Railway-R1-SFT1 model.

[0028] Specifically, based on a dataset of railway industry instructions (after desensitization) compiled by this invention, the Deepseek R1 model was used to generate output with an inference chain through controlled sampling. A tagging template, "Analyze requirements → Retrieve knowledge → Step-by-step verification," was added to each instruction. The base model was then fine-tuned using SFT. Note that during fine-tuning, this invention added special tokens for railway terminology (e.g., <turnout>, <catenary>). A dynamic masking strategy was used to retain industry keywords. Fine-tuning based on this data and strategy resulted in the Railway-R1-SFT1 model.

[0029] Optionally, reinforced reasoning data is generated through the Railway-R1-SFT1 model, and the base model is subjected to second-stage fine-tuning training to obtain the Railway-R1-RL1 model, specifically including: inputting the railway industry instruction data set into the Railway-R1-SFT1 model, prompting the engineering through its thinking chain, and outputting reinforced reasoning data; inputting the reinforced reasoning data into the base model, setting dynamic penalty parameters and gradient clipping parameters, performing reinforcement learning of group relative strategy optimization on the base model, and performing second-stage fine-tuning training on the base model to obtain the Railway-R1-RL1 model; the reinforced reasoning data includes: natural language instructions and decision paths for railway scheduling scenarios.

[0030] Specifically, based on the Railway-R1-SFT1 model, the Chain of Thought (CoT) prompts the engineering to generate multi-step reasoning trajectories. Each data contains: natural language instructions for railway scheduling scenarios (such as "adjust the T45 train to alleviate congestion at Zhengzhou North Station"); decision paths generated by the model (including alternative plans, constraint analysis, etc.) (3) Finally, reasoning data is generated; then, the generated reasoning data is optimized using GRPO reinforcement learning, introducing: dynamic KL penalty: initial coefficient 0.05, automatically adjusted according to policy deviation; gradient clipping: using a dual threshold mechanism (parameter update gradient limited to [-1.5, 1.5], value function gradient limited to [-0.8, 0.8]), based on the above data and strategy, fine-tuning is performed to obtain the Railway-R1-RL1 model.

[0031] Optionally, the base model is subjected to a third stage of fine-tuning training to obtain a railway question-answering model, including: inputting the railway industry instruction dataset into the Railway-R1-RL1 model for data distillation, and outputting non-reasoning task data and reasoning task data with thought chains; inputting the non-reasoning task data with thought chains into the base model for supervised fine-tuning to obtain the Railway-R1-SFT2 model; inputting the reasoning task data with thought chains into the Railway-R1-SFT2 model, performing reinforcement learning fine-tuning training for group relative strategy optimization, and obtaining the railway question-answering model.

[0032] Specifically, based on the Railway-R1-RL1 model, data distillation is performed on the instruction dataset to generate data containing COT non-reasoning tasks and reasoning tasks. For COT non-reasoning task data generation, the "task prefix + instruction" format is used, and tasks such as classification / summarization are converted into chain steps through a template method. For example, the classification task will be reconstructed as: "Identify the category keywords in the sentence → Comprehensively judge the text type → Final answer"; For reasoning task data construction, the following strategies are adopted: (1) Counterfactual enhancement: Artificially construct error examples that violate the original reasoning path; Multiple solution paths: Annotate at least two different reasoning processes for the same problem. Finally, non-reasoning data and reasoning data with COT are obtained; Based on the non-reasoning data and reasoning data generated above, fine-tuning is performed as follows: SFT fine-tuning is performed on the COT non-reasoning task data to obtain the Railway-R1-SFT2 model; GRPO reinforcement learning with reward verification is performed on the Railway-R1-SFT2 model on the reasoning task data. The third stage of fine-tuning training is formed through multiple stages of fine-tuning training, and the Railway-R1 model is finally obtained.

[0033] S102: Collect various text data of the railway industry, integrate and quantify them, and obtain a railway industry knowledge vector library.

[0034] Specifically, the industry knowledge base will be constructed, with the following content and objectives: Clarify objectives and areas: Build a comprehensive, accurate, and user-friendly railway industry knowledge base to serve users such as railway practitioners, researchers, and students; improve the efficiency of railway industry knowledge acquisition, promote knowledge sharing and innovation, and provide knowledge support for the intelligent development of the railway industry. Areas include: rolling stock, communications and signaling, civil engineering, traction power supply, transport organization, safety assurance, and more. In addition, there are areas such as railway economics, railway regulations, railway history, and railway culture.

[0035] Specifically, information is collected and organized from the following sources: (a) Internal data: technical documents, research reports, case studies, and fault records from railway bureaus, stations, research institutes, and other institutions; (b) External data: publicly available industry standards, patents, academic papers, conference reports, and expert experience; and (c) Internet data: industry websites, forums, blogs, and social media. Information organization is performed through the following steps: (a) Data cleaning: Removing duplicate, erroneous, and incomplete data; (b) Data classification: Categorizing and organizing data according to a knowledge framework; (c) Data annotation: Annotating data with keywords, topics, and sources to facilitate retrieval and analysis; and (d) Data updating: Regularly reviewing and updating the knowledge base. This ensures the timeliness and accuracy of information. Through this series of steps, a dynamic, multi-dimensional, and practical railway industry knowledge base is established, providing continuous knowledge support and driving innovation for the railway industry. Furthermore, intelligent optimization of the knowledge base enhances the user experience and enables fast and accurate information retrieval and question-and-answer services.

[0036] Specifically, a knowledge framework is established, including: (a) Knowledge representation: using ontology, semantic network, knowledge graph and other technologies to represent railway industry knowledge and build the relationship between knowledge; (b) Knowledge classification: classifying knowledge according to the dimensions of railway industry professional fields, knowledge types, application scenarios, etc.; (c) Knowledge system: building a railway industry knowledge system with clear hierarchy, structure and content.

[0037] S103: Based on the railway industry knowledge vector library, the trained Railway-Embedding model is used to perform vector retrieval on the text vector corresponding to the user question, and the BM25 algorithm is used to perform full-text retrieval to obtain vector retrieval results and full-text retrieval results. The vector retrieval results and full-text retrieval results are then sorted to obtain the target retrieval results.

[0038] Optionally, based on the railway industry knowledge vector library, a vector search is performed on the text vector corresponding to the user question, specifically including: based on the railway industry knowledge vector library, calculating the knowledge vector in the railway industry knowledge vector library whose similarity with the user question is greater than a preset threshold; extracting the corresponding text content based on the searched knowledge vector to obtain multiple vector search results with different relevance indicators.

[0039] Specifically, the ANN nearest neighbor similarity retrieval algorithm is used to calculate the knowledge vectors in the railway industry knowledge vector library that are similar to the user's question (the threshold is set to 0.85 after experimental determination), and the corresponding text content is extracted to obtain multiple vector retrieval results with different relevance indicators.

[0040] Optionally, obtaining vector retrieval results and full-text retrieval results specifically includes: performing a first sorting on the vector retrieval results and the full-text retrieval results by an inverted fusion sorting method, combining multiple retrieval results with different relevance indicators into a single retrieval result set; inputting the single retrieval result set into a Railway-Rerank model to perform a second sorting on the retrieval results, and obtaining the final vector retrieval results and the full-text retrieval results.

[0041] Optionally, the training process of the Railway-Embedding model specifically includes: Deepseek R1 is used to extract user questions from the railway industry knowledge vector library and generate questions similar to and dissimilar to the user questions. The format is: {"query": str, "pos": List[str], "neg":List[str],"prompt": str, "type": str}; where query is the user question to be queried, pos is the positive text whose similarity to the user question to be queried meets the similarity threshold, neg is the negative text whose similarity to the user question to be queried meets the similarity threshold, and prompt is used as the query prompt to cover the query retrieval instruction; difficult negative samples are generated through the FlagEmbedding framework; user questions, questions similar to and dissimilar to the user questions, and difficult negative samples are used to train the Railway-Embedding model to obtain the trained Railway-Embedding model.

[0042] Specifically, when a user asks a question, the Elasticsearch vector database is searched for relevant text content based on the user's question. When a user initiates a conversation, the system also inputs the user's conversation into the Railway-Embedding model to generate a vector. This vector is then placed in the vector database and the existing policy for querying. Based on the Elasticsearch framework, an ANN nearest neighbor similarity search algorithm (to improve search efficiency) is used to calculate the most similar domain knowledge vector and extract the text content of the corresponding vector. The present invention also uses BM25 for full-text search and uses the Reciprocal Rank Fusion (RRF) method to sort the two search results (full-text search and vector search) to combine multiple result sets with different relevance metrics into a single result set. Since RRF does not require tuning, different relevance metrics do not need to be correlated to obtain high-quality results. This method's advantage is that it does not utilize relevance scores, relying solely on ranking calculations. Furthermore, the Railway-Rerank model is used to perform secondary ranking of search results to optimize search accuracy.

[0043] Optionally, the vector retrieval results and the full-text retrieval results are sorted, specifically including: performing a first sorting of multiple vector retrieval results and full-text retrieval results of different relevance indicators by inverted fusion sorting, and constructing them into a retrieval result set; performing a second sorting of the retrieval results in the retrieval result set by the Railway-Rerank model to obtain the target retrieval results.

[0044] S104: Input the user question, vector search results, and full-text search results into the trained railway question-answering model to obtain the thinking process and answer content corresponding to the user question.

[0045] Optionally, configure a large language model environment and deploy an OpenAI-style API service based on the large language model framework. After deployment, output the user's question and the answer output by the railway question-answering model. The output format is: content = "Question: " + query + "\nThink first, then answer the user's question in the following format: <think> Reasoning content< / think> content".

[0046] The application deployment module 105 is used to deploy the OpenAI-based API service and generate a template for outputting user questions, thinking processes, and answer content.

[0047] Specifically, the fine-tuned model is deployed based on the vllm framework. vLLM (Very Large Language Model Serving) is a high-performance, low-latency large language model (LLM) reasoning and serving framework. It is designed for large-scale production-level deployment and is particularly good at handling ultra-long contexts (such as 8k+ tokens) and high-concurrency requests, while significantly optimizing graphics memory utilization. It has the following advantages: high throughput: using advanced server throughput technology; memory management: efficient management of attention key and value memory through PagedAttention; request batching: supporting continuous batching of incoming requests; model execution: using CUDA / HIP graph to achieve fast model execution; quantization technology: supporting GPTQ, AWQ, INT4, INT8 and FP8 quantization; optimized kernel: including integration with FlashAttention and FlashInfer; other features: supporting speculative decoding and block pre-filling. The specific process includes: (1) configuring the vllm-related environment, including NVIDIA drivers, CUDA, cuddnn, pytorch, etc. (2) deploying an OpenAI-style API service based on the vllm framework. Pay attention to setting parameters such as max-model-len, max-num-batched-tokens, tensor-parallel-size, and max-num-seqs. After deployment, use the following template for output: content = "Question: " + query + "\nThink first, then answer the user's question in the following format: <think> Reasoning content< / think> content".

[0048] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present invention.

[0049] The above is a railway industry question-answering method based on Deepseek reinforcement learning provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding railway industry question-answering device based on Deepseek reinforcement learning, such as Figure 3 shown.

[0050] Figure 3 A schematic diagram of a railway industry question-answering system based on Deepseek reinforcement learning provided by the present invention includes: Railway industry instruction data set construction module 301, used to construct a railway industry instruction data set; Railway question-answering model construction module 302, configured to use the large language model as a base model and perform multi-stage fine-tuning on the base model using a railway industry instruction dataset to obtain a railway question-answering model; Railway industry knowledge vector library building module 303, used to collect various text data of the railway industry, integrate and quantify them, and obtain the railway industry knowledge vector library; Question retrieval module 304 is used to perform vector retrieval and full-text retrieval on the text vector corresponding to the user question based on the railway industry knowledge vector library using the trained Railway-Embedding model to obtain vector retrieval results and full-text retrieval results; and sort the vector retrieval results and full-text retrieval results; The answer generation module 305 is used to input the user question, vector search results and full-text search results into the trained railway question-answering model to obtain the thinking process and answer content corresponding to the user question.

[0051] Regarding the specific definition of a railway industry question-answering device based on Deepseek reinforcement learning, please refer to the definition of a railway industry question-answering method based on Deepseek reinforcement learning above, which will not be repeated here. The various modules in the above-mentioned railway industry question-answering device based on Deepseek reinforcement learning can be implemented in whole or in part through software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0052] The present invention also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 A railway industry question-answering method based on Deepseek reinforcement learning is provided.

[0053] The present invention also provides Figure 3 The structural diagram of the computer equipment shown in FIG. Figure 3 As mentioned above, at the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 A railway industry question-answering method based on Deepseek reinforcement learning is provided.

[0054] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media used in the embodiments provided herein may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

Claims

1. A railway industry question answering method based on Deepseek reinforcement learning, characterized by: include: Build a railway industry instruction dataset and railway question-answering model; The railway question-answering model is based on a large language model and is obtained by multi-stage fine-tuning the base model using a railway industry instruction dataset. The method includes: using DeepseekR1 to distill the railway industry instruction dataset to generate reasoning data with thought chains, and performing a first-stage fine-tuning training on the base model to obtain the Railway-R1-SFT1 model; generating reinforced reasoning data using the Railway-R1-SFT1 model, and performing a second-stage fine-tuning training on the base model to obtain the Railway-R1-RL1 model; based on the Railway-R1-RL1 model, performing data distillation on the instruction dataset to generate non-reasoning task data and reasoning task data with thought chains, and performing a third-stage fine-tuning training on the base model to obtain the railway question-answering model. Collect various text data from the railway industry, integrate them and quantify them to obtain a railway industry knowledge vector library; Based on the railway industry knowledge vector library, the trained Railway-Embedding model is used to perform vector retrieval on the text vectors corresponding to the user's questions. The BM25 algorithm is then used to perform full-text retrieval to obtain vector retrieval results and full-text retrieval results. The vector retrieval results and full-text retrieval results are then sorted to obtain the target retrieval results. The user questions and target retrieval results are input into the trained railway question-answering model to obtain the thinking process and answer content corresponding to the user questions.

2. The railway industry question-answering method based on Deepseek reinforcement learning according to claim 1, characterized in that: The construction of the railway industry instruction dataset specifically includes: Based on knowledge fragments from multiple dimensions of the railway industry, Deepseek R1 is used to extract railway questions and generate multiple answers corresponding to railway questions; Each railway question and its corresponding answer are input into the Deepseek R1 model configured with vLLM to obtain the scores of multiple answers and rank them; All railway questions and their top-ranked answers are constructed as an instruction dataset.

3. The railway industry question-answering method based on Deepseek reinforcement learning according to claim 1, characterized in that: The railway industry instruction dataset is distilled using Deepseek R1 to generate reasoning data with thought chains, and the base model is fine-tuned for the first phase to obtain the Railway-R1-SFT1 model, which specifically includes: Input the railway industry instruction dataset into Deepseek R1, set the controlled sampling parameters, and output the thinking chain reasoning data; Specific railway terminology tokens are added to the thinking chain reasoning data and input into the base model for supervised fine-tuning training to obtain the Railway-R1-SFT1 model.

4. The railway industry question-answering method based on Deepseek reinforcement learning according to claim 1, characterized in that: The Railway-R1-SFT1 model is used to generate enhanced inference data, and the base model is fine-tuned for the second phase to obtain the Railway-R1-RL1 model, which specifically includes: Input the railway industry instruction dataset into the Railway-R1-SFT1 model, prompt engineering through its thought chain, and output enhanced reasoning data; the enhanced reasoning data includes natural language instructions and decision paths for railway scheduling scenarios; The reinforcement inference data is input into the base model, the dynamic penalty parameters and gradient clipping parameters are set, the base model is subjected to reinforcement learning for group relative strategy optimization, and the base model is subjected to the second stage of fine-tuning training to obtain the Railway-R1-RL1 model.

5. The railway industry question-answering method based on Deepseek reinforcement learning according to claim 1, characterized in that: The third stage of fine-tuning training of the base model to obtain the railway question-answering model specifically includes: The railway industry instruction dataset is input into the Railway-R1-RL1 model for data distillation, which outputs non-reasoning task data and reasoning task data with thought chains. The non-reasoning task data with thought chain is input into the base model for supervised fine-tuning to obtain the Railway-R1-SFT2 model; The reasoning task data with thought chains was input into the Railway-R1-SFT2 model, and reinforcement learning fine-tuning training with group relative strategy optimization was performed to obtain the railway question-answering model.

6. The railway industry question-answering method based on Deepseek reinforcement learning according to claim 1, characterized in that: The vector retrieval of the text vector corresponding to the user question specifically includes: Based on the railway industry knowledge vector library, calculate the knowledge vector in the railway industry knowledge vector library whose similarity with the user question is greater than a preset threshold; The corresponding text content is extracted based on the searched knowledge vectors to obtain multiple vector retrieval results with different relevance indicators.

7. The railway industry question-answering method based on Deepseek reinforcement learning according to claim 1, characterized in that: The training process of the Railway-Embedding model specifically includes: Deepseek R1 is used to extract user questions from the railway industry knowledge vector library and generate questions similar to the user questions and questions that are not similar to the user questions. The format is: {"query": str, "pos": List[str], "neg":List[str], "prompt": str, "type":str}; Among them, query is the user question to be queried, pos is the positive text whose similarity with the user question to be queried meets the similarity threshold, neg is the negative text whose similarity with the user question to be queried meets the similarity threshold, prompt is used for the prompt of query to cover the query retrieval instruction; Generate difficult negative samples through the FlagEmbedding framework; The Railway-Embedding model is trained with user questions, questions similar to user questions, questions dissimilar to user questions, and difficult negative samples to obtain a trained Railway-Embedding model.

8. The railway industry question-answering method based on Deepseek reinforcement learning according to claim 1, characterized in that: The vector search results and the full-text search results are sorted to obtain the target search results and the target full-text search results, specifically including: The multiple vector search results and full-text search results of different relevance indicators are sorted for the first time through inverted fusion sorting, and constructed into a search result set; The Railway-Rerank model is used to perform secondary sorting on the search results in the search result set to obtain the target search results.

9. The railway industry question-answering method based on Deepseek reinforcement learning according to claim 1, characterized in that: The method further comprises: Configure the large language model environment and deploy OpenAI-style API services based on the large language model framework; After deployment is complete, the user questions and the answers output by the railway question-answering model are output in the following format: content = "Question:" + query + "\nThink first, then answer the user's question in the following format: <think>Reasoning content< / think> content".

10. A railway industry question-answering system based on Deepseek reinforcement learning, characterized by: include: Railway industry instruction dataset construction module, used to construct railway industry instruction dataset; A railway question-answering model construction module, which uses a large language model as a base model and performs multi-stage fine-tuning on the base model using a railway industry instruction dataset to obtain a railway question-answering model; Railway industry knowledge vector library construction module, used to collect various text data of the railway industry, integrate and quantify them, and obtain the railway industry knowledge vector library; The question retrieval module is used to perform vector retrieval and full-text retrieval on the text vectors corresponding to user questions based on the railway industry knowledge vector library and the trained Railway-Embedding model, obtaining vector retrieval results and full-text retrieval results; and sorting the vector retrieval results and full-text retrieval results; The answer generation module is used to input user questions, vector retrieval results, and full-text retrieval results into the trained railway question-answering model to obtain the thinking process and answer content corresponding to the user questions.

Citation Information

Cited By

  • Table data analysis large model training and application method based on agent interaction reinforcement learning

    CN121478940A

  • AI intelligent aid decision-making model construction method for advanced intervention work of railway construction

    CN122088711A