A document question-answering method, apparatus, electronic device, and computer storage medium.

By fine-tuning the model based on the document organization structure and semantic expression in the railway field, high-quality sample fragments are generated and the document retrieval and ranking model is optimized. This solves the problems of poor relevance and comprehensiveness of answers in the railway field of traditional question answering systems, realizes the generation of technical questions, and improves the effectiveness and usability of question answering systems.

CN120821819BActive Publication Date: 2025-12-02BEIJING TRAFFIC & TRANSPORT TECH CORP LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511340391.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-12-02
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

Traditional question-answering systems in the railway sector suffer from problems such as poor relevance of answers, incomplete answers, and answer illusions due to the lack of model training and fine-tuning for document organization and semantic expression, which affects efficiency and cost.

Method used

By acquiring sample documents in the railway field, aligning them to generate sample fragments, and combining semantic relevance annotations to fine-tune the document retrieval model and semantic ranking model, a generative large model is generated to optimize the performance of the question answering system.

Benefits of technology

It improves the accuracy and usability of the question-and-answer system in the railway sector, reduces user usage and verification costs, increases work efficiency, and meets the needs of the railway industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821819B_ABST
    Figure CN120821819B_ABST
Patent Text Reader

Abstract

This invention relates to a document question-answering method, apparatus, electronic device, and computer storage medium. The method includes: aligning each sample document based on a preset segment length to obtain multiple sample segments; semantically labeling any two sample segments to obtain first labeled data; fine-tuning a document retrieval model based on the first labeled data to obtain a target document retrieval model; determining second labeled data based on multiple sample segments and semantic relevance; fine-tuning a semantic ranking model based on the second labeled data to obtain a target semantic ranking model; constructing fourth labeled data based on multiple questions and at least one sample segment corresponding to each question; fine-tuning a basic large model based on the fourth labeled data to obtain a generative large model; and determining the target answer corresponding to the question to be processed based on the three fine-tuned models. This invention fundamentally solves the problems frequently encountered in general large-model question answering on railway domain datasets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing technology, and more specifically, to a document question-answering method, apparatus, electronic device, and computer storage medium. Background Technology

[0002] The railway industry has a large number of non-public documents and knowledge bases. With the development of generative artificial intelligence, non-public document question answering is usually implemented using a retrieval-enhancement-generation technology framework (RAG framework). However, railway documents usually have specific document organization structures and semantic expressions. Due to the non-public nature of documents and data, traditional question answering systems use open-source or self-developed RAG frameworks. The core modules and models (such as segmentation-retrieval-ranking models) have not been trained and fine-tuned for the railway industry. Therefore, document question answering systems implemented directly using the RAG framework are prone to problems such as poor answer relevance, incomplete answers, and answer illusions, resulting in incorrect answer results, incorrect document citations, and other obvious quality and effectiveness issues. This increases the actual use and verification costs, and the overall usability of the question answering system is low, which in turn affects the work efficiency of users and increases the cost of troubleshooting and locating actual railway problems. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a document question-and-answer method, apparatus, electronic device and computer storage medium, which aims to solve at least one of the above-mentioned technical problems.

[0004] Firstly, the technical solution of the present invention to solve the above-mentioned technical problems is as follows: a document question-and-answer method, the method comprising:

[0005] Obtain sample documents in different document formats in the railway field, and align each sample document based on a preset segment length to obtain multiple sample segments;

[0006] Semantic relevance is labeled for any two sample segments to obtain the first labeled data. The document retrieval model is then fine-tuned based on the first labeled data to obtain the target document retrieval model.

[0007] Based on multiple sample fragments and the semantic relevance between each sample fragment, the second labeled data is determined, and the semantic ranking model is fine-tuned according to the second labeled data to obtain the target semantic ranking model. The second labeled data includes multiple questions and at least one sample fragment corresponding to each question.

[0008] Based on multiple questions and at least one sample fragment corresponding to each question, question-answer pairs are generated through a basic large model.

[0009] Based on question-answer pairs and third-labeled data, fourth-labeled data is constructed. The basic large model is then fine-tuned based on the fourth-labeled data to obtain a generative large model. The third-labeled data includes instructions, questions, and answers generated for the railway sector.

[0010] The target document retrieval model, the target semantic ranking model, and the generative large model are used to process the problem and obtain the target answer corresponding to the problem.

[0011] The beneficial effects of this invention are: by fine-tuning the model to address the unique characteristics of the document organization structure and semantic expression in the railway field, the accuracy and usability of the document question-answering system can be effectively improved. Specifically, by acquiring sample documents from the railway field and aligning them, high-quality sample fragments are generated. Then, the document retrieval model and semantic ranking model are fine-tuned in conjunction with semantic relevance annotation, enabling the model to better understand and process the document content in the railway field. Simultaneously, the basic large-scale model is fine-tuned based on the generated question-answer pairs and annotated data, further optimizing the performance of the generative large-scale model. Finally, by combining the target document retrieval model, the target semantic ranking model, and the generative large-scale model to process the problem, more accurate and comprehensive target answers can be generated. This method effectively solves the problems of poor answer relevance, incomplete answers, and answer illusions that occur in traditional question-answering systems for railway field documents, thereby reducing user usage and verification costs, improving user efficiency, reducing the cost of investigation and location when actual railway problems occur, and better meeting the needs of the railway industry for document question-answering systems.

[0012] Based on the above technical solution, the present invention can be further improved as follows.

[0013] Furthermore, the above-mentioned target document retrieval model, target semantic ranking model, and generative large model are used to process the problem to obtain the target answer corresponding to the problem, including:

[0014] Retrieve pending issues, which are specific to the railway sector;

[0015] Identify the multiple document fragments corresponding to the problem to be addressed;

[0016] The semantic relevance between document fragments is calculated using the target document retrieval model.

[0017] Based on the semantic relevance between document fragments, all document fragments are ranked using a target semantic ranking model;

[0018] Based on all sorted document fragments, a generative large model is used to determine the target answer to the question to be processed.

[0019] The beneficial effect of adopting the above-mentioned further approach is that it utilizes a target document retrieval model, a target semantic ranking model, and a generative large model to handle problems in the railway field, generating high-quality target answers. This process not only improves the accuracy and reliability of the question-answering system but also ensures the comprehensiveness and professionalism of the generated answers through semantic relevance calculation and ranking optimization.

[0020] Furthermore, the semantic relevance annotation of any two sample segments obtained above yields the first annotation data, including:

[0021] Convert each sample fragment into a vector of a specified dimension;

[0022] Determine the semantic relevance between any two vectors;

[0023] Based on the semantic relevance between various vectors, semantic relevance is labeled for any two sample segments to obtain the first labeled data.

[0024] The beneficial effect of adopting the above-mentioned further approach is that converting sample fragments into vectors of a set dimension and calculating the semantic relevance between vectors transforms the semantic information of the text into quantifiable numerical values, thereby avoiding the subjective errors of manual annotation and significantly improving the accuracy of semantic relevance annotation. The document retrieval model, fine-tuned based on this annotated data, can more accurately identify and retrieve document fragments related to the question, thus providing a more reliable basis for generating high-quality answers.

[0025] Furthermore, based on a preset fragment length, each sample document is aligned to obtain multiple sample fragments, including:

[0026] Each document in all sample documents that is longer than the preset segment length is segmented to obtain multiple first segments;

[0027] For each document in all sample documents that is shorter than the preset segment length, the document length is extended by using contextual paragraph linking and increasing the repetition coefficient before and after the document to obtain multiple second segments;

[0028] Multiple sample segments are obtained based on multiple first segments and multiple second segments.

[0029] The beneficial effects of adopting the above-mentioned further scheme are that, for documents longer than the preset segment length, they are split into multiple first segments that meet the length requirements; for documents shorter than the preset length, they are expanded by using contextual paragraph linking relationships and increasing the repetition coefficient to generate second segments that meet the length requirements. This strategy not only solves the problem of inconsistent paragraph block lengths in railway documents, but also enhances the semantic richness and coherence of document segments through the introduction of contextual linking relationships and overlapping content, enabling the generated sample segments to better preserve the semantic information of the original text. This alignment processing method provides high-quality input data for subsequent model training and question answer generation, thereby significantly improving the retrieval recall effect and answer generation accuracy of the question answering system, further enhancing the applicability and practicality of the question answering system in the railway field, and reducing the question answering error rate caused by improper document segment processing.

[0030] Furthermore, based on multiple sample fragments and the semantic relevance between them, the second labeled data is determined, including:

[0031] Multiple questions are generated based on multiple sample fragments and the semantic relevance between them.

[0032] For each question, determine the question relevance between the question and each sample segment;

[0033] For each problem, based on the first target segment and the second target segment corresponding to the problem, labeled data corresponding to the problem is generated. The first target segment is any sample segment whose problem relevance is greater than a first set value, and the second target segment is any sample segment whose problem relevance is less than a second set value.

[0034] Based on the labeled data corresponding to all questions, the second labeled data is determined.

[0035] The beneficial effect of adopting the above-mentioned further approach is that it generates questions based on multiple sample fragments and their semantic relevance, and determines the relevance of each question to the questions in each sample fragment. By selecting fragments with high and low semantic relevance as labeled data, the model can be effectively trained to distinguish document fragments with different semantic relevance, thereby improving the accuracy of the model in ranking document fragments in real-world question-answering scenarios.

[0036] Furthermore, the question types for multiple questions include parameter extraction type, clause listing type, numerical calculation type, and artificial scenario type. The second labeled data includes multiple questions corresponding to each question type and at least one sample fragment corresponding to each question.

[0037] The beneficial effect of adopting the above-mentioned further scheme is that the method also covers a variety of problem types (such as parameter extraction type, clause listing type, numerical calculation type, and artificial scenario type), further enriching the diversity of labeled data, enabling the model to better adapt to the complex and diverse problem forms in the railway field, and thus generating high-quality answers in different scenarios.

[0038] Secondly, in order to solve the above-mentioned technical problems, the present invention also provides a document question-and-answer device, the device comprising:

[0039] The acquisition module is used to acquire sample documents in different document formats in the railway field, and align each sample document based on a preset segment length to obtain multiple sample segments.

[0040] The document retrieval model fine-tuning module is used to perform semantic relevance annotation on any two sample segments to obtain the first annotation data, and to fine-tune the document retrieval model based on the first annotation data to obtain the target document retrieval model;

[0041] The semantic ranking model fine-tuning module is used to determine the second labeled data based on multiple sample fragments and the semantic relevance between each sample fragment, and to fine-tune the semantic ranking model according to the second labeled data to obtain the target semantic ranking model. The second labeled data includes multiple questions and at least one sample fragment corresponding to each question.

[0042] The question-answer pair generation module is used to generate question-answer pairs based on multiple questions and at least one sample fragment corresponding to each question, using a basic large model.

[0043] The basic large model fine-tuning module is used to construct fourth labeled data based on question-answer pairs and third labeled data, and to fine-tune the basic large model based on the fourth labeled data to obtain a generative large model. The third labeled data includes instructions, questions and answers generated for the railway field.

[0044] The processing module is used to process the problem to be processed based on the target document retrieval model, the target semantic ranking model, and the generative large model to obtain the target answer corresponding to the problem to be processed.

[0045] Thirdly, in order to solve the above-mentioned technical problems, the present invention also provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the document question-and-answer method of the present application.

[0046] Fourthly, in order to solve the above-mentioned technical problems, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the document question-and-answer method of the present application.

[0047] Additional aspects and advantages of this application will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of this application. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments of the present invention will be briefly introduced below.

[0049] Figure 1 A flowchart illustrating a document question-and-answer method according to an embodiment of the present invention;

[0050] Figure 2 A flowchart illustrating yet another document question-and-answer method provided in an embodiment of the present invention;

[0051] Figure 3 This is a schematic diagram of the structure of a document question-and-answer device according to an embodiment of the present invention;

[0052] Figure 4 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present invention. Detailed Implementation

[0053] The principles and features of the present invention are described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0054] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.

[0055] The solution provided in this invention can be applied to any application scenario requiring document-based question-and-answer functionality for issues in the railway field. The solution provided in this invention can be executed by any electronic device, such as a user's terminal device, including at least one of the following: smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, smart TV, or smart in-vehicle device.

[0056] This invention provides a possible implementation, such as... Figure 1 The diagram shows a flowchart of a document question-answering method. This method can be executed by any electronic device, such as a terminal device, or jointly by a terminal device and a server. For ease of description, the method provided in this embodiment will be described below using a terminal device as the execution subject as an example. Figure 1The flowchart shown indicates that the method may include the following steps:

[0057] S10: Obtain sample documents in different document formats in the railway field, and align each sample document based on a preset segment length to obtain multiple sample segments;

[0058] S20, semantic relevance is labeled for any two sample segments to obtain the first labeled data, and the document retrieval model is fine-tuned based on the first labeled data to obtain the target document retrieval model;

[0059] S30, based on multiple sample fragments and the semantic relevance between each sample fragment, determine the second labeled data, and fine-tune the semantic ranking model according to the second labeled data to obtain the target semantic ranking model. The second labeled data includes multiple questions and at least one sample fragment corresponding to each question.

[0060] S40 generates question-answer pairs based on multiple questions and at least one sample fragment corresponding to each question, using a basic large model.

[0061] S50, based on question-answer pairs and third-annotated data, constitutes fourth-annotated data. The basic large model is then fine-tuned based on the fourth-annotated data to obtain a generative large model. The third-annotated data includes instructions, questions, and answers generated for the railway sector.

[0062] S60 processes the problem to be processed based on the target document retrieval model, the target semantic ranking model, and the generative large model to obtain the target answer corresponding to the problem to be processed.

[0063] The method of this invention effectively improves the accuracy and usability of document question-answering systems by fine-tuning the model to address the unique characteristics of document organization and semantic expression in the railway sector. Specifically, by acquiring and aligning sample documents from the railway sector, high-quality sample fragments are generated. These fragments are then fine-tuned using semantic relevance annotation to improve the document retrieval and semantic ranking models, enabling the models to better understand and process railway-related document content. Simultaneously, the basic large-scale model is fine-tuned based on the generated question-answer pairs and annotated data, further optimizing the performance of the generative large-scale model. Finally, by combining the target document retrieval model, the target semantic ranking model, and the generative large-scale model to process the problem, more accurate and comprehensive target answers can be generated. This method effectively solves problems such as poor answer relevance, incomplete answers, and answer illusions encountered by traditional question-answering systems in railway-related documents, thereby reducing user usage and verification costs, improving user efficiency, reducing the cost of troubleshooting and locating actual railway problems, and better meeting the railway industry's needs for document question-answering systems.

[0064] The present invention will be further described below with reference to the following specific embodiments. In these embodiments, the present invention is combined with... Figure 2 This embodiment introduces a document question-and-answer method, which may include the following steps:

[0065] S10: Obtain sample documents in different document formats in the railway field, and align each sample document based on a preset segment length to obtain multiple sample segments;

[0066] The different document formats include doc, docx, pdf, html, and txt formats. The pdf format can include text-based pdfs and scanned pdfs.

[0067] Sample documents in different formats can specifically include railway industry documents such as technical specifications, emergency plans, and equipment manuals. For documents of different formats, corresponding parsers can be used to parse them, resulting in sample documents of different formats. Specifically, for .doc format, parsing tools can be used to convert it to .docx format; for .docx format, open-source parsers can be used for recognition; for text-based .pdf, layout and content recognition APIs can be used; for scanned .pdf, OCR models and text recognition software can be used; for .html and .txt formats, text content can be extracted and .html tags removed.

[0068] The preset segment length can be the mode of the average length, that is, the value of the average length of the segment that occurs most often, and can be preset.

[0069] The above-mentioned method of aligning each sample document based on a preset fragment length to obtain multiple sample fragments is as follows:

[0070] S101, For each document in all sample documents that is longer than the preset segment length, segment it to obtain multiple first segments. The length of each first segment is equal to the preset segment length. Here, a document that is longer than the preset segment length refers to a document whose text length is longer than the preset segment length. For a document, due to the uncertainty of the text length, not every segment after segmentation has a length equal to the preset segment length. For example, the length of the last segment after the document is segmented may be shorter than the preset segment length. For this segment, the second half of the content of the previous segment can be merged with the segment based on the preset segment length so that the length of the merged segment is equal to the preset segment length.

[0071] S102, for each document in all sample documents that is shorter than the preset segment length, the document length is extended by using contextual paragraph linking relationships and increasing the repetition coefficient before and after to obtain multiple second segments; wherein, for each second segment, the length of the second segment is not less than the preset segment length.

[0072] S103, based on multiple first segments and multiple second segments, multiple sample segments are obtained, that is, multiple sample segments include multiple first segments and multiple second segments.

[0073] By employing the above processing methods, we can ensure that the overall document parsing segment lengths are relatively consistent, reduce interference, improve semantic richness and semantic expression dimension differentiation, and enhance the overall retrieval and recall effect of the railway field document retrieval module.

[0074] Alternatively, one implementation of S102 above is as follows:

[0075] S1021, For each document in all sample documents that is shorter than the preset segment length, the document is divided into multiple third segments based on the preset segment length, with overlapping content between every two third segments;

[0076] The overlapping content helps the model better understand the context of each segment, especially when dealing with long documents, where contextual information is crucial for semantic understanding. The overlap between every two third segments serves the following purposes:

[0077] Enhanced contextual coherence: By overlapping, each segment can contain some of the content of the previous segment, thus maintaining contextual coherence.

[0078] Avoid information loss: When splitting a document, some key information may happen to be located on the boundary of the fragment. Overlapping can ensure that this information is not missed.

[0079] Improve model performance: When dealing with overlapping segments, the model is able to better understand the relationships between segments, thereby improving the overall semantic understanding ability.

[0080] S1022, For each third segment, construct the context chain of the third segment based on the third segment and its adjacent segments;

[0081] The context chain refers to the contextual relationship between the third segment and its adjacent segments.

[0082] S1023, for each third segment, based on the context chain of the third segment, the third segment is merged with the adjacent segments of the third segment to obtain the second segment, the length of the second segment is not less than the preset segment length.

[0083] Specifically, a context chain is constructed for each third segment to ensure that each segment contains context information before and after it. The specific method of S1023 is as follows:

[0084] Forward Linking: For each third segment, merge the third segment with the last few sentences (or a paragraph) of the previous third segment to obtain the second segment.

[0085] Backward Linking: For each third segment, the third segment is combined with the first few sentences (or a paragraph) of the next third segment to obtain the second segment.

[0086] S20, semantic relevance is labeled for any two sample segments to obtain the first labeled data, and the document retrieval model is fine-tuned based on the first labeled data to obtain the target document retrieval model;

[0087] The target document retrieval model is used to calculate the semantic relevance between document fragments.

[0088] One implementation of S20 above is as follows:

[0089] S201, convert each sample fragment into a vector of a set dimension (e.g., 512 dimensions); specifically, each sample fragment can be converted into a vector of a set dimension through a vectorization model.

[0090] S202, determine the semantic relevance between any two vectors;

[0091] Specifically, the semantic relevance between multiple vectors can be expressed through spatial representation (e.g., cosine). Vectors with high relevance have higher cosine values, with a maximum value of 1, while vectors with low relevance have lower cosine values, with a minimum value of 0. For relatively fixed-length paragraphs (sample fragments) parsed by the railway document parsing module, the semantic relevance between each sample fragment is labeled and distinguished by 0 / 1. By adjusting the 0-1 scoring results between paragraphs, the semantic relevance is represented, thereby affecting the retrieval model performance.

[0092] S203, based on the semantic relevance between various vectors, performs semantic relevance labeling on any two sample segments to obtain the first labeled data.

[0093] Among them, semantic relevance labeling for any two sample segments can be generated by manual annotation or large language model based on the semantic relevance between various vectors.

[0094] As an example, when two sample segments (e.g., the n1th segment of document A and the m2th segment of document B) have similar semantic expressions, they are labeled as (txt_n1, txt_m2, 1), indicating that the semantic expressions of these two segments are related. When two sample segments have almost no semantic correlation (e.g., the n1th segment of document A and the m2th segment of document B), they are not similar in semantic expression, and are labeled as (txt_n1, txt_m2, 0), indicating that the semantic expressions of these two sample segments are not related. Since both txt_n and txt_m are documents in the railway industry, the model fine-tuned after this stage of labeling incorporates railway knowledge representation, and its relevance results on railway professional documents are better than the publicly available model (trained using a general public dataset).

[0095] Based on the first set of labeled data, the document retrieval model was fine-tuned. During the fine-tuning process, the ratio of the training set, validation set, and test set in the sample dataset was 7:1:2.

[0096] S30, based on multiple sample fragments and the semantic relevance between each sample fragment, determine the second labeled data, and fine-tune the semantic ranking model according to the second labeled data to obtain the target semantic ranking model. The second labeled data includes multiple questions and at least one sample fragment corresponding to each question.

[0097] In this model, each question can correspond to at least one sample fragment, and each question can be generated based on various sample fragments. The target semantic ranking model is used to rank the document fragments and determine which fragments are more helpful in answering the questions.

[0098] Optionally, the determination of the second labeled data based on multiple sample segments and the semantic relevance between them may include:

[0099] S301 generates multiple questions based on multiple sample fragments and the semantic relevance between them;

[0100] Among these methods, a basic large model can be used to automatically generate questions based on the content of each sample segment, and then the questions can be manually reviewed and corrected to ensure their quality and relevance.

[0101] S302, For each question, determine the question relevance between the question and each sample segment;

[0102] Among them, question relevance refers to the textual relevance between the question and the sample fragment. The higher the relevance, the greater the likelihood that the answer to the question comes from the corresponding sample fragment.

[0103] S303, For each problem, based on the first target segment and the second target segment corresponding to the problem, generate labeled data corresponding to the problem. The first target segment is any sample segment whose problem relevance is greater than a first set value, and the second target segment is any sample segment whose problem relevance is less than a second set value.

[0104] S304. Based on the labeled data corresponding to all questions, determine the second labeled data.

[0105] Semantic ranking models typically use real-time document scoring models (in contrast, document retrieval models have a fixed vector representation for the same text segment, while semantic ranking models do not). The input is usually (q, txt1, txt2), and the score is (score1, score2), also known as question relevance, expressing the probability that two documents txt can answer a question. Therefore, for the same txt1, different questions q will lead to different (q, txt1) scores. For a relatively fixed-length sample segment txt_1 parsed by the railway document parsing module, by labeling the possible questions in the sample segment, question pairs (q1 q2…qn) are obtained. Simultaneously, a segment txt_2 is randomly selected from other irrelevant sample segments as a negative sample. The final labeling format is (q1, txt1, txt2), where the first position is the question, the second position is the highly relevant sample segment (first target segment), and the third position is the low-relevance sample segment (second target segment).

[0106] In this application, multiple questions are generated based on multiple sample fragments and the semantic relevance between each sample fragment. These questions can be constructed using either manual annotation or a generative large model.

[0107] In this application's scheme, since q1 is generated from txt1, txt1 can answer question q1 relatively well. However, txt2 is obtained by randomly selecting paragraphs from other irrelevant sample fragments, so txt2 usually cannot answer question q1. The final annotation is in the format (q1, txt1, txt2) as the second annotation data. The generative model understands that txt1 and txt2 contribute differently to answering question q1, thus enabling better document ranking in a practical question-answering system. Since txt_n and txt_m are both railway industry documents, the ranking model fine-tuned after this annotation stage incorporates railway knowledge representation, resulting in better relevance performance on railway professional documents compared to the publicly available model (trained using a general public dataset).

[0108] Optionally, the document retrieval model is fine-tuned based on the second labeled data. During the fine-tuning process, the ratio of the training set, validation set, and test set of the sample dataset is 7:1:2.

[0109] Optionally, the question types of the above multiple questions include parameter extraction type, clause listing type, numerical calculation type, and artificial scenario type. The second labeled data includes multiple questions corresponding to each question type and at least one sample fragment corresponding to each question.

[0110] Based on the characteristics of actual documents in the railway field and the common questioning methods of questioners, the question types are divided into parameter extraction type, clause listing type, numerical calculation type, and artificial scenario type. Second labeled data are constructed for each question type and reviewed and verified to improve the diversity of labeled data and enhance the model's ability in different scenarios.

[0111] Specifically, the second annotation data corresponding to various question types can be:

[0112] (1) The characteristic of parameter extraction type questions is that the answer to the question usually exists directly in the original text. For example, for the original text [A West Station is located in a certain district. The center mileage of this station is 24km 700m on the up line of a certain line and 22km 850m on the down line of a certain line. It belongs to the jurisdiction of A West Station of Company X], we can construct the question: Which station does A West Station belong to? The corresponding answer is a text fragment.

[0113] (2) The clause listing type of question is a comprehensive summary of multiple clauses. For example, for the original text [2.2.4 Main responsibilities of each member department: (1) Security Department: coordinate each workshop (station) to do a good job in fire protection knowledge and fire accident emergency response training; coordinate each department to do a good job in emergency response work, assist fire accident investigation and security work; assist relevant departments to do a good job in relevant matters after the fire accident emergency response is completed; cooperate in formulating the station fire accident emergency plan. (2) Office: responsible for communicating and contacting the group company and local government in accordance with the requirements of the fire accident emergency leadership group, assisting the station fire emergency leadership group to do a good job in emergency response work; coordinate the building construction department and power supply department to ensure production, living water supply and power supply during the emergency rescue process, and do a good job in heating and ventilation work; according to the rescue needs, provide food, drinking water, accommodation and other logistical support for the on-site working group and rescue personnel; cooperate with the labor and personnel department to coordinate medical rescue work; cooperate in formulating the station fire accident emergency plan. The legal counsel is responsible for the guidance, coordination and legal consultation of the handling of legal disputes caused by fire accidents. (3) Safety Department: In accordance with relevant regulations and based on the fire accident investigation report issued by the fire accident investigation team, it is responsible for organizing and conducting internal investigations and analyses of fire accidents and proposing qualitative handling opinions. It supervises, inspects, and guides the emergency response work for fire accidents in each workshop (station) and supervises the implementation process of the emergency plan. (4) Information Technology Department: It is responsible for formulating information system technical support plans to ensure smooth information system operation when the emergency plan is activated; and assists in the formulation of the station fire accident emergency plan. 】Combined with the context, a question can be constructed: In the fire emergency plan of West Station in Area A, what are the main responsibilities of each department? The corresponding answer is to summarize multiple points. This type of question tests the model's ability to provide comprehensive answers from multiple documents and to list multiple clauses without losing anything.

[0114] (3) Numerical calculation type questions require answers derived from mathematical calculations using data from the original document; the answers are not directly presented in the original text. For example, for the original text "[The combination between DS6-11 computer interlocking and outdoor signaling equipment currently uses relay circuits. These mainly include signal lighting circuits, turnout control circuits, track circuits, 64D semi-automatic block signaling, automatic block signaling circuits, and other combined circuits. The retained relays and signal circuits include LXJ, DXJ, YXJ, TXJ, ZXJ, FXJ, DJ, and 2DJ;]", a question can be constructed: How many types of relays are retained in the DS6-11 signaling circuit? The model should automatically summarize the 8 types of relays based on the data.

[0115] (4) Artificial scenario-based questions are based on the original document description. Inputting a question requires analyzing the original text, thinking, judging the corresponding requirements, and then obtaining the answer. For example, for the original text "[The distance between the edge of the high platform where passenger trains stop and the center line of the track is 1,750 mm, and the safety marking is 1,000 mm from the edge of the platform. The distance between the safety marking and the edge of the platform for non-high platforms is: 1,000 mm when the train speed is no more than 120 km / h; 1,500 mm when the train speed is between 120 km / h and 160 km / h; 2,000 mm when the train speed is between 160 km / h and 200 km / h]", a question can be constructed: A station is a non-high platform where trains with a speed of 180 km / h will pass. What are the requirements for the distance between the platform safety marking and the edge of the platform? The model should then determine the range of the original text [2,000 mm when the train speed is between 160 km / h and 200 km / h] based on the given scenario (a train will pass by at 180 km / h), and then provide a further answer.

[0116] S40 generates question-answer pairs based on multiple questions and at least one sample fragment corresponding to each question, using a basic large model.

[0117] S50, based on question-answer pairs and third-annotated data, constitutes fourth-annotated data. The basic large model is then fine-tuned based on the fourth-annotated data to obtain a generative large model. The third-annotated data includes instructions, questions, and answers generated for the railway sector.

[0118] Instructions are commands or hints that guide the model on how to process questions and generate answers. They are typically brief descriptions telling the model the type of task to be performed. For example:

[0119] "Extract key parameters based on the document content." (Parameter extraction type)

[0120] "Based on the document content, list all relevant clauses." (Clause List Type)

[0121] "Calculate the results based on the data in the document." (Numerical calculation type)

[0122] "Analyze and answer the questions based on the document description." (Artificial Context Type)

[0123] A question is a specific question that the user needs to answer from a document. The question type can be parameter extraction, clause listing, numerical calculation, or artificial context.

[0124] The answer is a response generated by the model based on the question and document content. The quality of the answer depends on each model's understanding of the question and its ability to parse the document content.

[0125] The generative large model, as the core model in the question-answering system, needs to be fine-tuned according to the document expression style, organizational form, and core terminology of the railway industry to better adapt to the document context and content characteristics of the railway field. Supervised fine-tuning is typically used for this large model. Each set of data containing an instruction-question-answer pair constitutes a fine-tuning data point. Multiple fine-tuning data points are generated to perform full or partial parameter fine-tuning of the large model. The fourth set of labeled data can contain multiple sets of instruction-question-answer pairs. Question-answer pairs consist of multiple pairs of questions and answers. Supplementing the question-answer pairs with the third set of labeled data yields multiple sets of data containing instruction-question-answer pairs, which serve as the fourth set of labeled data.

[0126] By using a basic model and combining it with railway sector instructions and roles, the basic model generates relatively high-quality question-and-answer pairs. These pairs are then manually reviewed, corrected, and filtered to select a set of high-quality question-and-answer pairs. These sets serve as the fourth set of labeled data for fine-tuning the basic model. The format of the fourth set of labeled data is typically (instruction, question, answer, and context paragraph).

[0127] The context paragraphs refer to the relevant paragraphs in the document corresponding to the answer.

[0128] The construction of the fourth labeled data should ensure that the number of questions in each category is approximately equal, avoiding excessive concentration of data in any one category, and preventing model performance degradation or imbalance. Based on the fourth labeled data, the generative large model is fine-tuned. During the fine-tuning process, the ratio of the training set, validation set, and test set in the sample dataset is 7:1:2.

[0129] S60 processes the problem to be processed based on the target document retrieval model, the target semantic ranking model, and the generative large model to obtain the target answer corresponding to the problem to be processed.

[0130] Alternatively, one implementation of the above S60 is as follows:

[0131] S601, Obtain the problem to be processed, which is a problem related to the railway sector;

[0132] S602, identify multiple document fragments corresponding to the question to be processed, wherein the multiple document fragments refer to document fragments that may contain the target answer corresponding to the question to be processed;

[0133] S603, calculate the semantic relevance between document fragments through the target document retrieval model;

[0134] S604, based on the semantic relevance between document fragments, sorts all document fragments using the target semantic ranking model. Specifically, the sorting can be based on the degree of semantic relevance. The higher the semantic relevance, the greater the likelihood that the document fragment corresponds to the target answer.

[0135] S605, based on all sorted document fragments, uses a generative large model to determine the target answer to the question to be processed.

[0136] Optionally, in the solution of this application, in S1021 above, for each document in all sample documents that is shorter than the preset segment length, the document is divided into multiple third segments based on the preset segment length, and the overlapping content between each two third segments is achieved by increasing the repetition coefficient between the front and back segments.

[0137] As an example, the original text (assuming a maximum of 3 sentences per paragraph) is as follows:

[0138] Training deep learning models requires a large amount of data. Data quality directly affects model performance. Therefore, data cleaning is a crucial step. Cleaned data then needs to be standardized. Standardization methods include Z-score and Min-Max.

[0139] If you are performing segmentation without overlap, that is, segmenting using methods that do not involve overlapping content:

[0140] Segment 1: "Training deep learning models requires a large amount of data. Data quality directly affects model performance. Therefore, data cleaning is a crucial step." (Segmentation point: end of the third sentence)

[0141] Segment 2: "The cleaned data still needs to undergo standardization. Standardization methods include Z-score and Min-Max."

[0142] This segmentation has the following problems: the second paragraph suddenly mentions "cleaned data" at the beginning, which lacks a connection with the preceding text.

[0143] If you are performing paragraphing with overlap (one overlapping sentence), that is, using paragraphing methods that involve overlapping content:

[0144] Segment 1: "Training deep learning models requires a large amount of data. Data quality directly affects model performance. Therefore, data cleaning is a crucial step."

[0145] Segment 2: "Therefore, data cleaning is a crucial step. The cleaned data then needs to undergo standardization. Standardization methods include Z-score and Min-Max."

[0146] When there is no overlap, the phrase "cleaned data" in the second paragraph seems abrupt.

[0147] When there is overlap, the model can clearly know that "standardization" is applied to the results after "data cleaning".

[0148] After obtaining the third segment by increasing the repetition coefficient, for each third segment, a context chain of the third segment is constructed based on the third segment and its adjacent segments; then, based on the context chain of the third segment, the third segment is merged with its adjacent segments to obtain the second segment, the length of which is not less than the preset segment length.

[0149] As an example, suppose the original document is segmented into segments abcd. Since a single segment b cannot express a complete semantic meaning, the content before and after the third segment b [abc] is used as the basic unit through document zipping. This is then merged to obtain the second segment, which serves as the input for subsequent strategy modules (such as the basic large model).

[0150] Paragraph zippers are a further optimization of Overlap segmentation. They generate more efficient text segments by merging, crossing, or deduplicating, reducing redundancy, enhancing coherence, and adapting to model input constraints.

[0151] Optionally, in the scheme of this application, in S101, each document in all sample documents that is longer than the preset segment length is segmented to obtain multiple first segments, and the length of each first segment is equal to the preset segment length, may further include taking one document in all sample documents that is longer than the preset segment length as the document to be processed and performing the following processing:

[0152] The first method involves segmenting the document into paragraphs. For any paragraph obtained after segmentation, the paragraph is further segmented into sentences based on periods or semicolons. The sentence length of the segmented sentences is determined in real time. Based on the sentence length and the preset segment length, each sentence is further segmented to obtain multiple first segments with a text length equal to the preset segment length.

[0153] The specific implementation process for segmenting each statement into multiple text segments based on its statement length and a preset segment length is as follows:

[0154] For each statement, if the statement length is greater than the preset segment length, the corresponding statement is divided into a first number of first text segments, each with a length equal to the preset segment length. The total length of all first text segments is less than the statement length. The text segments in the statement other than all first text segments are taken as second text segments. The statement length of the second text segments is usually not greater than the preset segment length. In this case, the second text segment can be directly taken as a new first text segment. Alternatively, the second text segment can be concatenated with the next adjacent statement of the statement to obtain a third text segment of the preset segment length. The third text segment is taken as a new first text segment. All first text segments (including the new first text segments) are taken as multiple first segments corresponding to the statement.

[0155] If the next adjacent statement of a statement has already been partially occupied by the third text segment of the statement, then the part of the adjacent statement other than that part can be treated as a new statement and segmented in the manner described above.

[0156] The second method involves segmenting the document into paragraphs. A fixed segmentation length is determined based on a preset segment length. Each paragraph is then segmented according to this fixed length to obtain at least two segments. It's important to note that during the segmentation process, if any resulting segments contain commas or pauses, the segmentation must be repeated. Specifically, segments must be made based on the comma or pause, and then re-segmented starting with the first text following the comma or pause, again according to the fixed segmentation length. Each segmented result is then sliced ​​using the slicing method described in the first method to obtain the first segment corresponding to each sentence.

[0157] The solution of this invention, through fine-tuning the model strategy based on railway industry documents and constructing a document-based question-answering method based on the fine-tuned model, can integrate railway documents, terminology, semantic expressions, and features into each core strategy model node in a deeper way (rather than the traditional RAG approach of simply attaching a database). This fundamentally solves the obvious quality and effectiveness problems that often occur in general large-scale model question answering on railway domain datasets, such as poor answer relevance, incomplete answers, answer illusions, incorrect answer results, and incorrect document citations. It improves answer accuracy and usability, reduces user usage and verification costs, and increases user work efficiency. The question-answering system provides professional, accurate, and comprehensive answers, better leveraging the advantages of the railway large-scale model in various scenarios such as management efficiency improvement, emergency management, manual reference, and vocational education.

[0158] Based on and Figure 1 Based on the same principle as the method shown, this embodiment of the invention also provides a document question-and-answer device 20, such as... Figure 3 As shown, the document question-answering device 20 may include an acquisition module 210, a document retrieval model fine-tuning module 220, a semantic ranking model fine-tuning module 230, a question-answer pair generation module 240, a basic large model fine-tuning module 250, and a processing module 260, wherein:

[0159] The acquisition module 210 is used to acquire sample documents in different document formats in the railway field, and to perform alignment processing on each sample document based on a preset segment length to obtain multiple sample segments.

[0160] The document retrieval model fine-tuning module 220 is used to perform semantic relevance annotation on any two sample segments to obtain the first annotation data, and to fine-tune the document retrieval model based on the first annotation data to obtain the target document retrieval model;

[0161] The semantic ranking model fine-tuning module 230 is used to determine the second labeled data based on multiple sample fragments and the semantic relevance between each sample fragment, and to fine-tune the semantic ranking model according to the second labeled data to obtain the target semantic ranking model. The second labeled data includes multiple questions and at least one sample fragment corresponding to each question.

[0162] The question-answer pair generation module 240 is used to generate question-answer pairs based on multiple questions and at least one sample fragment corresponding to each question, using a basic large model.

[0163] The basic large model fine-tuning module 250 is used to construct fourth labeled data based on question-answer pairs and third labeled data, and to fine-tune the basic large model based on the fourth labeled data to obtain a generative large model. The third labeled data includes instructions, questions and answers generated for the railway field.

[0164] The processing module 260 is used to process the problem to be processed based on the target document retrieval model, the target semantic ranking model and the generative big model to obtain the target answer corresponding to the problem to be processed.

[0165] Optionally, when the processing module 260 processes the problem to be processed based on the target document retrieval model, the target semantic ranking model, and the generative large model to obtain the target answer corresponding to the problem to be processed, it is specifically used for:

[0166] Retrieve pending issues, which are specific to the railway sector;

[0167] Identify the multiple document fragments corresponding to the problem to be addressed;

[0168] The semantic relevance between document fragments is calculated using the target document retrieval model.

[0169] Based on the semantic relevance between document fragments, all document fragments are ranked using a target semantic ranking model;

[0170] Based on all sorted document fragments, a generative large model is used to determine the target answer to the question to be processed.

[0171] Optionally, when the document retrieval model fine-tuning module 220 performs semantic relevance annotation on any two sample segments to obtain the first annotation data, it is specifically used for:

[0172] Convert each sample fragment into a vector of a specified dimension;

[0173] Determine the semantic relevance between any two vectors;

[0174] Based on the semantic relevance between various vectors, semantic relevance is labeled for any two sample segments to obtain the first labeled data.

[0175] Optionally, when the acquisition module 210 performs alignment processing on each sample document based on a preset fragment length to obtain multiple sample fragments, it is specifically used for:

[0176] Each document in all sample documents that is longer than the preset segment length is segmented to obtain multiple first segments;

[0177] For each document in all sample documents that is shorter than the preset segment length, the document length is extended by using contextual paragraph linking and increasing the repetition coefficient before and after the document to obtain multiple second segments;

[0178] Multiple sample segments are obtained based on multiple first segments and multiple second segments.

[0179] Optionally, when the acquisition module 210 expands the document length of each document shorter than the preset segment length in all sample documents to obtain multiple second segments by using contextual paragraph linking relationships and increasing the repetition coefficient, it is specifically used for:

[0180] For each document in all sample documents that is shorter than the preset segment length, the document is divided into multiple third segments based on the preset segment length, with overlapping content between every two third segments;

[0181] For each third segment, construct the context chain of the third segment based on the third segment and its adjacent segments;

[0182] For each third segment, based on the context chain of the third segment, the third segment is merged with the adjacent segments of the third segment to obtain the second segment. The length of the second segment is not less than the preset segment length.

[0183] Optionally, when determining the second labeled data based on multiple sample fragments and the semantic relevance between each sample fragment, the aforementioned semantic ranking model fine-tuning module 230 is specifically used for:

[0184] Multiple questions are generated based on multiple sample fragments and the semantic relevance between them.

[0185] For each question, determine the question relevance between the question and each sample segment;

[0186] For each problem, based on the first target segment and the second target segment corresponding to the problem, labeled data corresponding to the problem is generated. The first target segment is any sample segment whose problem relevance is greater than a first set value, and the second target segment is any sample segment whose problem relevance is less than a second set value.

[0187] Based on the labeled data corresponding to all questions, the second labeled data is determined.

[0188] Optionally, the question types of the multiple questions include parameter extraction type, clause listing type, numerical calculation type, and artificial scenario type. The second labeled data includes multiple questions corresponding to each question type and at least one sample fragment corresponding to each question.

[0189] The document question-and-answer device of this invention can execute the document question-and-answer method provided in this invention. The implementation principle is similar. The actions performed by each module and unit in the document question-and-answer device in each embodiment of this invention correspond to the steps in the document question-and-answer method in each embodiment of this invention. For detailed functional descriptions of each module of the document question-and-answer device, please refer to the descriptions in the corresponding document question-and-answer methods shown above. They will not be repeated here.

[0190] The aforementioned document question-and-answer device can be a computer program (including program code) running on a computer device, such as an application software; the device can be used to execute the corresponding steps in the method provided in the embodiments of the present invention.

[0191] In some embodiments, the document question-answering device provided in this invention can be implemented using a combination of hardware and software. As an example, the document question-answering device provided in this invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the document question-answering method provided in this invention. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0192] In other embodiments, the document question-answering device provided in this invention can be implemented in software. Figure 3 A document question-answering device stored in a memory is shown. It can be software in the form of programs and plug-ins, and includes a series of modules, including an acquisition module 210, a document retrieval model fine-tuning module 220, a semantic ranking model fine-tuning module 230, a question-answer pair generation module 240, a basic large model fine-tuning module 250, and a processing module 260, for implementing the document question-answering method provided in the embodiments of the present invention.

[0193] The modules described in the embodiments of the present invention can be implemented in software or hardware. The names of the modules are not, in some cases, limiting the scope of the module itself.

[0194] Based on the same principles as the methods shown in the embodiments of the present invention, the embodiments of the present invention also provide an electronic device, which may include, but is not limited to: a processor and a memory; the memory for storing computer programs; and the processor for executing the methods shown in any embodiment of the present invention by invoking the computer programs.

[0195] In one alternative embodiment, an electronic device is provided, such as Figure 4 As shown, Figure 4The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present invention.

[0196] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this invention. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0197] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0198] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.

[0199] The memory 4003 stores application code (computer program) for executing the present invention, and its execution is controlled by the processor 4001. The processor 4001 executes the application code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.

[0200] Among these, electronic devices can also be terminal devices. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0201] This invention provides a computer-readable storage medium storing a computer program that, when run on a computer, enables the computer to execute the corresponding content in the aforementioned method embodiments.

[0202] According to another aspect of the present invention, a computer program product or computer program is also provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various embodiments described above.

[0203] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0204] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0205] The computer-readable storage medium provided in this invention can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0206] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.

[0207] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.

Claims

1. A document question-and-answer method, characterized in that, include: Sample documents in different document formats in the railway field are obtained, and each sample document is aligned based on a preset segment length to obtain multiple sample segments. Semantic relevance is labeled for any two sample segments to obtain first labeled data, and the document retrieval model is fine-tuned based on the first labeled data to obtain the target document retrieval model; Based on multiple sample fragments and the semantic relevance between each sample fragment, second labeled data is determined, and the semantic ranking model is fine-tuned according to the second labeled data to obtain the target semantic ranking model. The second labeled data includes multiple questions and at least one sample fragment corresponding to each question. Based on multiple questions and at least one sample fragment corresponding to each question, question-answer pairs are generated through a basic large model. Based on the question-and-answer pairs and the third annotation data, a fourth annotation data is constructed. The basic large model is then fine-tuned based on the fourth annotation data to obtain a generative large model. The third annotation data includes instructions, questions, and answers generated for the railway sector. Based on the target document retrieval model, the target semantic ranking model, and the generative large model, the problem to be processed is obtained to obtain the target answer corresponding to the problem to be processed. The step of aligning each sample document based on a preset segment length to obtain multiple sample segments includes: Each document in all sample documents that is longer than the preset segment length is segmented to obtain multiple first segments; For each document in all sample documents that is shorter than the preset segment length, the document length is extended by using contextual paragraph linking and increasing the repetition coefficient before and after the document to obtain multiple second segments; Multiple sample segments are obtained based on multiple first segments and multiple second segments.

2. The method according to claim 1, characterized in that, The process of processing the problem to be processed based on the target document retrieval model, the target semantic ranking model, and the generative large model to obtain the target answer corresponding to the problem to be processed includes: Obtain the problems to be processed, which are problems related to the railway sector; Identify multiple document fragments corresponding to the problem to be processed; The semantic relevance between the document fragments is calculated using the target document retrieval model. Based on the semantic relevance between the document fragments, all document fragments are sorted using the target semantic ranking model; Based on all sorted document fragments, the target answer corresponding to the problem to be processed is determined through the generative large model.

3. The method according to claim 1, characterized in that, The step of semantically relevance labeling for any two sample segments to obtain first labeled data includes: Each of the sample fragments is converted into a vector of a specified dimension; Determine the semantic relevance between any two vectors; Based on the semantic relevance between the vectors, semantic relevance is labeled for any two sample segments to obtain the first labeled data.

4. The method according to claim 1, characterized in that, For each document in all sample documents that is shorter than the preset segment length, the document length is extended by using contextual paragraph linking relationships and increasing the repetition coefficient before and after the paragraph to obtain multiple second segments, including: For each document in all sample documents that is shorter than the preset segment length, the document is divided into multiple third segments based on the preset segment length, with overlapping content between every two third segments; For each of the third segments, a context chain for the third segment is constructed based on the third segment and its adjacent segments; For each of the third segments, based on the context chain of the third segment, the third segment is merged with the adjacent segments of the third segment to obtain a second segment, the length of the second segment being not less than the preset segment length.

5. The method according to any one of claims 1 to 3, characterized in that, The determination of the second labeled data based on multiple sample fragments and the semantic relevance between each sample fragment includes: Multiple questions are generated based on multiple sample fragments and the semantic relevance between the sample fragments. For each of the problems, determine the problem relevance between the problem and each of the sample segments; For each of the problems, labeled data corresponding to the problem is generated based on the first target segment and the second target segment corresponding to the problem. The first target segment is any sample segment whose problem relevance is greater than a first set value, and the second target segment is any sample segment whose problem relevance is less than a second set value. Based on the labeled data corresponding to all questions, the second labeled data is determined.

6. The method according to any one of claims 1 to 3, characterized in that, The problem types of the multiple problems include parameter extraction type, clause listing type, numerical calculation type, and artificial scenario type. The second labeled data includes multiple problems corresponding to each problem type and at least one sample fragment corresponding to each problem.

7. A document question-and-answer device, characterized in that, The document question-answering method according to claim 1, wherein the apparatus comprises: The acquisition module is used to acquire sample documents in different document formats in the railway field, and to align each sample document based on a preset segment length to obtain multiple sample segments. The document retrieval model fine-tuning module is used to perform semantic relevance labeling on any two sample segments to obtain first labeled data, and to fine-tune the document retrieval model based on the first labeled data to obtain the target document retrieval model; The semantic ranking model fine-tuning module is used to determine second labeled data based on multiple sample fragments and the semantic relevance between each sample fragment, and to fine-tune the semantic ranking model according to the second labeled data to obtain the target semantic ranking model. The second labeled data includes multiple questions and at least one sample fragment corresponding to each question. The question-answer pair generation module is used to generate question-answer pairs based on multiple questions and at least one sample fragment corresponding to each question, using a basic large model. The basic large model fine-tuning module is used to construct fourth annotation data based on the question-answer pair and the third annotation data, and to fine-tune the basic large model based on the fourth annotation data to obtain a generative large model. The third annotation data includes instructions, questions and answers generated for the railway field. The processing module is used to process the problem to be processed based on the target document retrieval model, the target semantic ranking model, and the generative big model to obtain the target answer corresponding to the problem to be processed.

8. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method of any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Domain-adaptive retrieval enhancement generation method and system

    CN119669400A