Railway whole-process consultation document digitization and retrieval method based on large model

By adopting a large-scale model-based digital method for railway whole-process consulting documents, the problem of low correlation between multi-source information was solved, the accurate extraction and integration of key information was achieved, the document processing efficiency and information correlation were improved, and high-quality project consulting and investment control were supported.

CN120821758AActive Publication Date: 2025-10-21CHINA RAILWAY ECONOMIC & PLANNING RES INST
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511332075.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-10-21
Estimated Expiration
2045-09-18

AI Technical Summary

Technical Problem

In railway engineering consulting, the low correlation between multi-source information, inaccurate extraction of core information, and semantic bias lead to low document processing efficiency and difficulty in achieving effective parsing and correlation of data across stages.

Method used

A large-scale model-based approach is adopted to classify and grade railway consultation documents throughout the entire process, construct a virtual environment, use a large language model for semantic segmentation and vector storage encoding, combine structured and vector databases for retrieval, determine semantic similarity through the large language model and perform text fusion, construct digital extraction standards, and achieve accurate extraction and fusion of key information.

Benefits of technology

It improves document processing efficiency, ensures the accuracy and comprehensiveness of key information, enables effective data correlation across stages, and supports high-quality project consulting and investment control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821758A_ABST
    Figure CN120821758A_ABST
Patent Text Reader

Abstract

The invention discloses a railway whole-process consultation document digitization and retrieval method based on a large model. The method comprises the steps that S1, railway whole-process consultation documents are classified and graded; s2, constructing a virtual environment required by large model deployment; s3, constructing a digital retrieval database; s4, constructing a digital extraction standard for the classified and graded railway whole-process consultation documents; s5, document retrieval is carried out; s5, determining which type of documents the reserved paragraphs belong to according to the retrieval result of S5, and extracting corresponding contents according to the digital extraction standard constructed in S4; and S7, fusing the content extracted in the step S6 with the document paragraphs retrieved in the step S5 so as to realize accurate answer to the retrieval requirements of the user. According to the method, original document content is extracted, key content is reserved, key information omission in the summarizing process of a large model and content splitting of a traditional segmentation mode are avoided, real-time analysis and overall control of the project progress are achieved, and the accuracy of whole-process consultation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of railway engineering management, and in particular to a large-scale model-based railway whole-process consulting document digitization and retrieval method. Background Art

[0002] Railway full-process engineering consulting runs through all stages of railway construction project decision-making, survey and design, engineering implementation, completion acceptance and operation and maintenance. It can solve the fragmentation problem of engineering consulting service methods through integrated innovation of consulting work models at each stage.

[0003] Railway full-process engineering consulting provides clients with comprehensive engineering consulting services, led by project consulting and management, with specialized consulting services as specific content. Project consulting and management, guided by the overall goals of full-process engineering consulting, focuses on schedule management, quality management, and investment management. It organically integrates specialized consulting services throughout the entire railway construction project, and is the core and key to realizing the value of full-process engineering consulting.

[0004] Railway full-process engineering consulting, as a comprehensive, integrated intellectual service, involves numerous document types and volumes. Data from each phase is diverse and subject to inconsistent standards. For example, structured data from the survey and design phase and unstructured reports from construction and maintenance are highly heterogeneous. Unstructured data, such as technical reports, meeting minutes, and drawing annotations, accounts for a high proportion and has low semantic relevance. In collaborative consulting across disciplines and phases, structured design drawings and unstructured text lack effective data parsing methods, making them challenging to access and connect.

[0005] Document digitization refers to the use of information technology to digitally process documents to improve the efficiency of document storage, retrieval, and use. This significantly increases the speed and accuracy of information retrieval, and through rights management, ensures document security. The core value of document digitization lies in its ability to transform fragmented and inefficient traditional document processing into a new model driven by knowledge interconnection and intelligence.

[0006] In view of the problems faced by railway full-process consulting, such as the large number of document types and the large number of documents, a document digitization method is urgently needed to achieve unified management of multi-source heterogeneous data in the survey and design and construction / operation and maintenance stages, eliminate the barriers of data between stages, realize the association of information and knowledge across stages, and improve the processing efficiency of railway full-process engineering consulting documents. Summary of the Invention

[0007] The purpose of the present invention is to provide a large-scale model-based method for digitizing railway full-process consulting documents to solve the problems of low correlation between multi-source information, inaccurate extraction of core information, and deviation in semantic understanding in the existing technology.

[0008] To this end, the present invention adopts the following technical solutions:

[0009] A method for digitizing and retrieving railway whole-process consulting documents based on a large model includes the following steps:

[0010] S1, Classification and grading of railway whole process consulting documents:

[0011] The railway full-process consulting documents are classified into folders according to project and object types. The folder directories are divided into two levels: project level and classification level. The classification level folders are classified according to the different stages of the railway full-process consulting project;

[0012] S2, builds the virtual environment required for large model deployment;

[0013] S3, building a digital retrieval database: sequentially read the contents of all railway full-process consulting documents in each project-level folder, and perform semantic segmentation on the read contents; perform text structured storage encoding on the segmented text contents and store them in a structured database; perform vector storage encoding on the segmented text contents and store them in a vector database;

[0014] S4: Establish digital extraction standards for railway whole-process consulting documents classified and graded in S1;

[0015] S5, document retrieval: extract keywords from the search content input by the user, obtain the top 1 to 5 vectors with similarity to the search content from the vector database through the search enhancement generation method, and match the vectors to the corresponding paragraphs in the vector database; perform a noun search on the structured database, compare the search results of the structured database with the matching results of the vector database, and retain consistent results;

[0016] S6, based on the search results retained in S5, determines which category of documents the retained paragraph belongs to, and then extracts the corresponding content based on the digital extraction criteria constructed in S4;

[0017] S7, merges the content extracted by S6 with the document paragraphs retained by S5 to achieve accurate answers to user search needs.

[0018] In the above step S2, the langchain and transformer framework modules required for the deployment of the large language model are configured in the virtual environment, and the Baichuan large language model is deployed locally.

[0019] In the above step S1, the formats of the railway full-process consultation documents include Word, PDF and Excel formats, and the document types include engineering reporting documents, standard specifications, document issuance documents, and full-process consultation weekly and monthly reports.

[0020] Preferably, in the above step S3, the content of the railway full-process consulting document is read by Python programming.

[0021] The segmentation method in step S3 is as follows: using a large language model, a prompt method is used to determine the semantic similarity between two adjacent natural segments. If they are similar, a number 1 is output; if they are not similar, a number 0 is output. Based on the output number, the semantically similar paragraphs are merged into one segment.

[0022] In the above step S3, the method for text structured storage encoding is: the text content after segmentation remains unchanged, and the segmented text is labeled by constructing the label field "content"; the method for vector storage encoding is: the segmented text content is encoded through the natural language processing model.

[0023] In step S4, for engineering report documents, the digital extraction standards are: nouns, key parameters, and corresponding data; for standard specification documents, the digital extraction standards are: proper nouns, corresponding expressions, and corresponding specifications; for issued documents, the digital extraction standards are: keywords, content summaries; for full-process consulting weekly and monthly reports, the digital extraction standards are: weekly / monthly report issue number, context with the words "key points", "important", and "difficult and difficult points", and numbers other than dates.

[0024] Preferably, in step S5, a noun search is performed on the structured database through an Elasticsearch module.

[0025] The present invention establishes a set of digital storage and query methods for railway construction project documents. It adopts a natural segment adaptive segmentation method based on a large language model to construct a parallel dual storage mode of text storage and vector storage, defines the digital extraction standard for the construction of classified document materials, and realizes key information extraction and information fusion based on the large language model.

[0026] Compared with the prior art, the present invention has the following beneficial effects:

[0027] 1. This invention implements database queries in parallel through character matching and semantic query, and defines key content within a paragraph based on digital extraction criteria. This not only refines the original document content but also retains the key content of the original document, preventing large models from missing key information in the original document during the summarization process.

[0028] 2. This invention uses a semantic segmentation approach, using a large language model to determine the semantic similarity between segments, grouping semantically consistent content into the same small segment, ultimately achieving semantic segmentation. The large language model is also used to determine the semantic similarity between adjacent segments, grouping semantically similar paragraphs into the same segment, thus avoiding the fragmentation of content often associated with traditional segmentation approaches.

[0029] 3. This invention uses a large language model to digitally search and extract documents from the entire railway construction project process, resolving issues such as inaccurate key information and incomplete summarization caused by traditional search and extraction. It effectively improves the processing efficiency of all consulting documents and enhances the accuracy and comprehensiveness of key information acquisition.

[0030] 4. This invention establishes digital extraction standards for classified railway full-process consulting documents. Documents of different formats are categorized according to various stages of consulting reports, including engineering submissions, standards and specifications, issued documents, weekly and monthly full-process consulting reports, and other categories. Digital extraction standards for railway full-process consulting documents are proposed based on the core content of each document category, specifying the key content of each document category.

[0031] 5. The method of the present invention can be used to accurately control the quality of the entire consultation process of railway construction projects. It can accurately extract key and difficult issues in engineering construction from design documents and provide accurate consultation on these key and difficult issues, which helps to achieve high-quality construction of the project. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 Flowchart of a method for digitizing and retrieving railway whole-process consulting documents based on a large model in an embodiment of the present invention;

[0033] Figures 2 to 4 This is the result obtained when using the deepseek official website large model to summarize important work in the embodiment of the invention. DETAILED DESCRIPTION

[0034] The following is a clear and complete description of the technical solutions of the present invention with reference to the accompanying drawings and embodiments. It is apparent that the following embodiments are only partial embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by persons of ordinary skill in the art without inventive effort are within the scope of protection of the present invention.

[0035] Example

[0036] See also Figure 1 The present invention provides a large-scale model-based railway whole-process consulting document digitization and retrieval method, comprising the following steps:

[0037] S1, Classification and grading of railway whole process consulting documents:

[0038] Railway full-process consulting documents are available in formats such as Word, PDF, and Excel. Document types include engineering reporting documents, standards and specifications, document issuance documents, and weekly and monthly full-process consulting reports. The present invention categorizes railway full-process consulting documents by project and object type using folders. The folder directories are divided into two levels: project level and classification level. Project-level folders categorize different railway full-process consulting projects, while classification-level folders categorize them according to the different stages of each railway full-process consulting project. Files in each classification-level folder have different formats and types.

[0039] S2, builds the virtual environment required for large model deployment:

[0040] In this embodiment, conda is used to build the virtual environment. The construction method is: conda create -ndigitization python=3.10 (the Python version can be continuously upgraded according to the development of subsequent versions); activate the virtual environment, conda activate digitization.

[0041] In addition, the langchain and transformer framework modules required for large language model deployment are configured in the virtual environment, and the Baichuan (2-13B) large language model is deployed locally.

[0042] S3, building a digital retrieval database, includes the following steps:

[0043] S31, through Python programming, reads the contents of railway full-process consulting documents of all formats and types in each project-level folder in turn, segments the read text content, and splits multiple long documents into multiple paragraphs of text.

[0044] Existing segmentation methods segment content based on natural segments, while this invention uses semantic segmentation. Using a large language model to determine the semantic similarity between adjacent natural segments, content with consistent semantics is merged into one segment, ultimately achieving semantic segmentation.

[0045] When using a large language model to determine the semantic similarity between two adjacent paragraphs, a prompt is used. The prompt content is: "Please determine the semantic similarity between the two paragraphs. If similar, please output the number 1; if not similar, output the number 0." Based on the output number, paragraphs with similar semantics are merged into one paragraph.

[0046] S32, constructing a database for the segmented text content. The details are as follows:

[0047] Encode the text content after S31 segmentation. The encoding methods are divided into text structured storage (electsearch) encoding method and vector storage encoding method. The two encoding methods correspond to different retrieval methods. Among them:

[0048] The text structured storage encoding method is: the text content after segmentation remains unchanged, and the segmented text is labeled by constructing the tag field "content" and then stored in a structured database.

[0049] The vector storage encoding method involves encoding (embedding) the segmented text content using a natural language processing model. In this embodiment, based on the proposed document semantic retrieval requirements, a semantic encoding model is used to encode the segmented results, obtaining fixed-length (e.g., 1024) text vectors, which are automatically numbered and stored in a vector database.

[0050] S4: Construct digital extraction standards for documents classified and graded in S1.

[0051] For engineering reporting documents, the digital extraction standards are: noun, key parameters, and corresponding data;

[0052] For the standard specification category, the digital extraction criteria are: proper noun, corresponding expression, and corresponding specification;

[0053] For published documents, the digital extraction criteria are: keywords, content summary;

[0054] For the weekly and monthly reports of the whole process consultation, the digital extraction standards are: the issue number of the weekly / monthly report, the context with the words "key points", "important", "difficult points", and numbers other than the date.

[0055] S5, document retrieval. Details are as follows:

[0056] The user describes the content to be retrieved in natural language. The present invention has no restrictions on the expression mode of the user input content. In this embodiment, the user input is "What are the important tasks of XX railway in X month?"

[0057] Keywords are extracted from the content input by the user, and three keywords are obtained: XX railway, X month, and important work.

[0058] Through the retrieval enhancement generation method, the top three vectors with similarity to the search content are obtained, and the vectors are matched to the corresponding paragraphs in the vector database.

[0059] In order to ensure the reliability of the search results, the structured database is searched for terms using the Elasticsearch (character query algorithm) module, and the work documents of XX Railway in X month are located using the terms "XX Railway" and "X Month" among the three keywords.

[0060] Then, the search results of the structured database are compared with the search results of the vector database matched by the search enhancement generation method, and the consistent results are retained.

[0061] S6, perform standardized extraction on the results retained by S5: determine which type of document the retained paragraph belongs to based on the search results retained by S5 (engineering reporting documents, standards and specifications, issued documents, or full-process consultation weekly and monthly reports), and then extract the corresponding content based on the digital extraction standards constructed in S4.

[0062] In this embodiment, the results obtained in S5 determine that the retrieved document paragraphs belong to the category of full-process consulting weekly and monthly reports. The extraction method uses a large language model. Since the extraction criteria for weekly and monthly report documents are: weekly / monthly report issue number, context with the words "key points", "important", and "key and difficult points", and numbers other than dates, the content obtained is the following three paragraphs:

[0063] "Key tasks are progressing. Railway X Institute has completed the preparation of documents for the design changes for XXX Station (changes to the station front and XXXX due to adjustments to the overall layout) on X / X / X. Design and cost consulting opinions have been sought during the process, and the project is now ready for review."

[0064] "Controlling and key projects. XXX Tunnel: 195 meters of excavation were completed at the 1# inclined shaft this month, with a cumulative total of 722 meters (53%); 189.5 meters of excavation were completed at the 2# inclined shaft this month, with a cumulative total of 919.1 meters (69%); 26.4 meters of excavation were completed at the entrance this month, with a cumulative total of 70.2 meters (4%); this is about two months behind the guidance schedule, mainly due to construction organization and excavation progress indicators not meeting expectations. XX Tunnel: 86 meters of excavation were completed at the entrance this month, with a cumulative total of 529.5 meters (25%); 71 meters of excavation were completed at the exit section this month, with a cumulative total of 172 meters (5%); 4.8 meters of excavation were completed at the main section this month, with a cumulative total of 130.4 meters (58%); - 2 - The overall construction group is about 1.5 months behind the guidance schedule, mainly due to the restrictions on nighttime construction and the late arrival of mechanized supporting equipment. XXXX Bridge: The pile foundation construction of the main piers of the 219#~220# rotating bridge has been completed, and the turntable under the 220# pier has been completed; the 3.5~4.5m foundation of the pier cap under the 220# pier has been completed; Concrete pouring is complete; progress is not delayed. The overall project progress is shown in the table below. "II. Key coordination and focus items for this month (I) Important meetings. On XX / XX / XX, XXXXX's Deputy General Manager visited the XX Railway for research and guidance. He visited the XXX beam fabrication yard, the XX tunnel entrance, and other construction sites. After listening to the reports, he affirmed the project's interim achievements and put forward four requirements: First, fully coordinate the strengths of all parties involved in the project, focus on key and difficult control projects, and work together to resolve difficulties and bottlenecks to create conditions for construction; second, stick to the goals and fully promote the progress of Phases 1 and 2 of the XX Railway as planned, and deliver on the construction schedule; third, make overall plans and organize operational intervention and external environment remediation in advance; fourth, strictly adhere to the bottom line and create a safe, century-long, high-quality project."

[0065] "(III) Key and Difficult Points XXXX. 1. XX County XXXXX Mine. On X month XX day, the XX County Railway Department held a meeting on the disposal of the XXXXX Mine. XX City XX Office, XXXX Company, XXXX Command, and the XX Project Department participated. After discussion among all parties, the meeting reached the following consensus: First, risks still exist in the slope of the mine and hidden dangers should be eliminated promptly; second, XX County XX should identify the mining rights of the mine and clarify the ownership of the mine in writing to provide a basis for subsequent XXXXXX; third, after the safety hazards of the mine are eliminated and the ownership is confirmed, the "case-by-case discussion" coordination and disposal will be adopted in accordance with the requirements of Document No. XXXX [XXXX] XXX. 2. XX County XX Water Plant XX. On X month XX day, XX On the day, XXXX Company organized a meeting to coordinate and promote the reservoir of XX Company in XX County. The meeting agreed to the overall XX and XX plans for the reservoir of XX Company on the south side, and agreed that the XXXXXX Department would be responsible for the design, implementation, and policy handling of the XX plan. It is planned to obtain XX approval in X month and complete the reservoir XX in XX month. The XXX Bureau should coordinate the railway construction schedule; it should be involved in the entire consultation, design, and price review process, review and control the plans and costs, and control investment. 3. Road XX in XX Village, XX County. On XX day of X month, XXXX Company organized an on-site survey to study the XX plan for the road in XX Village, XX County. Regarding the problems of one XXXX road that cannot be connected smoothly and two XX walls located within the XX land boundary, the safety distance between the parallel road and railway, the setting of railway fences, and road drainage should be comprehensively considered to further optimize the longitudinal and cross-sectional plans of the road. Subsequently, opinions will be solicited from local governments, construction agents, and operating units to clarify the XX plan. "Investment X.XX will be completed this month." RMB XX billion was invested, with a cumulative total of RMB XX billion completed, accounting for 36.4% of the annual target of RMB XX billion. We recommend further strengthening the coordination of external factors and urging relevant local departments to submit XXXX filing materials to ensure the smooth completion of the annual investment.

[0066] S7, merges the content extracted by S6 with the document paragraphs retained by S5 to achieve accurate answers to user search needs.

[0067] In this example, a large model approach is used to integrate the content through prompts. The prompt in this example is, "Extract and integrate the key themes of the <paragraph> content, retaining the weekly and monthly report numbers, context with the words 'key points,' 'important,' and 'difficult points,' and, except for the date, retain the above content strictly according to the original text." The semantic understanding of the large model is used to understand and integrate the captured content.

[0068] The fusion result of this embodiment, that is, the query result, is:

[0069] "Key tasks are progressing. Railway X Institute has been organized to complete the preparation of documents for the design changes for XXX Station (due to changes in the station front and four power engineering projects caused by adjustments to the overall layout). Design and cost consulting opinions have been sought during the process, and the conditions for review are met."

[0070] "Regarding the progress of the control project: 195 meters of excavation have been completed this month for the 1# inclined shaft of the XXX Tunnel, with a cumulative total of 722 meters (53%). 189.5 meters of excavation have been completed this month for the 2# inclined shaft, with a cumulative total of 919.1 meters (69%). 26.4 meters of excavation have been completed this month for the entrance, with a cumulative total of 70.2 meters (4%). This is about two months behind the guidance schedule, mainly due to construction organization and excavation progress indicators failing to meet expectations. 86 meters of excavation have been completed this month for the entrance of the XX Tunnel, with a cumulative total of 529.5 meters (25%). 71 meters of excavation have been completed this month for the small section of the exit work area, with a cumulative total of 172 meters (5%). 4.8 meters of excavation have been completed this month for the large section, with a cumulative total of 130.4 meters (58%). This is an overall delay of about 1.5 months, mainly due to restrictions on nighttime construction and late arrival of equipment. The pile foundations of the main piers of the 219#~220# rotating bridge of the XXX Bridge have been completed, the turntable under the 220# pier has been completed, and the concrete pouring of the 3.5~4.5m depth of the pedestal under the 220# pier has been completed. There is no progress delay."

[0071] "XXXXX's Deputy General Manager visited the XXXX project site (XXX beam fabrication site, XX tunnel entrance, XXX tunnel entrance, and other construction sites) for research. After affirming the project's interim achievements, he put forward four requirements: First, we must coordinate the joint efforts of all parties to overcome difficulties and focus on key projects to resolve difficulties and bottlenecks; second, we must stick to the goals and make every effort to ensure that the first and second phases of the project are completed on schedule; third, we must plan ahead for operational intervention and external environmental remediation; and fourth, we must strictly adhere to the bottom line of safety and quality."

[0072] "Key and difficult tasks XX involve three construction sites: 1. XXXX Mine in XX County, 2. XX Water Plant in XX County, and 3. XX Road in XX Village in XX County." "This month, we completed an investment of XX billion yuan, bringing the annual cumulative investment to XX billion yuan, accounting for 36.4% of the annual target of XX billion yuan. We recommend further strengthening coordination of external factors and urging relevant local units to submit XXXX filing materials to ensure the smooth completion of the annual investment."

[0073] For comparison, we use the deepseek official website model to summarize important work, and ask the same questions as in this example. The results are as follows: Figures 2 to 4 As shown (sensitive information has been blocked in the picture).

[0074] Comparing the query results of our invention with those of DeepSeek reveals that, based on the key focus of the full-process consultation document, the results obtained by our method fully retain the project progress data, fully extract investment data, accurately extract quality-related requirements and speeches, and filter out non-key focus content such as "cleaning up the slag pile problem in the first section; strictly monitoring the progress of lagging work sites" summarized by DeepSeek. This shows that the answer of our invention not only refines the original document content but also retains the key content, avoiding the problem of large models omitting key information during the summary process.

[0075] It can be seen from the precise consultation results obtained in this embodiment that the method of the present invention can accurately extract investment data from multiple dimensions such as the entire line, districts and counties, and sections for investment control in the whole process consultation, thereby realizing overall control of investment consultation and optimization of investment control; for progress control in the whole process consultation, by extracting and analyzing progress information of key projects, real-time analysis and overall control of project progress can be achieved, thereby comprehensively improving the accuracy of the whole process consultation of railway construction projects and effectively promoting the high-quality construction of railway construction projects.

Claims

1. A method for digitizing and retrieving railway whole-process consulting documents based on a large model, characterized in that: The following steps are involved: S1, Classification and grading of railway whole process consulting documents: The railway full-process consulting documents are classified into folders according to project and object types. The folder directories are divided into two levels: project level and classification level. The classification level folders are classified according to the different stages of the railway full-process consulting project; S2, builds the virtual environment required for large model deployment; S3, building a digital retrieval database: sequentially read the contents of all railway full-process consulting documents in each project-level folder, and perform semantic segmentation on the read contents; perform text structured storage encoding on the segmented text contents and store them in a structured database; perform vector storage encoding on the segmented text contents and store them in a vector database; S4: Establish digital extraction standards for railway whole-process consulting documents classified and graded in S1; S5, document retrieval: extract keywords from the search content input by the user, obtain the top 1 to 5 vectors with similarity to the search content from the vector database through the search enhancement generation method, and match the vectors to the corresponding paragraphs in the vector database; perform a noun search on the structured database, compare the search results of the structured database with the matching results of the vector database, and retain consistent results; S6, based on the search results retained in S5, determines which category of documents the retained paragraph belongs to, and then extracts the corresponding content based on the digital extraction criteria constructed in S4; S7, merges the content extracted by S6 with the document paragraphs retained by S5 to achieve accurate answers to user search needs.

2. The large-scale model-based railway whole-process consulting document digitization and retrieval method according to claim 1 is characterized by: In S2, the langchain and transformer framework modules required for the deployment of the large language model are configured in the virtual environment, and the Baichuan large language model is deployed locally.

3. The large-scale model-based railway whole-process consulting document digitization and retrieval method according to claim 1 is characterized by: The formats of the railway full-process consulting documents described in S1 include Word, PDF and Excel formats, and the document types include engineering reporting documents, standard specifications, issuance documents, and full-process consulting weekly and monthly reports.

4. The large-scale model-based railway whole-process consulting document digitization and retrieval method according to claim 3 is characterized by: In S3, the content of the railway full process consultation document is read through Python programming.

5. The method for digitizing and retrieving railway whole-process consulting documents based on a large model according to claim 4 is characterized in that: The segmentation method in S3 is as follows: using a large language model, the prompt method is used to determine the semantic similarity of two adjacent natural paragraphs. If they are similar, the number 1 is output; if they are not similar, the number 0 is output; based on the output number, the paragraphs with similar semantics are merged into one paragraph.

6. The method for digitizing and retrieving railway whole-process consulting documents based on a large model according to claim 3 is characterized in that: In step S3: The method of text structured storage encoding is as follows: the text content after segmentation remains unchanged, and the segmented text is labeled by constructing the tag field "content"; The method of performing vector storage encoding is to encode the segmented text content through a natural language processing model.

7. The method for digitizing and retrieving railway whole-process consulting documents based on a large model according to claim 3 is characterized in that: In step S4: For engineering report documents, the digital extraction standards are: nouns, key parameters, and corresponding data; For standard and regulatory documents, the digital extraction criteria are: proper nouns, corresponding expressions, and corresponding specifications; For the documents of published documents, the digital extraction standards are: keywords, content summary; For the weekly and monthly reports of the full-process consultation, the digital extraction standards are: the issue number of the weekly / monthly report, the context with the words "key points", "important", "difficult points", and numbers other than the date.

8. The method for digitizing and retrieving railway whole-process consulting documents based on a large model according to claim 1 is characterized in that: In S5, a noun search is performed on the structured database through an Elasticsearch module.

Citation Information

Patent Citations

  • Security industry data asset directory construction method and device and security industry data asset directory retrieval method and device

    CN119377450A

  • RAG knowledge base construction method based on layout analysis and query generation

    CN119441507A

  • Railway investigation and design standard specification retrieval method based on large model

    CN119988637A

  • Intelligent question answering system method for air traffic control communication business knowledge

    CN120407745A

  • Work report generation system and method based on multi-path retrieval and large model

    CN120596650A