Railway investigation and design standard specification retrieval method based on large model

By building a knowledge graph for railway survey and design and using large models to extract key information, the problem of inefficient search methods for existing railway survey and design specifications is solved, and the efficiency, accuracy and stability of standard and specification search is achieved, and the quality and safety of project are ensured.

CN119988637APending Publication Date: 2025-05-13CHINA RAILWAY DESIGN GRP CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411821612.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing railway survey and design specification search methods are inefficient and rely on manual reviews, making it difficult to retrieve the required standard documents comprehensively and accurately, and there is a lack of effective semantic similarity judgment methods.

Method used

The railway survey and design standard specification search method is adopted based on large models. By extracting standard specification clauses in the specification document, a knowledge graph for railway survey and design is constructed, keywords and overview are extracted using the large model, and multiple recalls and sorts are combined with knowledge graphs and vector databases to achieve efficient screening and sorting of standard clauses.

Benefits of technology

It improves the comprehensiveness, accuracy and efficiency of standard and specification search results, ensures the stability and semantic correlation of search results, can deliver design results on schedule, reduces the project construction cycle and cost, and ensures the quality and safety of the project.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988637A_ABST
    Figure CN119988637A_ABST
Patent Text Reader

Abstract

The invention discloses a railway investigation and design standard specification retrieval method based on a large model. The railway investigation and design standard specification retrieval method based on the large model comprises the steps that 1, the railway investigation and design standard specification retrieval method based on the large model is provided; s2, utilizing the large model to construct a knowledge graph for railway investigation and design standard specification retrieval; s3, using a vector database to carry out vectorization storage on the article information; s4, extracting key information in the user question by using the large model; s5, screening the standard specification articles through a multi-path recall mechanism; and S6, performing secondary sorting by using the large model and outputting a result. According to the method, accurate retrieval of railway investigation and design standard specifications is realized, and effective technical guarantee is provided for improving the quality of railway investigation and design results; constructing a knowledge graph for railway survey design standard specification retrieval; and a multi-path recall and large model secondary sorting method is adopted, so that the comprehensiveness, the accuracy and the stability of standard article retrieval are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of rail transit information technology, and in particular to a large-model-based railway survey and design standard specification retrieval method. Background Art

[0002] Improving the quality of railway survey and design results is an important goal of railway engineering. During the survey and design process, as well as when the survey and design results are delivered, quality control of the design results and quality feedback play a vital role in ensuring the smooth completion and even safe operation of the railway project. At present, the quality control of railway survey and design results mostly relies on manual review of each item, manually reading a large number of relevant specifications, checking whether the design results meet the requirements of the specifications, and finally forming a review report. This form of review is inefficient and has many shortcomings. For example, manually reading the text of the specification documents one by one will consume a lot of human resources and time, resulting in errors in the comparison of the specification clauses, and the low review efficiency will lead to the failure to deliver the design results on time. To a certain extent, it affects the construction period and construction cost of the project, and even endangers the quality and safety of the project.

[0003] Railway survey and design specifications contain a large amount of engineers' experience and knowledge and researchers' experimental knowledge, which are important standards for controlling the quality of design results. How to accurately and comprehensively retrieve the required survey and design specifications according to user needs has become a demand. The existing retrieval methods are mainly based on NLP methods, which retrieve the text of the specification documents that are most similar to the user's questions from the aspects of word frequency and vector similarity. However, these methods ignore the reference relationship between clauses, making the retrieval results incomplete. On the other hand, when evaluating the relevance to user questions, the evaluation method is single and lacks an effective semantic level similarity judgment method. Summary of the invention

[0004] In order to solve the problems in the background technology, the present invention provides a railway survey and design standard specification retrieval method based on a large model, which has comprehensive, accurate, reliable and efficient retrieval results.

[0005] To this end, the present invention adopts the following technical solutions:

[0006] A railway survey and design standard specification retrieval method based on a large model includes the following steps:

[0007] S1, extract standard specification clauses from specification documents:

[0008] First, use Python's docx library to process the full text of the railway survey and design specification document in doc or docx format, and obtain the content of each paragraph one by one;

[0009] Then, a regular expression is used to determine whether the article number appears at the beginning of the paragraph text;

[0010] If it appears, record a clause information and use it as the current standard specification clause;

[0011] If it does not appear, and there is a current standard specification, then the text is added to the content of the current standard specification; otherwise, no processing is performed;

[0012] Finally, a clause list of the specification document is obtained, wherein the clause list includes all standard specification clauses marked with clause numbers;

[0013] S2, using the big model to build a knowledge graph for railway survey and design standards and specifications retrieval, including the following steps:

[0014] S21, using the large model prompt word engineering to extract keywords for each standard specification clause in the clause list in S1;

[0015] S22, using a large model prompt word project to obtain an overview of the standard specification clauses;

[0016] S23, using the large model prompt word project, by writing corresponding prompt words, extracting the reference relationship of each standard specification article from the article;

[0017] S24, in order to improve the accuracy of the reference relationship, first manually process the reference relationship of some of the standard specification articles, and combine the prompt words in S23 to construct a large model fine-tuning dataset in the form of question and answer pairs; then use the large model fine-tuning dataset and the QLora fine-tuning method to fine-tune the large model to obtain a reference relationship large model, and use the reference relationship large model to extract the remaining standard specification articles to obtain the reference relationship of the corresponding standard specification articles;

[0018] S25, storing the standard specification clauses and the reference relationship obtained in S24 in a graph database to obtain a knowledge graph for standard specification retrieval;

[0019] S3 uses a vector database to vectorize and store the article information:

[0020] The keywords and summaries in S2 are converted into vectors, and the vectorized keywords and summaries are stored in a vector database to obtain a clause keyword vector database and a clause summary vector database;

[0021] S4, using the big model to extract key information from user questions:

[0022] Using the large model prompt word project, key information is extracted from the text input by the user, and the key information includes: project objects and object attributes;

[0023] S5, screening of standard specification clauses through a multi-channel recall mechanism:

[0024] For the engineering objects and object attributes extracted in step S4, the top k approximate keywords are obtained from the article keyword vector database obtained in step S3 according to the vector distance, and then a graph database query statement is constructed for the approximate keywords. The standard specification articles and their reference relationships containing these keywords are screened out from the knowledge graph obtained in S2, and the overall duplication is removed as the initial screening result of the articles;

[0025] Sort the standard and regulatory clauses in the preliminary screening results by using the BM25 text matching algorithm, select the top n standard and regulatory clauses, and use them as the results of the graph query recall;

[0026] Then, the text input by the user in S4 is directly vectorized using the large model, and the first n standard specification clauses most relevant to the text input by the user are searched from the summary vector database of standard specification clauses obtained in step S3, and the searched content is used as the result of recall based on vector similarity;

[0027] Finally, the results based on graph query recall and the results based on vector similarity recall are deduplicated, and the remaining standard specification clauses are used as the results of standard clause screening;

[0028] S6, use the large model to perform secondary sorting and output the results:

[0029] Use the large model and the corresponding prompt words to calculate the correlation between the standard articles screened out in step S5 and the text entered by the user, and calculate the similarity value for each standard article after multiple recalls; then sort all the standard articles after multiple recalls according to the similarity value, and return the first n articles as the final query result.

[0030] The regular expression in S1 is "'^[0-9]+\.[0-9]+\.[0-9a-zA-Z]+'"; the article number includes numbers, decimal points, and uppercase and lowercase letters.

[0031] The reference relationship in step S2 includes: the name of the referenced external standard document, the article number of the referenced external standard specification text and other article numbers in the referenced current standard document; the external standard document represents other standard documents other than the current standard document.

[0032] The similarity value in step S6 is between 0 and 100.

[0033] In step S25, the reference relationship obtained in step S24, the keywords in step S21 and the overview of the standard specification article in step S22 are used to construct three types of triple relationships for each standard specification article: <Article (overview), including, keywords>, <Article (overview), reference, article (overview)>, <Article (overview), reference, standard document> and <Standard document, including, article (overview)>, thereby forming a knowledge graph for standard specification retrieval.

[0034] Preferably, the prompt word is written using Markdown syntax.

[0035] Preferably, in step S4, if the information extracted by the large model includes text other than key information, a regular expression '(?<={).*? (?=})' is used to filter out irrelevant information to ensure that the large model outputs accurate key information.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] 1. The method of the present invention focuses on the reference relationship between standard specification articles, and uses large model prompt words and fine-tuning technology to extract the reference relationship between standard document articles, as well as the keywords and overviews contained in the articles, and constructs a knowledge graph for standard specification retrieval, improves the comprehensiveness and accuracy of standard specification retrieval results, and provides a reliable source of information for subsequent retrieval processes.

[0038] 2. The method of the present invention comprehensively utilizes multiple retrieval and sorting methods such as knowledge graph query, BM25 sorting and vector similarity retrieval to perform multi-way recall retrieval and sorting of standard specification clauses, ensuring the stability, comprehensiveness and accuracy of standard specification retrieval results.

[0039] 3. The method of the present invention utilizes the ability of the large model to understand natural language, and uses the prompt word technology to score the relevance of the initial screening results of standard and specification clauses to achieve secondary sorting, ensuring from the level of semantic understanding that standard and specification clauses that are more relevant to user questions can be displayed first. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 A flowchart of the standard specification retrieval method of the present invention;

[0041] Figure 2 It is the overall framework of the standard specification retrieval method of the present invention;

[0042] Figure 3 An example of a standard canonical knowledge graph constructed for the present invention. DETAILED DESCRIPTION

[0043] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0044] like Figure 1 and Figure 2 As shown, the railway survey and design standard specification retrieval method based on a large model of the present invention comprises the following steps:

[0045] S1, extract standard specification clauses from specification documents:

[0046] First, use Python's docx library to process the railway survey and design specification document in doc or docx format and obtain the content of each paragraph one by one.

[0047] Then, use regular expressions to filter. The regular expression used is: "'^[0-9]+\.[0-9]+\.[0-9a-zA-Z]+'". The meaning of the expression is: filter whether the article number appears at the beginning of the text of the paragraph. The article number includes numbers, decimal points, and uppercase and lowercase letters;

[0048] If there is an article number in the text, a new article information will be recorded and used as the current article;

[0049] If the paragraph does not contain a clause number and the current standard specification clause exists, the paragraph is added to the content of the current standard specification clause; otherwise, no processing is performed. The above process is repeated until all paragraphs in the standard document are read.

[0050] Finally, a clause list of the specification document is obtained, which includes all standard specification clauses marked with clause numbers.

[0051] S2, using the big model to build a knowledge graph for railway survey and design standards and specifications retrieval, including the following steps:

[0052] S21, use the big model prompt word project to extract keywords from the standard specification clauses obtained in step S1. In the prompt words, the extraction principle of the big model keyword extraction and the output form of the results are agreed upon, and a small number of examples are given to guide the big model to extract keywords from the standard specification clauses. The specific prompt words are as follows:

[0053] -------------

[0054] Now you are a word segmentation expert, requirements:

[0055] 1. Your answer only contains word segmentation results;

[0056] 2. The extracted words are all engineering objects or object attributes;

[0057] 3. The word segmentation results are output in the format of [word 1; word 2; word 3]. Here are some examples of word segmentation. Please refer to these examples for word segmentation:

[0058] (1) Example 1:

[0059] Input: 3.4.5 The fire protection distance between civil buildings with a building height greater than 100m and adjacent buildings shall comply with Articles 3.4.5, 3.5.3 and 4.2.1 of this Code.

[0060] Output: [building height; civil building; adjacent building; fire protection distance].

[0061] (2) Example 2:

[0062] Input: Storage tanks with the same or similar fire hazard categories should be arranged in each firebreak. Boiling oil storage tanks should not be arranged in the same firebreak as non-boiling oil storage tanks. Above-ground and semi-underground storage tanks should not be arranged in the same firebreak as underground storage tanks.

[0063] Output: [firebreak; fire hazard; storage tank; boiling oil storage tank; non-boiling oil storage tank; above-ground tank area; semi-underground tank area; underground tank].

[0064] (3) Example 3:

[0065] Input: 4.3.7 The fire protection distance between liquid hydrogen and liquid ammonia storage tanks and buildings, storage tanks, yards, etc. can be determined by reducing the fire protection distance of liquefied petroleum gas storage tanks of corresponding volume in Article 4.4.1 of this Code by 25%.

[0066] Output: [liquid hydrogen; liquid ammonia storage tank; building; storage tank; yard; fire protection distance; liquefied petroleum gas storage tank].

[0067] When the input is {input}, what should the output be?

[0068] -------------

[0069] When using the above prompt words to extract keywords, you only need to replace "{input}" in the prompt words with the standard specification clauses of the keywords to be extracted, and send the entire prompt words to the big model, and you can get the keywords output in the specified format.

[0070] S22, using the large model prompt word engineering, briefly summarize the main contents of the standard specification provisions. The content after the summary is required to not exceed 200 words, and include the engineering objects and object attributes concerned in the original standard specification provisions, so as to obtain an overview of the standard specification provisions.

[0071] S23, based on the relevant technologies of large model prompt word engineering, the prompt words are written, and the reference relationship of each standard specification article is extracted from the article. The reference relationship includes: the name of the referenced external standard document, the article number of the referenced external standard specification text, and the article number of the referenced other articles in the current standard document. The external standard document refers to other standard documents other than the current standard document. The relevant technology includes writing prompt words in Markdown syntax. The specific prompt word content is as follows: -------------

[0072] #background

[0073] You are an information extraction expert. Please extract the specification documents referenced by the statement and the related standard specification clauses from the known standard specification clauses and compose the output. The output format is:

[0074] {"Normative document":["Document 1":{"Related articles":["Article 1 of Document 1","Article 2 of Document 1'"]},"Document 2":{"Related articles":["Article 1 of Document 2","Article 2 of Document 2'"]}],"Related articles":["Article 1 of this document","Article 2 of this document"]}.

[0075] #Require

[0076] The following requirements must be met when extracting information:

[0077] ##1. Each selected regulatory document and standard regulatory clause must be the original text of known information;

[0078] ##2. If the known information does not contain the standard specification article number, specification document, and related articles, then the output is: {"standard document": [],"related articles": []}.

[0079] ##3. Please output the result in json format.

[0080] ##4. The involved clauses must appear in the text in the form of "Article xxx" or "Article xxx, yyy", where "xxx" and "yyy" are the results of the involved clauses to be extracted. If there is no such description, the "involved clauses" in the output will be empty. In addition, if the involved clauses are repeated, only one will be retained.

[0081] ##5. The standard document must start with the book title quotation marks, including the book title quotation marks and the content therein and the code that follows it. If the input standard document does not include the standard document with the book title quotation marks, the "standard document" in the output will be empty.

[0082] ##6. If the input standard specification clauses mention referenced specification documents, such as specific clauses of "Document 1", these clauses will be added to the list of related clauses of "Document 1" as "Clause xxx of Document 1".

[0083] ##7. If multiple standard documents with the same name are extracted, only one will be retained.

[0084] ##8. When the referenced standard specification clauses are expressed as a range such as "Article xxx to Article yyy", please calculate and list all the standard specification clauses within the range one by one as the result of the clauses involved.

[0085] ##9. The output results should strictly conform to the Json format.

[0086] #example

[0087] When the known information is: "6.1.2 When the waiting area and distribution hall of a railway passenger station meet the following conditions, the building area of ​​each fire compartment should not be greater than 10,000 m':

[0088] 1 It is located on the first floor, a single-story elevated floor, or the second floor with half the number of direct external evacuation exits and an indoor enclosed stairwell.

[0089] 2 An automatic sprinkler fire extinguishing system, smoke exhaust facilities and automatic fire alarm system shall be installed in accordance with the provisions of Article 3.4.1 of this Code.

[0090] 3 The interior decoration design complies with the relevant provisions of Articles 1.2.5 and 1.5.6 of the "Code for Fire Protection Design of Interior Decoration of Buildings" GB50222. ", your output should be:

[0091] {"Specification document":["Code for Fire Protection Design of Building Interior Decoration GB50222":{"Related clauses":["1.2.5","1.5.6"]}],"Related clauses":["3.4.1"]}".

[0092] Based on the above, when the input is "{input}", what should your output be? Please think step by step and just output the result.

[0093] -------------

[0094] When using the above prompt words to extract reference relationships, "{input}" needs to be replaced with the text of the standard specification clause to be processed. In addition, the above prompt words use descriptions that conform to Markdown syntax to represent the hierarchical relationship of the standard documents in the prompt words, so that the large model can understand the user's intention more accurately.

[0095] It can be seen that through the above prompt words, the large model can be used to extract and obtain the reference relationship between each standard specification article and the articles in this specification and other specification documents, names and article numbers.

[0096] S24, in order to further improve the accuracy of extracting reference relationships between standard specification articles, we first use the prompt words in step S23 as instructions and manually organize them to construct a fine-tuning dataset. The dataset is a series of question-answer pairs stored in json format. The specific form of the question-answer pairs is as follows:

[0097] ------------

[0098] {

[0099] "instruction": "You are an information extraction expert. Please extract the normative documents and related articles referenced by the known standard specification clauses and form the output. The output format is: {"normative document": ["document 1": {"related articles": ["article 1 of document 1", "article 2 of document 1'"]},"document 2": {"related articles": ["article 1 of document 2", "article 2 of document 2'"]}],"related articles": ["article 1 of this document", "article 2 of this document"]}.

[0100] #The following requirements must be met when extracting information:

[0101] ##1. Each selected regulatory document and standard regulatory clause must be the original text of known information;

[0102] ##2. If there is no article number, standard document and related articles in the known information, then output: {"standard document": [],"related articles": []}.

[0103] ##3. Please output the result in json format.

[0104] ##4. The involved clauses must appear in the text in the form of "Article xxx" or "Article xxx, yyy", where "xxx" and "yyy" are the results of the involved clauses to be extracted. If there is no such description, the "involved clauses" in the output will be empty. In addition, if the involved clauses are repeated, only one will be retained.

[0105] ##5. The standard document must start with the book title quotation marks, including the book title quotation marks and the content therein and the code that follows it. If the input standard document does not include the standard document with the book title quotation marks, the "standard document" in the output will be empty.

[0106] ##6. If the input standard specification clauses mention referenced specification documents, such as specific clauses of "Document 1", these clauses will be added as "Clause x of Document 1" to the list of related clauses of "Document 1".

[0107] ##7. If multiple standard documents with the same name are extracted, only one will be retained.

[0108] ##8. When the referenced standard specification clauses are expressed as a range such as "Article xxx to Article yyy", please calculate and list all the standard specification clauses within the range one by one as the result of the clauses involved.

[0109] ##9. The output result should strictly conform to the Json format. ";

[0110] "input": "Considering that some buildings have high decoration standards and need to use combustible materials for decoration, refer to 6.5.2 of this code. In order to meet actual needs without reducing overall safety performance, it is stipulated that fire-fighting facilities should be installed to make up for the problem of insufficient combustion level of decoration materials. According to the provisions of Articles 5.6.1 and 5.2.3 of the American standard "Life Safety Code" NFPA 101, if automatic fire extinguishing measures are taken, the combustion performance level of the decoration materials used can be reduced by one level. This article is formulated with reference to the above provisions.";

[0111] "output": "{"Specification document": ["NFPA 101 Life Safety Code": {"Related clauses": ["5.6.1", "5.2.3"]}],"Related clauses": ["6.5.2"]}"

[0112] },

[0113] ------------

[0114] The 7b Qwen2.5 large model was fine-tuned by LoRA using a large model fine-tuning dataset containing several question-answer pairs to obtain a large model for extracting reference relationships between standard and specification clauses. The reference relationship large model was then used to extract the remaining standard and specification clauses to obtain the reference relationships of the corresponding standard and specification clauses.

[0115] S25, save as knowledge graph:

[0116] By using the reference relationship of the standard specification clauses obtained in step S24, the keywords in step S21 and the overview of the standard specification clauses in step S22, three types of triple relationships can be constructed for each standard specification clause: <clause (overview), including, keywords>, <clause (overview), reference, clause (overview)>, <clause (overview), reference, specification document> and <specification document, including, clause (overview)>, thereby forming a knowledge graph for standard specification retrieval, such as Figure 3 shown.

[0117] S3 uses a vector database to vectorize and store the article information:

[0118] All keywords of all standard specification articles and the overview of standard specification articles are read from the standard specification knowledge graph obtained in step S2, and the extracted keywords and article overviews are embedded in text using the large model, and the keywords and the overviews of standard specification articles in natural language are converted into vectors. Then, the vectorized keywords and the overviews of standard specification articles are stored respectively using the vector database to obtain the article keyword vector database and the article overview vector database.

[0119] S4, using the big model to extract key information from user questions:

[0120] Using the large model prompt words, extract the key information involved in the text entered by the user. The key information includes engineering objects and object attributes. The prompt words used are as follows:

[0121] -------------

[0122] #Background You are an expert in sentence component analysis. Please extract the main engineering objects and their attributes described in the sentence from the input sentence.

[0123] #Require

[0124] ##1. Please output the extraction results in json format.

[0125] ##2. The extracted "engineering object" must be an engineering entity. If the input does not contain the attributes of the engineering object, the attributes in the output should be "".

[0126] ##3. The extracted "project objects" and "attributes" must be the original text in the input statement.

[0127] #Example

[0128] ##Example 1 When the input is "How should the fire partition of a railway station be designed?", your output should be

[0129] "{"Project object": railway station building, "Attribute": fire partition}"

[0130] ##Example 2 When the input is "How should a railway station be designed", your output should be "{"Project Object":"Railway Station,"Attribute":""}.

[0131] ##Example 3 When the input is "During the design of railway station building, if the fire partition is designed on the first floor, how should the fire partition area be calculated?", your output should be "{"Project Object":"Railway Station Building,"Attribute":"Fire Partition Area"}.

[0132] Please think about it step by step below. When my input is "{input}", what should your output be?

[0133] -------------

[0134] When using the above prompt words to extract engineering objects and object attributes, you need to first replace "{input}" with the user's question and input the replaced prompt words into the large model. Since the description "...please think step by step..." is added to the prompt words using the thinking chain technology, the output may contain text other than the target Json text. You need to use a regular expression to filter the redundant text. The regular expression used is: '(?<={).*? (?=})'. Through this regular expression, you can get the target Json text from the output of the large model.

[0135] S5, screening the standard specification clauses through a multi-channel recall mechanism, including the following steps:

[0136] S51, obtain the results based on graph query recall:

[0137] First, the engineering objects and object attributes extracted in step S4 are vectorized using the large model. The top k similar keywords are queried in the article keyword vector database obtained in S3 using the vectorized engineering objects and object attributes. Then, for each queried keyword, a graph query statement is generated, and all standard specification articles containing the keyword are queried from the recognition graph constructed in step S2. Assuming that a certain keyword is "Keyword 1", taking Cypher language as an example, the generated graph query statement should be: MATCH (a: Rule) - [: `Include`] -> (b: Keyword {text: 'Keyword 1'}) RETURN a. Among them, "(a: Rule)" represents the node with the category of "Standard Article" in the knowledge graph, and "(b: Keyword {text: 'Keyword 1'})" represents the node with the keyword text of "Keyword 1" in the keyword node of the knowledge graph. Then use the graph query statement to query the clauses referenced by the above standard specification clauses. Taking Cypher language as an example, the statement to query the clauses referenced by a certain standard specification clause should be: MATCH (a: Rule)-[:`reference`]->(b: Rule) RETURN b. Among them, b is the clause referenced by standard specification clause a. Finally, all the queried standard specification clauses are deduplicated as a whole, and the preliminary screening results of the clauses that may be related to the user's problem can be preliminarily screened.

[0138] Then, the BM25 algorithm is used to sort the preliminary screening results of the clauses in S51, and the top n standard specification clauses are selected as the results based on graph query recall.

[0139] S52, obtain the result based on vector similarity recall:

[0140] The large model is then used to directly vectorize the text input by the user, and the first n standard specification clauses most relevant to the user's question are queried from the clause summary vector database obtained in step S3. The queried content is used as the result of vector similarity-based recall.

[0141] S53, deduplicate the 2n standard specification clauses obtained from S51 and S52 to obtain the standard specification clauses after multi-channel recall.

[0142] S6, secondary sorting using a large model:

[0143] The prompt word project is used to score the relevance of the standard specification clauses after the multi-channel recall obtained in step S5 and the text input by the user. The prompt words used for scoring are as follows:

[0144] -------------

[0145] Please give each of the following content a relevance score between 0 and 100 based on its relevance to "{query}", with 100 being most relevant and 0 being completely irrelevant.

[0146] #Require

[0147] 1. The output results are sorted in order of relevance, and the original sequence number of the data is retained;

[0148] 2. The output results are organized in json format, for example: {[{entry number: "original entry number", relevance score: "relevance score of this entry"}, {entry number: "original entry number", relevance score: "relevance score of this entry"}]}

[0149] #Content Items

[0150] ##1: {rule1}

[0151] ##2: {rule2}

[0152] ##3: {rule3}

[0153] ##4: {rule4} ........

[0155] ##n:{rulem}

[0156] -------------

[0157] When using the above-mentioned prompt words to score the relevance of the standard specification clauses after multiple recalls, it is necessary to replace "{query}" with the text entered by the user, and "rule1~m" with the original texts of m different standard specification clauses after multiple recalls or their corresponding summaries. The setting of m should ensure that the total length of the entire prompt word does not exceed the context length that the large model used can accept.

[0158] After obtaining the relevance scores between all the standard specification clauses and the user questions after multi-way recall, the standard specification clauses are sorted according to the relevance scores, and finally n standard specification clauses that are most relevant to the user questions are output as the final query results.

Claims

1. A railway survey and design standard specification retrieval method based on a large model, characterized in that: The following steps are involved: S1, extract standard specification clauses from specification documents: First, use Python's docx library to process the full text of the railway survey and design specification document in doc or docx format, and obtain the content of each paragraph one by one; Then, a regular expression is used to determine whether the article number appears at the beginning of the paragraph text; If it appears, record a clause information and use it as the current standard specification clause; If it does not appear, and there is a current standard specification, then the text is added to the content of the current standard specification; otherwise, no processing is performed; Finally, a clause list of the specification document is obtained, wherein the clause list includes all standard specification clauses marked with clause numbers; S2, using the big model to build a knowledge graph for railway survey and design standards and specifications retrieval, including the following steps: S21, using the large model prompt word engineering to extract keywords for each standard specification clause in the clause list in S1; S22, using a large model prompt word project to obtain an overview of the standard specification clauses; S23, using the large model prompt word project, by writing corresponding prompt words, extracting the reference relationship of each standard specification article from the article; S24, in order to improve the accuracy of the reference relationship, first manually process the reference relationship of some of the standard specification articles, and combine the prompt words in S23 to construct a large model fine-tuning dataset in the form of question and answer pairs; then use the large model fine-tuning dataset and the QLora fine-tuning method to fine-tune the large model to obtain a reference relationship large model, and use the reference relationship large model to extract the remaining standard specification articles to obtain the reference relationship of the corresponding standard specification articles; S25, storing the standard specification clauses and the reference relationship obtained in S24 in a graph database to obtain a knowledge graph for standard specification retrieval; S3 uses a vector database to vectorize and store the article information: The keywords and summaries in S2 are converted into vectors, and the vectorized keywords and summaries are stored in a vector database to obtain a clause keyword vector database and a clause summary vector database; S4, using the big model to extract key information from user questions: Using the large model prompt word project, key information is extracted from the text input by the user, and the key information includes: project objects and object attributes; S5, screening of standard specification clauses through a multi-channel recall mechanism: For the engineering objects and object attributes extracted in step S4, the top k approximate keywords are obtained from the article keyword vector database obtained in step S3 according to the vector distance, and then a graph database query statement is constructed for the approximate keywords. The standard specification articles and their reference relationships containing these keywords are screened out from the knowledge graph obtained in S2, and the overall duplication is removed as the initial screening result of the articles; Sort the standard and regulatory clauses in the preliminary screening results by using the BM25 text matching algorithm, select the top n standard and regulatory clauses, and use them as the results of the graph query recall; Then, the text input by the user in S4 is directly vectorized using the large model, and the first n standard specification clauses most relevant to the text input by the user are searched from the summary vector database of standard specification clauses obtained in step S3, and the searched content is used as the result of recall based on vector similarity; Finally, the results based on graph query recall and the results based on vector similarity recall are deduplicated, and the remaining standard specification clauses are used as the results of standard clause screening; S6, use the large model to perform secondary sorting and output the results: Use the large model and the corresponding prompt words to calculate the correlation between the standard articles screened out in step S5 and the text entered by the user, and calculate the similarity value for each standard article after multiple recalls; then sort all the standard articles after multiple recalls according to the similarity value, and return the first n articles as the final query result.

2. The large model-based railway survey and design standard specification retrieval method according to claim 1 is characterized by: The regular expression in S1 is "'^[0-9]+\.[0-9]+\.[0-9a-zA-Z]+'"; the article number includes numbers, decimal points, and uppercase and lowercase letters.

3. The railway survey and design standard specification retrieval method based on a large model according to claim 1 is characterized by: The reference relationship in step S2 includes: the name of the referenced external standard document, the article number of the referenced external standard specification text and other article numbers in the referenced current standard document; the external standard document represents other standard documents other than the current standard document.

4. The railway survey and design standard specification retrieval method based on a large model according to claim 1 is characterized by: The similarity value in step S6 is between 0 and 100.

5. The railway survey and design standard specification retrieval method based on a large model according to claim 1 is characterized by: In step S25, the reference relationship obtained in step S24, the keywords in step S21 and the overview of the standard specification article in step S22 are used to construct three types of triple relationships for each standard specification article: <Article (overview), including, keywords>, <Article (overview), reference, article (overview)>, <Article (overview), reference, standard document> and <Standard document, including, article (overview)>, thereby forming a knowledge graph for standard specification retrieval.

6. The railway survey and design standard specification retrieval method based on a large model according to claim 1 is characterized by: The prompt words are written using Markdown syntax.

7. The railway survey and design standard specification retrieval method based on a large model according to claim 1 is characterized by: In step S4, if the information extracted by the large model includes text other than key information, the regular expression '(?<={).*? (?=})' is used to filter out irrelevant information to ensure that the large model outputs accurate key information.

Citation Information

Cited By

  • Railway whole-process consultation document digitization and retrieval method based on large model

    CN120821758A

  • Hybrid reasoning BIM (Building Information Modeling) design achievement auditing method for complex standard specification provisions

    CN121352012A

  • Railway survey and design standard specification retrieval method based on large model

    WO2026123845A1