A rail transit knowledge base construction method and system based on large language model

By building a rail transit knowledge base based on large language models, the problem of weak inference ability of large language models in the field of rail transit is solved, efficient and accurate knowledge retrieval services are achieved, and the rail transit industry's demand for professional information is met.

CN118733740BActive Publication Date: 2025-05-13BEIJING URBAN CONSTRUCTION DESIGN & DEVELOPMENT GROUP CO LIMITED
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410943012.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-15
Publication Date
2025-05-13
Estimated Expiration
2044-07-15

AI Technical Summary

Technical Problem

The general-purpose large-language model has weak inference ability in the field of rail transit and cannot obtain the latest engineering data and construction plans in real time, resulting in low accuracy, high cost and long time for fine-tuning, and does not meet the promotion conditions. The traditional knowledge base question and answer system has shortcomings in semantic understanding and intelligent processing of retrieval results, and its query accuracy and efficiency are low.

Method used

By obtaining text data in the field of rail transit, using text extraction tools to convert it into plain text files, perform word segmentation, part-of-speech annotation and keyword extraction, construct an inverted index, and generate a knowledge base. Preprocess the user's Query, extract keywords, vectorize them, calculate the similarity to keywords in the knowledge base, and locate the most relevant text block ID.

Benefits of technology

An intelligent question-and-answer platform has been realized, which improves the efficiency and accuracy of question-and-answer, can provide more intelligent and efficient knowledge retrieval services, meets the rail transit industry's needs for professional and accurate engineering information, and supports the quality and safety management of engineering practitioners.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118733740B_ABST
    Figure CN118733740B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for constructing a rail transit knowledge base based on a large language model. The method comprises: acquiring text data in the field of rail transit, converting the text data in different formats into plain text files through a text extraction tool, removing irrelevant formatting marks, headers and footers, charts and / or images in the plain text files, performing word segmentation and part-of-speech tagging on the plain text files, and generating multiple text blocks; extracting keywords, setting a filtering threshold to exclude keywords whose position weights are lower than the filtering threshold, and sorting the remaining keywords according to the position weights, generating a keyword list, vectorizing the keyword list, constructing an inverted index, associating keyword vectors with text block IDs, and generating a knowledge base; extracting keywords of user queries, converting the keywords of user queries into vector representations, calculating the similarity between the vectors of the keywords of the user queries and the keyword vectors in the knowledge base, and locating the most relevant text block IDs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of knowledge base construction, and more specifically, relates to a rail transit knowledge base construction method and system based on a large language model. Background Art

[0002] The general large language model has powerful reasoning logic and can achieve fast and accurate reasoning for simple problems. However, in the rail transit industry, the project volume is large, the professions are complex, and the engineering characteristics are obvious. The large model cannot obtain the latest engineering data construction plan in real time, resulting in the large model's weak reasoning ability and low accuracy for complex problems in the rail transit field. Although the reasoning ability can be improved through fine-tuning technology combined with data sets, the process requires rich experience and huge computing power, high cost, long time cycle, and is not suitable for promotion. The traditional knowledge base question-answering system has gone through technologies such as database query, file search, and knowledge graph. Although the accuracy and ease of use have been improved, the semantic understanding of users and the search results lack intelligent processing, and the query accuracy and efficiency are low.

[0003] Although general large language models perform well in dealing with simple problems and have fast and accurate reasoning capabilities, their reasoning capabilities are limited when facing the rail transit industry, a field with strong professionalism, large engineering volume, and rapid data updates. Rail transit projects involve many professional fields, each of which has its own complexity, and engineering data and construction plans are constantly updated, which poses a challenge to the real-time data acquisition capabilities of large models. Although the reasoning ability of the model can be improved by fine-tuning technology and combining specific data sets, this process requires not only professional knowledge and rich experience, but also huge computing resources, resulting in high costs and long cycles, making it difficult to widely apply.

[0004] Although traditional knowledge base question-answering systems have made certain progress in database query, file search, knowledge graph, etc., and improved the accuracy and ease of use of the system, they still have shortcomings in semantic understanding and intelligent processing of retrieval results, resulting in unsatisfactory query accuracy and efficiency. The LLM+local knowledge base solution has the problem of low Embedding-Search recall accuracy in the scenario of multi-knowledge point aggregation processing. Summary of the invention

[0005] In order to solve the above technical problems, the present invention proposes a method for constructing a rail transit knowledge base based on a large language model, comprising:

[0006] Acquire text data in the field of rail transit, convert the text data in different formats into plain text files through a text extraction tool, remove irrelevant formatting marks, headers and footers, charts and / or images in the plain text files, and perform word segmentation and part-of-speech tagging on the plain text files to generate multiple text blocks;

[0007] Perform word segmentation, word frequency statistics and position weight analysis on the text block, extract keywords, set a filtering threshold to exclude keywords with a position weight lower than the filtering threshold, sort the remaining keywords by position weight, generate a keyword list, vectorize the keyword list, build an inverted index, associate the keyword vector with the text block ID, and generate a knowledge base;

[0008] Preprocess the user query, remove punctuation, lowercase, remove stop words, and extract the keywords of the user query. Convert the keywords of the user query into vector representation, calculate the similarity between the vector of the keyword of the user query and the keyword vector in the knowledge base, locate the most relevant text block ID, and provide the text block list corresponding to the most relevant text block ID to the user.

[0009] Furthermore, the text data in different formats are converted into a plain text file through a text extraction tool, wherein the plain text file is a plain text file encoded in UTF-8.

[0010] Furthermore, before performing word segmentation and part-of-speech tagging, the method further includes: converting all text in the plain text file into lowercase.

[0011] Furthermore, after the plain text file is segmented and POS tagged, the method further includes: detecting and removing duplicate data of the tagged text file through a hash algorithm based on text fingerprint, so as to ensure the uniqueness of the data in the knowledge base.

[0012] Furthermore, after extracting the keywords of the user's query, it also includes: identifying the intention of the user's query through a deep neural network, and determining the type of knowledge base that the user wants to retrieve.

[0013] The present invention also proposes a rail transit knowledge base construction system based on a large language model, comprising:

[0014] A text block generation module is used to obtain text data in the field of rail transit, convert the text data in different formats into plain text files through a text extraction tool, remove irrelevant formatting marks, headers and footers, charts and / or images in the plain text files, and perform word segmentation and part-of-speech tagging on the plain text files to generate multiple text blocks;

[0015] Generate a knowledge base module, which is used to perform word segmentation, word frequency statistics and position weight analysis on the text block, extract keywords, set a filtering threshold to exclude keywords with a position weight lower than the filtering threshold, and sort the remaining keywords by position weight, generate a keyword list, vectorize the keyword list, build an inverted index, associate the keyword vector with the text block ID, and generate a knowledge base;

[0016] A knowledge module is provided to pre-process user queries, remove punctuation, lowercase, remove stop words, extract keywords of user queries, convert keywords of user queries into vector representations, calculate the similarity between the vectors of keywords of user queries and the keyword vectors in the knowledge base, locate the most relevant text block ID, and provide the user with a list of text blocks corresponding to the most relevant text block IDs.

[0017] Furthermore, the text data in different formats are converted into a plain text file through a text extraction tool, wherein the plain text file is a plain text file encoded in UTF-8.

[0018] Furthermore, before performing word segmentation and part-of-speech tagging, the method further includes: converting all text in the plain text file into lowercase.

[0019] Furthermore, after the plain text file is segmented and POS tagged, the method further includes: detecting and removing duplicate data of the tagged text file through a hash algorithm based on text fingerprint, so as to ensure the uniqueness of the data in the knowledge base.

[0020] Furthermore, after extracting the keywords of the user's query, it also includes: identifying the intention of the user's query through a deep neural network, and determining the type of knowledge base that the user wants to retrieve.

[0021] Compared with the prior art, the above technical solution conceived by the present invention has the following beneficial effects:

[0022] The present invention uses advanced technologies such as natural language processing (NLP) technology, vector databases and large language models to create an intelligent question-and-answer platform, aiming to improve the efficiency and accuracy of question-and-answering, thereby effectively assisting on-site decision-making. The system is optimized through the following key technologies: data preprocessing, knowledge base management, keyword extraction, vector knowledge base construction, question retrieval, large model processing, performance monitoring and optimization. Through the integration and application of these technologies, the rail transit knowledge base question-and-answer system of the present invention can provide more intelligent and efficient knowledge retrieval services, meet the rail transit industry's demand for professional and accurate engineering information, and effectively support engineering practitioners in on-site quality and safety management. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a flow chart of the method of embodiment 1 of the present invention;

[0024] Figure 2 It is a system structure diagram of embodiment 2 of the present invention. DETAILED DESCRIPTION

[0025] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0026] The method provided by the present invention can be implemented in the following terminal environment, and the terminal may include one or more of the following components: a processor, a storage medium, and a display screen. The storage medium stores at least one instruction, and the instruction is loaded and executed by the processor to implement the method described in the following embodiment.

[0027] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts in the entire terminal, and executes various functions of the terminal and processes data by running or executing instructions, programs, code sets or instruction sets stored in the storage medium, and calling data stored in the storage medium.

[0028] The storage medium may include a random access memory (RAM) or a read-only memory (ROM). The storage medium may be used to store instructions, programs, codes, code sets or instructions.

[0029] The display is used to show the interactive cross-section of each application.

[0030] All subscripts in the formulas of the present invention are only used to distinguish parameters and have no actual meaning.

[0031] In addition, those skilled in the art can understand that the structure of the above terminal does not constitute a limitation on the terminal, and the terminal may include more or fewer components, or combine certain components, or arrange the components differently. For example, the terminal also includes components such as a radio frequency circuit, an input unit, a sensor, an audio circuit, and a power supply, which will not be described in detail here.

[0032] Example 1

[0033] like Figure 1 As shown, the present invention proposes a method for constructing a rail transit knowledge base based on a large language model, comprising:

[0034] Step 1: Collect text materials in the field of rail transit from multiple sources, including PDF, TXT, Word, and MD formats, to ensure the diversity and comprehensiveness of the data.

[0035] Step 2: Use text extraction tools to convert documents in different formats into UTF-8 encoded plain text files for subsequent processing.

[0036] Step 3: Remove irrelevant formatting marks, headers and footers, charts and images from the text to ensure the purity of the text content.

[0037] Specifically, this embodiment designs a new algorithm for removing irrelevant formatting marks, headers and footers, charts and images in text, as shown below:

[0038]

[0039] Among them, score1(P i ) is the i-th paragraph P i The first evaluation score is used to identify paragraphs that contain a large number of special characters, unusual length, or unusual number of lines. These paragraphs may be headers, footers, or other formatting tags. s (P i ) is the i-th paragraph P i α is the first adjustment factor of the first evaluation score, β is the second adjustment factor of the first evaluation score, γ is the third adjustment factor of the first evaluation score, L(P i )) is the i-th paragraph P i The length of the paragraph, N l (P i ) is the i-th paragraph P i number of rows.

[0040]

[0041] Among them, score2(P i ) is the i-th paragraph P i The second evaluation score is used to identify paragraphs with a high proportion of digital characters, low keyword density, or abnormal average line length. These paragraphs may be tables, charts, or other non-text content. δ is the adjustment factor for the second evaluation score, R n (P i ) is the i-th paragraph P i The ratio of digital characters, D k (P i ) is the i-th paragraph P i Keyword density, A l (P i ) is the i-th paragraph P i The average line length.

[0042]

[0043] Among them, score3(P i ) is the i-th paragraph P iThe third evaluation score is used to comprehensively evaluate multiple aspects of the paragraph, so as to effectively identify paragraphs containing a large number of special characters, abnormal length, high proportion of numeric characters or low keyword density. α′ is the first adjustment factor of the third evaluation score, β′ is the second adjustment factor of the third evaluation score, γ′ is the third adjustment factor of the third evaluation score, and F w (P i ) is the i-th paragraph P i The frequency of a word.

[0044]

[0045] Among them, score4(P i ) is the i-th paragraph P i , δ′ is the first adjustment factor of the fourth evaluation score, ∈ is the second adjustment factor of the fourth evaluation score, η is the third adjustment factor of the fourth evaluation score, ζ is the fourth adjustment factor of the fourth evaluation score, R c (P i ) is the i-th paragraph P i is the proportion of uppercase letters in the fourth evaluation score, and ι is the fifth adjustment factor of the fourth evaluation score.

[0046] The lower the score, the more likely the paragraph is to contain irrelevant content (such as formatting marks, headers and footers, charts and images, etc.), while the higher the score, the more likely the paragraph is to be valuable body content.

[0047] The above scores are weighted averaged to obtain the final score, which is compared with the threshold to determine whether to delete the i-th paragraph P i .

[0048] Step 4: Convert all text to lowercase, unify punctuation, and perform word segmentation and part-of-speech tagging.

[0049] Step 5: Use a hash algorithm based on text fingerprints to detect and remove duplicate data to ensure the uniqueness of the data in the knowledge base.

[0050] Step 6: Implement the creation, update, deletion, and query functions of the knowledge base, and build various types of knowledge bases, such as national standards, specifications, guidelines, laws and regulations, etc.

[0051] Step 7: Use the HanLP algorithm to perform word segmentation, word frequency statistics and position weight analysis on the text block, and use the TF-IDF weighting mechanism to extract keywords.

[0052] Step 8: Set the filtering threshold to exclude low-weight words, and sort the remaining keywords by weight to generate a keyword list.

[0053] Step 9: Use the BERT model to vectorize the keyword list and build an inverted index to associate the keyword vector with the text block ID.

[0054] Step 10: Implement a mechanism to synchronously update the vector representation and inverted index when the knowledge base is updated.

[0055] Step 11: Preprocess the user query by removing punctuation, lowercasing, removing stop words, and extracting keywords.

[0056] Step 12: Identify the query intent through a deep neural network (DNN) to determine the type of knowledge base the user wants to retrieve.

[0057] Step 13: Use the BERT model to convert the keywords in the query into vector representations.

[0058] Step 14: Use cosine similarity to calculate the similarity between the query vector and the keyword list vector in the knowledge base, and quickly locate the most relevant text block ID. According to the calculated similarity, retrieve the corresponding text block list and provide it to the user.

[0059] Example 2

[0060] like Figure 2 As shown, the embodiment of the present invention further proposes a rail transit knowledge base construction system based on a large language model, including:

[0061] A text block generation module is used to obtain text data in the field of rail transit, convert the text data in different formats into plain text files through a text extraction tool, remove irrelevant formatting marks, headers and footers, charts and / or images in the plain text files, and perform word segmentation and part-of-speech tagging on the plain text files to generate multiple text blocks;

[0062] Generate a knowledge base module, which is used to perform word segmentation, word frequency statistics and position weight analysis on the text block, extract keywords, set a filtering threshold to exclude keywords with a position weight lower than the filtering threshold, and sort the remaining keywords by position weight, generate a keyword list, vectorize the keyword list, build an inverted index, associate the keyword vector with the text block ID, and generate a knowledge base;

[0063] A knowledge module is provided to pre-process user queries, remove punctuation, lowercase, remove stop words, extract keywords of user queries, convert keywords of user queries into vector representations, calculate the similarity between the vectors of keywords of user queries and the keyword vectors in the knowledge base, locate the most relevant text block ID, and provide the user with a list of text blocks corresponding to the most relevant text block IDs.

[0064] Specifically, the text data in different formats are converted into a plain text file through a text extraction tool, wherein the plain text file is a plain text file encoded in UTF-8.

[0065] Specifically, before performing word segmentation and part-of-speech tagging, the method further includes: converting all text in the plain text file into lowercase.

[0066] Specifically, after the plain text file is segmented and POS tagged, the method further includes: detecting and removing duplicate data of the tagged text file through a hash algorithm based on text fingerprints to ensure the uniqueness of the data in the knowledge base.

[0067] Specifically, after extracting the keywords of the user's query, the process also includes: identifying the user's query intention through a deep neural network and determining the type of knowledge base the user wants to retrieve.

[0068] Example 3

[0069] An embodiment of the present invention further proposes a storage medium storing a plurality of instructions, wherein the instructions are used to implement the method for constructing a rail transit knowledge base based on a large language model.

[0070] Optionally, in this embodiment, the above storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0071] Optionally, in this embodiment, the storage medium is configured to store program code for executing the method steps: Step 1: Collect text materials in the field of rail transit from multiple sources, including PDF, TXT, Word, and MD formats, to ensure the diversity and comprehensiveness of the data.

[0072] Step 2: Use text extraction tools to convert documents in different formats into UTF-8 encoded plain text files for subsequent processing.

[0073] Step 3: Remove irrelevant formatting marks, headers and footers, charts and images from the text to ensure the purity of the text content.

[0074] Step 4: Convert all text to lowercase, unify punctuation, and perform word segmentation and part-of-speech tagging.

[0075] Step 5: Use a hash algorithm based on text fingerprints to detect and remove duplicate data to ensure the uniqueness of the data in the knowledge base.

[0076] Step 6: Implement the creation, update, deletion, and query functions of the knowledge base, and build various types of knowledge bases, such as national standards, specifications, guidelines, laws and regulations, etc.

[0077] Step 7: Use the HanLP algorithm to perform word segmentation, word frequency statistics and position weight analysis on the text block, and use the TF-IDF weighting mechanism to extract keywords.

[0078] Step 8: Set the filtering threshold to exclude low-weight words, and sort the remaining keywords by weight to generate a keyword list.

[0079] Step 9: Use the BERT model to vectorize the keyword list and build an inverted index to associate the keyword vector with the text block ID.

[0080] Step 10: Implement a mechanism to synchronously update the vector representation and inverted index when the knowledge base is updated.

[0081] Step 11: Preprocess the user query by removing punctuation, lowercasing, removing stop words, and extracting keywords.

[0082] Step 12: Identify the query intent through a deep neural network (DNN) to determine the type of knowledge base the user wants to retrieve.

[0083] Step 13: Use the BERT model to convert the keywords in the query into vector representations.

[0084] Step 14: Use cosine similarity to calculate the similarity between the query vector and the keyword list vector in the knowledge base, and quickly locate the most relevant text block ID. According to the calculated similarity, retrieve the corresponding text block list and provide it to the user.

[0085] Example 4

[0086] An embodiment of the present invention also proposes an electronic device, including a processor and a storage medium connected to the processor, wherein the storage medium stores a plurality of instructions, and the instructions can be loaded and executed by the processor so that the processor can execute a method for constructing a rail transit knowledge base based on a large language model.

[0087] Specifically, the electronic device of this embodiment may be a computer terminal, and the computer terminal may include: one or more processors, and a storage medium.

[0088] Among them, the storage medium can be used to store software programs and modules, such as a method for constructing a rail transit knowledge base based on a large language model in an embodiment of the present invention, and the corresponding program instructions / modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the storage medium, that is, realizing the above-mentioned method for constructing a rail transit knowledge base based on a large language model. The storage medium may include a high-speed random storage medium, and may also include a non-volatile storage medium, such as one or more magnetic storage systems, flash memory, or other non-volatile solid-state storage media. In some instances, the storage medium may further include a storage medium remotely arranged relative to the processor, and these remote storage media may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0089] The processor can call the information and application programs stored in the storage medium through the transmission system to execute the method steps: Step 1: Collect text materials in the field of rail transit from multiple sources, including PDF, TXT, Word, MD and other formats to ensure the diversity and comprehensiveness of the data.

[0090] Step 2: Use text extraction tools to convert documents in different formats into UTF-8 encoded plain text files for subsequent processing.

[0091] Step 3: Remove irrelevant formatting marks, headers and footers, charts and images from the text to ensure the purity of the text content.

[0092] Step 4: Convert all text to lowercase, unify punctuation, and perform word segmentation and part-of-speech tagging.

[0093] Step 5: Use a hash algorithm based on text fingerprints to detect and remove duplicate data to ensure the uniqueness of the data in the knowledge base.

[0094] Step 6: Implement the creation, update, deletion, and query functions of the knowledge base, and build various types of knowledge bases, such as national standards, specifications, guidelines, laws and regulations, etc.

[0095] Step 7: Use the HanLP algorithm to perform word segmentation, word frequency statistics and position weight analysis on the text block, and use the TF-IDF weighting mechanism to extract keywords.

[0096] Step 8: Set the filtering threshold to exclude low-weight words, and sort the remaining keywords by weight to generate a keyword list.

[0097] Step 9: Use the BERT model to vectorize the keyword list and build an inverted index to associate the keyword vector with the text block ID.

[0098] Step 10: Implement a mechanism to synchronously update the vector representation and inverted index when the knowledge base is updated.

[0099] Step 11: Preprocess the user query by removing punctuation, lowercasing, removing stop words, and extracting keywords.

[0100] Step 12: Identify the query intent through a deep neural network (DNN) to determine the type of knowledge base the user wants to retrieve.

[0101] Step 13: Use the BERT model to convert the keywords in the query into vector representations.

[0102] Step 14: Use cosine similarity to calculate the similarity between the query vector and the keyword list vector in the knowledge base, and quickly locate the most relevant text block ID. According to the calculated similarity, retrieve the corresponding text block list and provide it to the user.

[0103] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0104] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0105] In the several embodiments provided by the present invention, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the system embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0106] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0107] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0108] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only storage medium (ROM, Read-Only Memory), random access storage medium (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.

[0109] Obviously, the above embodiments are merely examples for the purpose of clear explanation, and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the invention.

Claims

1. A method for constructing a rail transit knowledge base based on a large language model, characterized in that: include: Acquire text data in the field of rail transit, convert the text data in different formats into plain text files through a text extraction tool, remove irrelevant formatting marks, headers and footers, charts and / or images in the plain text files, and perform word segmentation and part-of-speech tagging on the plain text files to generate multiple text blocks; Removing extraneous formatting marks, headers and footers, charts and / or images from the plain text file includes: Calculate the first evaluation score, the second evaluation score, the third evaluation score and the fourth evaluation score of each paragraph, and perform weighted average to obtain a final score, and compare it with the threshold to determine whether to delete the current paragraph; Wherein, calculating the first evaluation score includes: Among them, score1(P i ) is the i-th paragraph P i The first evaluation score, R s (P i ) is the i-th paragraph P i α is the first adjustment factor of the first evaluation score, β is the second adjustment factor of the first evaluation score, γ is the third adjustment factor of the first evaluation score, L(P i )) is the uth paragraph P i The length of the paragraph, N l (P i ) is the i-th paragraph P i number of rows; Calculating the second assessment score includes: Among them, score2(P i ) is the i-th paragraph P i The second evaluation score of , δ is the adjustment factor of the second evaluation score, R n (P i ) is the i-th paragraph P i The ratio of digital characters, D k (P i ) is the i-th paragraph P i Keyword density, A l (P i ) is the i-th paragraph P i Average line length; Calculation of the third assessment score includes: Among them, score3(P i ) is the i-th paragraph P i The third evaluation score, α′ is the first adjustment factor of the third evaluation score, β′ is the second adjustment factor of the third evaluation score, γ′ is the third adjustment factor of the third evaluation score, F w (P i ) is the i-th paragraph P i The frequency of a word in Calculation of the fourth assessment score includes: Among them, score4(P i ) is the i-th paragraph P i , δ′ is the first adjustment factor of the fourth evaluation score, ∈ is the second adjustment factor of the fourth evaluation score, η is the third adjustment factor of the fourth evaluation score, ζ is the fourth adjustment factor of the fourth evaluation score, R c (P i ) is the i-th paragraph P i is the proportion of capital letters, ι is the fifth adjustment factor of the fourth assessment score; Perform word segmentation, word frequency statistics and position weight analysis on the text block, extract keywords, set a filtering threshold to exclude keywords with a position weight lower than the filtering threshold, sort the remaining keywords by position weight, generate a keyword list, vectorize the keyword list, build an inverted index, associate the keyword vector with the text block ID, and generate a knowledge base; Preprocess the user query, remove punctuation, lowercase, remove stop words, and extract the keywords of the user query. Convert the keywords of the user query into vector representation, calculate the similarity between the vector of the keyword of the user query and the keyword vector in the knowledge base, locate the most relevant text block ID, and provide the text block list corresponding to the most relevant text block ID to the user.

2. A rail transit knowledge base construction method based on a large language model as claimed in claim 1, characterized in that: The text data in different formats are converted into a plain text file by a text extraction tool, wherein the plain text file is a plain text file encoded in UTF-8.

3. A method for constructing a rail transit knowledge base based on a large language model as claimed in claim 1, characterized in that: Before performing word segmentation and part-of-speech tagging, the method further includes: converting all text in the plain text file into lowercase.

4. A method for constructing a rail transit knowledge base based on a large language model as claimed in claim 1, characterized in that: After the plain text file is segmented and POS tagged, the method further includes: detecting and removing duplicate data of the tagged text file through a hash algorithm based on text fingerprint, so as to ensure the uniqueness of the data in the knowledge base.

5. The method for constructing a rail transit knowledge base based on a large language model as claimed in claim 1, characterized in that: After extracting the keywords of the user's query, it also includes: identifying the intention of the user's query through a deep neural network and determining the type of knowledge base the user wants to retrieve.

6. A rail transit knowledge base construction system based on a large language model, characterized in that: include: A text block generation module is used to obtain text data in the field of rail transit, convert the text data in different formats into plain text files through a text extraction tool, remove irrelevant formatting marks, headers and footers, charts and / or images in the plain text files, and perform word segmentation and part-of-speech tagging on the plain text files to generate multiple text blocks; Removing extraneous formatting marks, headers and footers, charts and / or images from the plain text file includes: Calculate the first evaluation score, the second evaluation score, the third evaluation score and the fourth evaluation score of each paragraph, and perform weighted average to obtain a final score, and compare it with the threshold to determine whether to delete the current paragraph; Wherein, calculating the first evaluation score includes: Among them, score1(P i ) is the i-th paragraph P i The first evaluation score, R s (P i ) is the i-th paragraph P i α is the first adjustment factor of the first evaluation score, β is the second adjustment factor of the first evaluation score, γ is the third adjustment factor of the first evaluation score, L(P i )) is the i-th paragraph P i The length of the paragraph, N l (P i ) is the i-th paragraph P i number of rows; Calculating the second assessment score includes: Among them, score2(P i ) is the i-th paragraph P i The second evaluation score of , δ is the adjustment factor of the second evaluation score, R n (P i ) is the i-th paragraph P i The ratio of digital characters, D k (P i ) is the i-th paragraph P i Keyword density, A l (P i ) is the i-th paragraph P i Average line length; Calculation of the third assessment score includes: Among them, score3(P i ) is the i-th paragraph P i The third evaluation score, α′ is the first adjustment factor of the third evaluation score, β′ is the second adjustment factor of the third evaluation score, γ′ is the third adjustment factor of the third evaluation score, F w (P i ) is the i-th paragraph P i The frequency of a word in Calculation of the fourth assessment score includes: Among them, score4(P i ) is the i-th paragraph P i , δ′ is the first adjustment factor of the fourth evaluation score, ∈ is the second adjustment factor of the fourth evaluation score, η is the third adjustment factor of the fourth evaluation score, ζ is the fourth adjustment factor of the fourth evaluation score, P c (P i ) is the i-th paragraph P i is the proportion of capital letters, ι is the fifth adjustment factor of the fourth assessment score; Generate a knowledge base module, which is used to perform word segmentation, word frequency statistics and position weight analysis on the text block, extract keywords, set a filtering threshold to exclude keywords with a position weight lower than the filtering threshold, and sort the remaining keywords by position weight, generate a keyword list, vectorize the keyword list, build an inverted index, associate the keyword vector with the text block ID, and generate a knowledge base; A knowledge module is provided to pre-process user queries, remove punctuation, lowercase, remove stop words, extract keywords of user queries, convert keywords of user queries into vector representations, calculate the similarity between the vectors of keywords of user queries and the keyword vectors in the knowledge base, locate the most relevant text block ID, and provide the user with a list of text blocks corresponding to the most relevant text block IDs.

7. A rail transit knowledge base construction system based on a large language model as claimed in claim 6, characterized in that: The text data in different formats are converted into a plain text file by a text extraction tool, wherein the plain text file is a plain text file encoded in UTF-8.

8. A rail transit knowledge base construction system based on a large language model as claimed in claim 6, characterized in that: Before performing word segmentation and part-of-speech tagging, the method further includes: converting all text in the plain text file into lowercase.

9. A rail transit knowledge base construction system based on a large language model as claimed in claim 6, characterized in that: After the plain text file is segmented and POS tagged, the method further includes: detecting and removing duplicate data of the tagged text file through a hash algorithm based on text fingerprint, so as to ensure the uniqueness of the data in the knowledge base.

10. A rail transit knowledge base construction system based on a large language model as claimed in claim 6, characterized in that: After extracting the keywords of the user's query, it also includes: identifying the intention of the user's query through a deep neural network and determining the type of knowledge base the user wants to retrieve.

Citation Information

Patent Citations

  • Natural language question answering method and device based on large language model, medium and equipment

    CN117473053A