Track traffic industry document content extraction method and device, equipment and medium

By using pre-tuned instructions to prompt the propt and content extraction method in the rail transit industry, the problems of low accuracy, poor flexibility and low efficiency in document processing are solved, and document content extraction with high accuracy, flexibility and efficiency are achieved.

CN120068844APending Publication Date: 2025-05-30BEIJING UNISOUND INFORMATION TECH CO LTD +7
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510126944.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has problems such as low accuracy, poor flexibility, low efficiency and high update and maintenance costs in document processing in the rail transit industry.

Method used

Provide a method for extracting document content in the rail transit industry, using pre-tuned instructions to prompt the propt and pre-trained content extraction large models, to automatically extract key information from the to-process documents, including text names, table contents, quantity information, differential contents, compliance terms, etc.

Benefits of technology

It realizes high accuracy, flexibility and efficiency of document content extraction, reduces the dependence and cost of manual processing, improves the consistency and accuracy of document processing, and adapts to various document processing scenarios of different types and formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068844A_ABST
    Figure CN120068844A_ABST
Patent Text Reader

Abstract

The invention discloses a rail transit industry document content extraction method and device, equipment and a medium. Various kinds of key information can be accurately extracted from the to-be-processed document through the pre-optimized instruction prompt prompt in combination with the content extraction large model. The key information can be comprehensively extracted from the to-be-processed document in the rail transit industry through different types of prompts, a reliable data basis is provided for subsequent data analysis, decision making and other work, and the accuracy of judgment and decision making based on the document information is greatly improved. The whole extraction process is highly automatic, a large model is extracted by means of pre-trained content, and a large number of rail transit industry documents can be rapidly processed. The dependence on manual document processing is reduced, and the labor cost is reduced. Meanwhile, errors caused by factors such as fatigue and negligence possibly occurring in the manual processing process are avoided, and the document processing consistency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and data processing, and particularly to a method, apparatus, device, and medium for extracting document content in the rail transit industry. Background Art

[0002] With the rapid development of the rail transit industry, a large number of document materials have emerged continuously. These documents cover various aspects such as design drawings, operation reports, and safety specifications, and play a key role in the planning, construction, operation, and management of rail transit. As the core means of managing these materials, document processing technology is of great importance.

[0003] Currently, the technologies widely used in rail transit document processing mainly include natural language processing (NLP), rule engines, and manual review processes. NLP technology plays an important role in processing text documents by virtue of its capabilities of parsing text content, semantic understanding, information extraction, and summary generation; the rule engine processes specific types of documents according to predefined rules, such as matching specific format information through regular expressions; the manual review process, as the last line of defense to ensure accuracy, reviews the results of automated processing.

[0004] However, these existing technologies have exposed many problems in practical applications. In terms of accuracy, when faced with complex or non-standard format documents, NLP technology often fails to correctly identify and parse them; due to relying on fixed rules, the rule engine has serious deficiencies in flexibility when encountering new document types, requires reconfiguration, and lacks the adaptive ability to handle changes; when processing a large number of documents, traditional methods are inefficient and it is even difficult to meet the requirements in scenarios that require real-time response; moreover, with the rapid development of the rail transit industry, the number of new document types continues to increase, making the update and maintenance of existing processing systems difficult and costly. These problems have seriously hindered the process of efficiently processing documents in the rail transit industry. Therefore, it is urgent to develop a more accurate, flexible, efficient, and cost-controlled document processing technology. Summary of the Invention

[0005] This application provides a method, apparatus, device, and medium for extracting document content in the rail transit industry, which is used to solve the problems of low efficiency, poor flexibility, and poor accuracy in the current process of extracting document content in the rail transit industry, resulting in the need to consume manpower for review.

[0006] In a first aspect, this application provides a method for extracting document content in the rail transit industry, and the method includes:

[0007] Obtain any document to be processed in the rail transit industry;

[0008] Based on the to-be-processed document and each pre-tuned instruction prompt, input data is obtained; wherein, each of the prompts includes the following types: text name extraction instruction, table content extraction instruction, text quantity extraction instruction, content extraction instruction for the content with differences in the text and without relevant analysis, clause extraction instruction in compliance, integrity judgment instruction based on, and instruction for judging the relevance between fields and values; the text name extraction instruction is used to extract name information in the to-be-processed document; the table content extraction instruction is used to extract table content in the to-be-processed document; the text quantity instruction is used to extract the vocabulary and corresponding content representing quantity in the to-be-processed document; the content extraction instruction for the content with differences in the text and without relevant analysis is used to extract the content with differences but without detailed analysis in the to-be-processed document; the clause extraction instruction in compliance is used to extract the clauses related to compliance and their specific content in the to-be-processed document; the integrity judgment instruction based on is used to judge whether the basis given in the to-be-processed document is complete; the instruction for judging the relevance between fields and values is used to judge the relevance between fields and values in the to-be-processed document;

[0009] For each of the input data, through a pre-trained large content extraction model for rail transit industry documents, based on this input data, a special document extraction result corresponding to this input data is obtained; wherein, the document extraction result includes a list of document extraction content key-value pairs and document extraction content.

[0010] In a second aspect, the present application also provides a rail transit industry document content extraction device, and the device includes:

[0011] An acquisition unit, configured to acquire any to-be-processed document in the rail transit industry;

[0012] A processing unit, configured to obtain input data based on the document to be processed and each pre-tuned instruction prompt; wherein, each of the prompts includes the following types: text name extraction instruction, table content extraction instruction, text quantity extraction instruction, content extraction instruction for differences in text without relevant analysis, clause extraction instruction in compliance, basis integrity judgment instruction, and instruction for judging the relevance between fields and values; the text name extraction instruction is used to extract name information in the document to be processed; the table content extraction instruction is used to extract table content in the document to be processed; the text quantity instruction is used to extract the vocabulary and corresponding content representing quantity in the document to be processed; the content extraction instruction for differences in text without relevant analysis is used to extract the content with differences but without detailed analysis in the document to be processed; the clause extraction instruction in compliance is used to extract the clauses related to compliance and their specific content in the document to be processed; the basis integrity judgment instruction is used to judge whether the basis given in the document to be processed is complete; the instruction for judging the relevance between fields and values is used to judge the relevance between fields and values in the document to be processed;

[0013] A content extraction unit, configured to, for each of the input data, obtain a special document extraction result corresponding to the input data through a pre-trained content extraction large model for rail transit industry documents, based on the input data; wherein, the document extraction result includes a list of document extraction content key-value pairs and document extraction content.

[0014] In a third aspect, the present application provides a computer device, which includes a processor, and the processor is configured to implement the steps of the rail transit industry document content extraction method as described above when executing a computer program stored in a memory.

[0015] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program, and the computer program is configured to implement the steps of the rail transit industry document content extraction method as described above when executed by a processor.

[0016] The beneficial effects of the present application are as follows:

[0017] 1. By using each pre-tuned instruction prompt in combination with a content extraction large model, various types of key information can be accurately extracted from the document to be processed. Among them, each instruction prompt includes the following types: text name extraction instruction, table content extraction instruction, text quantity extraction instruction, content extraction instruction for the parts with differences in the text without relevant analysis, clause extraction instruction in compliance, basis integrity judgment instruction, and instruction for judging whether a field and its value are relevant. Through different types of prompts, key information can be comprehensively extracted from the documents to be processed in the rail transit industry, providing a reliable data basis for subsequent data analysis, decision-making, etc., and greatly improving the accuracy of judgments and decisions based on document information.

[0018] 2. The entire extraction process is highly automated. With the help of a pre-trained content extraction large model, a large number of rail transit industry documents can be quickly processed. Compared with the traditional manual processing method, it greatly saves labor and time costs. It reduces the dependence on manual document processing and lowers labor costs. At the same time, it avoids errors caused by factors such as fatigue and negligence that may occur during manual processing, improving the consistency and accuracy of document processing. When faced with a large number of operation reports, maintenance records and other documents, this method can complete the information extraction work in a short time, improving the speed of document processing, enabling relevant departments to obtain the required information in a timely manner and make a quick response, thus enhancing the operation and management efficiency of the entire rail transit system.

[0019] 3. Different types of prompts cover a variety of common document processing requirements. This diverse instruction design enables this method to adapt to various different types and formats of document processing scenarios in the rail transit industry. Whether it is a newly emerged document type or special document content, accurate information extraction can be achieved by adjusting and optimizing the corresponding prompts. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a schematic diagram of the process for extracting the content of rail transit industry documents provided by the embodiment of the present application;

[0022] Figure 2 It is a schematic diagram of the device structure for extracting the content of rail transit industry documents provided by the embodiment of the present application;

[0023] Figure 3 It is a schematic structural diagram of a computer device provided by an alternative embodiment of the present application. Specific embodiments

[0024] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0025] In order to extract the content of rail transit industry documents more accurately, flexibly, efficiently and with controllable costs, the present application provides a method, device, equipment and medium for extracting the content of rail transit industry documents.

[0026] Embodiment 1:

[0027] The present application provides a method for extracting the content of rail transit industry documents. Figure 1 It is a schematic diagram of the process of extracting the content of rail transit industry documents provided by an embodiment of the present application. The process includes:

[0028] S101: Obtain any document to be processed in the rail transit industry.

[0029] The method for extracting the content of rail transit industry documents provided by the present application is applied to a computer device. The computer device can be an intelligent terminal, such as a computer, a robot, etc., or a server, such as an application server, a business server, etc.

[0030] In the actual rail transit operation scenario, the sources of documents to be processed are very extensive. For example, the operation department generates detailed operation reports every day, which record information such as the departure time, arrival time, on-time / delayed situation, passenger flow data, and passenger flow distribution in different time periods of the trains on that day. These data are crucial for analyzing operation efficiency and service quality. The maintenance department will submit maintenance record documents, which detail information such as the time when the equipment fails, the failure phenomenon, the maintenance measures taken by the maintenance personnel, and the model and quantity of the replaced parts. These documents are of great significance for the maintenance and management of the equipment. In addition, the safety management department will issue safety specification documents, covering various safety operation procedures, emergency plans, etc. For any of the above-mentioned documents to be processed, the method for extracting the content of rail transit industry documents provided by the present application can be used to accurately, efficiently and flexibly extract the content of the document to be processed.

[0031] S102: Obtain input data based on the document to be processed and the pre-tuned instruction prompts (prompts). Among them, the prompts include the following types: text name extraction instruction, table content extraction instruction, text quantity extraction instruction, content extraction instruction for the differences in the text without relevant analysis, clause extraction instruction in compliance, basis integrity judgment instruction, and instruction for judging the relevance between fields and values. The text name extraction instruction is used to extract name information in the document to be processed. The table content extraction instruction is used to extract the table content in the document to be processed. The text quantity instruction is used to extract the vocabulary and corresponding content representing quantity in the document to be processed. The content extraction instruction for the differences in the text without relevant analysis is used to extract the content with differences but without detailed analysis in the document to be processed. The clause extraction instruction in compliance is used to extract the clauses related to compliance and their specific content in the document to be processed. The basis integrity judgment instruction is used to judge whether the basis given in the document to be processed is complete. The instruction for judging the relevance between fields and values is used to judge the relevance between fields and values in the document to be processed.

[0032] After obtaining the document to be processed based on the above embodiments, input data can be obtained based on the document to be processed and the pre-tuned instruction prompts (prompts). For example, for the pre-tuned instruction prompts (prompts), input data can be obtained by concatenating the document to be processed and the prompt.

[0033] In a possible implementation, any prompt includes: user role information, extraction requirements based on the user role, and reference input and output examples. Among them, each prompt includes multiple types, and each type has its specific function and use. The following is an explanation of each prompt:

[0034] 1. Text name extraction instruction

[0035] The text name extraction instruction is used to extract name information in the document to be processed, such as train name, station name, equipment name, etc. Each prompt includes user role information, extraction requirements based on the user role, and reference input and output examples. For example, the prompt for the text name extraction instruction may be as follows:

[0036]

User role information

[0037]

Extraction requirements

[0038] a) Please output the extraction result in a markdown table, and the columns of the table are respectively:

[0039] Target field name, name of the affiliated line segment or line plan, value of the target field, original text, name of the cited reference;

[0040] b) If the name of the line segment or line plan corresponding to the target field cannot be determined from the input text, the value of the name of the affiliated line segment or line plan is "undetermined";

[0041] c) If the target field content is not present in the original text, the value of the target field is "no corresponding value";

[0042] d) The content in the original text column is required to follow the extraction of the corresponding original text fragment without any modification;

[0043] e) The name of the cited reference refers to official issued regulatory documents such as construction plans and project approvals. If the current value of the target field is cited from a document, return the document name; otherwise, return "none";

[0044]

Reference input and output examples

[0045] Input: The 29th line project runs from Hongjian Road Station to Baoyang Road Station, with a line length of 44.5 kilometers and 2 main substations, namely Lingzhao Xincun Main Substation and Shanghai International Studies University Main Substation.

[0046] Output: Field|Value|Original Text|Reference File Name Mentioned in the Original Text|

[0047] --------|---|Project Name of Line 29|Line 29 Project|The 29th line project runs from Hongjian Road Station to Baoyang Road Station|none

[0048] |Starting Station of Line 29|Hongjian Road Station|The 29th line project runs from Hongjian Road Station to Baoyang Road Station|none|Terminating Station of Line 29|Baoyang Road Station|The 29th line project runs from Hongjian Road Station to Baoyang Road Station|none

[0049] |Name of Main Substation 1 of Line 29|Lingzhao Xincun Main Substation|The 29th line project runs from Hongjian Road Station to Baoyang Road Station, with a line length of 44.5 kilometers and 2 main substations, namely Lingzhao Xincun Main Substation and Shanghai International Studies University Main Substation|none

[0050] |Name of Main Substation 2 of Line 29|Shanghai International Studies University Main Substation|The 29th line project runs from Hongjian Road Station to Baoyang Road Station, with a line length of 44.5 kilometers and 2 main substations, namely Lingzhao Xincun Main Substation and Shanghai International Studies University|none

[0051] 2. Instructions for extracting table content

[0052] The table content extraction instruction is used to extract the table content from the document to be processed. The design idea of its prompt is similar to that of the text name extraction instruction, including user role information, extraction requirements, and reference input and output examples. For example:

[0053]

User role information

[0054]

Extraction requirements

[0055]

Reference input and output examples

[0056] Input: (Show a screenshot of a table containing train operation time and passenger flow data)

[0057] Output:

[0058] Train Number Departure Time Arrival Time Number of Boarding Passengers Number of Alighting Passengers Train No. 001 on Line 1 06:00 06:30 100 50 Train No. 002 on Line 1 06:10 06:40 120 60

[0059] 3. Text quantity extraction instruction

[0060] The text quantity extraction instruction is used to extract the vocabulary and corresponding content representing quantity in the document to be processed, such as the running mileage of trains, passenger flow data, the number of equipment repairs, etc. Its prompt example is as follows:

[0061]

User role information

[0062]

Extraction requirements

[0063]

Reference input and output examples

[0064] Input: "Today, the total running mileage of Line 1 trains is 1000 kilometers, the passenger flow reaches 50,000 person-times, and the number of equipment repairs is 3 times"

[0065] Output:

[0066]

[0067]

[0068] 4. Content extraction instruction for differences in text without relevant analysis

[0069] This instruction is used to extract the content with differences but without detailed analysis in the document to be processed. For example, in the operation report, it may be mentioned that there are significant differences in the passenger flow during different time periods, but the reasons for the differences are not analyzed. Its prompt example is as follows:

[0070]

User Role Information

[0071]

Extraction Requirements

[0072]

Reference Input and Output Examples

[0073] Input: "The passenger flow on Monday this week was 80,000 person-times, while the passenger flow on Tuesday this week was only 50,000 person-times. The difference is obvious, but the reason is not explained."

[0074] Output:

[0075]

[0076] 5. Clause Extraction Instructions in Compliance

[0077] The clause extraction instructions in compliance are used to extract the compliance-related clauses and their specific contents in the document to be processed. For example, in a safety specification document, there may be various compliance clauses related to safety operations, emergency plans, etc. The prompt example is as follows:

[0078]

User Role Information

[0079]

Extraction Requirements

[0080]

Reference Input and Output Examples

[0081] Input: "According to the 'Rail Transit Safety Management Regulations', when a fire accident occurs, the emergency plan should be immediately activated to evacuate passengers to a safe area."

[0082] Output:

[0083]

[0084] 6. Instruction for Judging Completeness of Basis

[0085] The instruction for judging completeness of basis is used to judge whether the basis given in the document to be processed is complete. For example, in a document about equipment maintenance plan, it is necessary to judge whether sufficient equipment failure information, basis for maintenance measures, etc. are provided. The prompt example is as follows:

[0086]

User Role Information

[0087]

Extraction Requirements

[0088]

Reference Input and Output Examples

[0089] Input: (Show part of the content of a device maintenance plan document)

[0090] Output:

[0091]

[0092] 7. Instruction for judging whether fields and values are relevant

[0093] This instruction is used to judge the relevance between fields and values in the document to be processed. For example, in a train operation data report, it is necessary to judge whether the departure time and arrival time of the train conform to the normal operation logic. Its prompt example is as follows:

[0094]

User Role Information

[0095]

Extraction Requirement

[0096]

Reference Input and Output Examples

[0097] Input: (Show part of the content of a train operation data report)

[0098] Output:

[0099]

[0100] In a possible implementation manner, after obtaining any document to be processed in the rail transit industry, before obtaining input data based on the document to be processed and each pre-optimized instruction prompt, the method further includes:

[0101] If it is determined that the format of the document to be processed does not conform to the preset standard document format, convert the document to be processed into any standard document format.

[0102] After obtaining the document to be processed, its format can be judged to determine whether it conforms to the preset standard document format. Among them, the preset standard document format can be determined according to the actual processing requirements and system compatibility. Common standard document formats include plain text format (.txt), XML format, etc. Plain text format is easy to process and analyze, and is suitable for most natural language processing tasks; XML format has good structural characteristics and is convenient for organizing and extracting data in the document. If it is found that the format of the document to be processed does not conform to the standard, format conversion is required to convert the document to be processed into a file in the standard document format, so as to facilitate subsequent processing.

[0103] In a possible implementation, before obtaining input data, the document to be processed may be structured to convert disordered text content into information with a clear structure and hierarchy, so as to better understand the content of the document and facilitate subsequent content extraction.

[0104] In this application, any prompt used in content extraction is obtained after multiple iterations of tuning. For example, any prompt is obtained after multiple iterations of tuning, and for any iteration:

[0105] Through the content extraction model, based on the input data carrying the example document sample and the prompt of the current iteration, the expected special extraction result corresponding to the input data is obtained;

[0106] According to the expected special extraction result and the standard special extraction result corresponding to the example document sample, the prompt of the current iteration is tuned to obtain a tuned prompt.

[0107] In each iteration, the content extraction model is first used to obtain the expected special extraction results corresponding to the input data based on the input data carrying the sample document sample and the prompt of the current iteration. For example, a historical operation report is selected as the sample document sample, which contains information such as the operation status of the train and passenger flow, and the standard special extraction results of the sample document sample are manually annotated. The text name class extraction instruction prompt of the current iteration is combined with the sample document sample into input data, which is input into the content extraction model to obtain the expected special extraction results. Then, according to the expected special extraction results and the standard special extraction results corresponding to the sample document sample, the prompt of the current iteration is tuned. Among them, the standard special extraction results can be generated by manual annotation by professionals, with high accuracy and authority. By comparing the expected special extraction results with the standard special extraction results, the differences are found and the reasons are analyzed. If you find that there are omissions or errors in the expected special extraction results, you can make targeted adjustments to the prompt, such as modifying the extraction requirements to make them clearer and more specific; improving the reference input and output examples, and adding more sample data to improve the generalization ability of the prompt. After multiple rounds of iterative tuning, the prompt can meet the preset performance indicators, such as accuracy and recall rate.

[0108] S103: For each of the input data, a pre-trained content extraction model supporting rail transit industry documents is used to obtain, based on the input data, a special document extraction result corresponding to the input data; wherein the document extraction result includes a document extraction content key-value pair list and document extraction content.

[0109] For each input data, through a pre-trained large content extraction model for rail transit industry documents, based on this input data, obtain the special document extraction result corresponding to this input data. Among them, this content extraction large model can adopt a pre-trained model based on a deep learning architecture, such as BERT, GPT, etc., which already has basic natural language understanding ability, and then perform fine-tuning training on a large number of document data in the rail transit industry to enable it to better adapt to the content extraction task of rail transit industry documents.

[0110] Taking the extraction of text name category as an example, after the input data is processed by the large content extraction model, the special document extraction results output may be as follows:

[0111] {

[0112] |Field|Value|Original Text|Reference File Name Mentioned in the Original Text

[0113] |Project Name of Line 29|Line 29 Project|The Line 29 project runs from Hongjian Road Station to Baoyang Road Station|None

[0114] |Starting Station of Line 29|Hongjian Road Station|The Line 29 project runs from Hongjian Road Station to Baoyang Road Station|None |Terminating Station of Line 29|Baoyang Road Station|The Line 29 project runs from Hongjian Road Station to Baoyang Road Station|None

[0115] |Name of the First Main Substation of Line 29|Lingzhao Xincun Main Substation|The Line 29 project runs from Hongjian Road Station to Baoyang Road Station, with a line length of 44.5 kilometers and 2 main substations, namely Lingzhao Xincun Main Substation and Shanghai International Studies University Main Substation|None

[0116] |Name of the Second Main Substation of Line 29|Shanghai International Studies University Main Substation|The Line 29 project runs from Hongjian Road Station to Baoyang Road Station, with a line length of 44.5 kilometers and 2 main substations, namely Lingzhao Xincun Main Substation and Shanghai International Studies University|None

[0117] }

[0118] In a possible implementation manner, the method further includes:

[0119] Generate and output a structured report based on the special extraction results respectively corresponding to each of the input data.

[0120] After obtaining the special extraction results corresponding to each input data, these special extraction results can be integrated to generate a structured report, so as to present the complex and diverse extraction results in a clear and intuitive manner, facilitating relevant personnel to view and analyze. Among them, the structure of the report can be designed according to actual needs and generally can be divided into multiple parts. First is the basic information part of the report, including information such as the time when the report is generated, the name of the document to be processed corresponding to the report, and the operating personnel who processed it. For example, the report generation time is "January 24, 2025, 15:00", the name of the document to be processed is "Operating Report.pdf on January 24, 2025", and the operating personnel is "Zhang San". Next is the summary part of the extraction results. In this part, various types of extraction results will be classified and summarized in the form of a table. For example, for the extraction results of text name types, the extracted train names, station names, equipment names, etc. will be listed; for the extraction results of table content, the extracted table data will be presented in a standardized table form in the report; for the extraction results of text quantity types, various quantity information and their units will be summarized and presented; for the extraction results of content with differences and no relevant analysis given, the difference descriptions and specific data will be clearly listed; for the extraction results of compliance clauses, the extracted compliance clauses and their categories will be listed; for the judgment results based on integrity, the judgment results and missing information will be shown; for the results of judging whether the fields and values are relevant, the correlation judgment results and reasons of the fields will be presented.

[0121] In addition to the result summary, the report can also include an analysis and suggestion part. Analyze based on the extraction results, identify potential problems or points that need attention, and give corresponding suggestions. For example, if it is found that the maintenance times of certain equipment are relatively high, it may be recommended to strengthen the maintenance and management of these equipment; if it is found that the passenger flow during a certain period is abnormal, it may be recommended to further analyze the reasons in order to optimize the operation strategy.

[0122] Finally, the report can also set an appendix part for storing some original data, intermediate results during the processing, etc., for subsequent reference and traceability.

[0123] The beneficial effects of this application are as follows:

[0124] 1. Through each pre-optimized instruction prompt and combined with a content extraction large model, various key information can be accurately extracted from the document to be processed. Among them, each instruction prompt includes the following types: text name extraction instruction, table content extraction instruction, text quantity extraction instruction, content extraction instruction for the text with differences and without relevant analysis, clause extraction instruction in compliance, integrity judgment instruction based on, and instruction for judging whether the field and value are relevant. Through different types of prompts, key information can be comprehensively extracted from the documents to be processed in the rail transit industry, providing a reliable data basis for subsequent data analysis, decision-making, etc., and greatly improving the accuracy of judgments and decisions based on document information.

[0125] 2. The entire extraction process is highly automated. With the help of a pre-trained content extraction large model, a large number of documents in the rail transit industry can be quickly processed. Compared with the traditional manual processing method, it greatly saves labor and time costs. It reduces the dependence on manual document processing and lowers the labor cost. At the same time, it avoids errors caused by factors such as fatigue and negligence that may occur during the manual processing process, improving the consistency and accuracy of document processing. When facing a large number of documents such as operation reports and maintenance records, this method can complete the information extraction work in a short time, improving the speed of document processing, enabling relevant departments to obtain the required information in a timely manner and make a quick response, thus enhancing the operation and management efficiency of the entire rail transit system.

[0126] 3. Different types of prompts cover a variety of common document processing requirements. This diverse instruction design enables this method to adapt to various document processing scenarios of different types and formats in the rail transit industry. Whether it is a newly emerging document type or special document content, accurate information extraction can be achieved by adjusting and optimizing the corresponding prompts.

[0127] Example 2:

[0128] Based on the same inventive concept, the present application also provides a device for extracting the content of rail transit industry documents. Figure 2 FIG. is a schematic structural diagram of a device for extracting the content of rail transit industry documents provided in an embodiment of the present application. The device includes:

[0129] An acquisition unit 21, configured to acquire any document to be processed in the rail transit industry;

[0130] A processing unit 22, configured to obtain input data based on the document to be processed and each pre-tuned instruction prompt; wherein, each of the prompts includes the following types: text name extraction instruction, table content extraction instruction, text quantity extraction instruction, content extraction instruction for the content with differences in the text and without relevant analysis, clause extraction instruction in compliance, basis integrity judgment instruction, and instruction for judging the relevance between fields and values; the text name extraction instruction is used to extract name information in the document to be processed; the table content extraction instruction is used to extract table content in the document to be processed; the text quantity instruction is used to extract the vocabulary and corresponding content representing quantity in the document to be processed; the content extraction instruction for the content with differences in the text and without relevant analysis is used to extract the content with differences but without detailed analysis in the document to be processed; the clause extraction instruction in compliance is used to extract clauses related to compliance and their specific content in the document to be processed; the basis integrity judgment instruction is used to judge whether the basis given in the document to be processed is complete; the instruction for judging the relevance between fields and values is used to judge the relevance between fields and values in the document to be processed.

[0131] A content extraction unit 23, configured to, for each of the input data, obtain a corresponding special document extraction result based on the input data through a pre-trained content extraction large model for rail transit industry documents; wherein, the document extraction result includes a list of document extraction content key-value pairs and document extraction content.

[0132] In some possible embodiments, any prompt includes: user role information, extraction requirements based on the user role, and reference input and output examples.

[0133] In some possible embodiments, the device further includes: a prompt tuning unit.

[0134] The prompt tuning unit is configured to obtain a tuned prompt; wherein, any prompt is obtained after multiple iterative tuning, and for any iteration:

[0135] Through the content extraction large model, based on the input data carrying the example document sample and the prompt of the current iteration, obtain the expected special extraction result corresponding to the input data.

[0136] According to the expected special extraction result and the standard special extraction result corresponding to the example document sample, tune the prompt of the current iteration to obtain a tuned prompt.

[0137] In some possible embodiments, the obtaining unit 21 is further configured to perform a structuring process on the document to be processed.

[0138] In some possible embodiments, the obtaining unit 21 is further configured to, if it is determined that the format of the document to be processed does not conform to a preset standard document format, convert the document to be processed into any standard document format.

[0139] In some possible embodiments, the apparatus further includes: an output unit;

[0140] The output unit is configured to generate and output a structured report based on the special extraction results respectively corresponding to the input data.

[0141] It should be noted that the principle of the rail transit industry document content extraction apparatus for solving problems is similar to that of the above-mentioned rail transit industry document content extraction method. For details, reference can be made to the above method embodiments, and no specific elaboration will be made here.

[0142] Embodiment 3:

[0143] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a computer device provided by an optional embodiment of the present application. As shown in Figure 3 , the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 3 In

[0144] Processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, processor 10 can further include a hardware chip. The above hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device can be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.

[0145] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.

[0146] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of a computer device presented by a kind of mini-program landing page, etc. In addition, the memory 20 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0147] The memory 20 may include volatile memory, for example, random access memory; the memory may also include non-volatile memory, for example, flash memory, a hard disk, or a solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0148] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected through a bus or other means. Figure 3 Taking the connection through the bus as an example.

[0149] The input device 30 may receive input digital or character information, and generate key signal inputs related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (for example, an LED), and a tactile feedback device (for example, a vibration motor), etc. The above display device includes but is not limited to a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.

[0150] Embodiment 4:

[0151] Based on the above embodiments, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program executable by a processor is stored. When the program runs on the processor, the processor is caused to execute the following steps when executed:

[0152] Obtain any to-be-processed document in the rail transit industry;

[0153] Based on the document to be processed and each pre-tuned instruction prompt, input data is obtained; wherein, each of the prompts includes the following types: text name extraction instruction, table content extraction instruction, text quantity extraction instruction, content extraction instruction for the content with differences in the text and without relevant analysis, clause extraction instruction in compliance, basis integrity judgment instruction, and instruction for judging the relevance between fields and values; the text name extraction instruction is used to extract name information in the document to be processed; the table content extraction instruction is used to extract table content in the document to be processed; the text quantity instruction is used to extract the vocabulary and corresponding content representing quantity in the document to be processed; the content extraction instruction for the content with differences in the text and without relevant analysis is used to extract the content with differences but without detailed analysis in the document to be processed; the clause extraction instruction in compliance is used to extract the clauses related to compliance and their specific content in the document to be processed; the basis integrity judgment instruction is used to judge whether the basis given in the document to be processed is complete; the instruction for judging the relevance between fields and values is used to judge the relevance between fields and values in the document to be processed.

[0154] For each of the input data, through a pre-trained large model for content extraction supporting rail transit industry documents, based on this input data, a special document extraction result corresponding to this input data is obtained; wherein, the document extraction result includes a list of document extraction content key-value pairs and document extraction content.

[0155] Since the principle of the above computer-readable storage medium for solving problems is similar to that of the rail transit industry document content extraction method, the implementation of the above computer-readable storage medium can refer to the embodiments of the method, and the repeated parts will not be elaborated.

[0156] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application also intends to include these changes and modifications.

Claims

1. A method for extracting document content in the rail transit industry, characterized in that: The method comprises: Get any pending documents in the rail transit industry; Based on the document to be processed and the pre-tuned instruction prompts, input data is obtained; wherein the prompts include the following types: text name extraction instructions, table content extraction instructions, text quantity extraction instructions, content extraction instructions with differences in the text and no relevant analysis, compliance clause extraction instructions, integrity judgment instructions, and instructions for judging whether fields and values ​​are related; the text name extraction instructions are used to extract name information in the document to be processed; the table content extraction instructions are used to extract table content in the document to be processed; the text quantity instructions are used to extract words and corresponding content representing quantity in the document to be processed; the text difference and no relevant analysis are used to extract content with differences in the document to be processed but no detailed analysis; the compliance clause extraction instructions are used to extract clauses and their specific contents related to compliance in the document to be processed; the integrity judgment instructions are used to judge whether the basis given in the document to be processed is complete; the instructions for judging whether fields and values ​​are related are used to judge the correlation between fields and values ​​in the document to be processed; For each of the input data, a pre-trained content extraction model supporting rail transit industry documents is used to obtain a special document extraction result corresponding to the input data based on the input data; wherein the document extraction result includes a document extraction content key-value pair list and document extraction content.

2. The method according to claim 1, characterized in that Any prompt includes: user role information, extraction requirements based on the user role, and reference input and output examples.

3. The method according to claim 1, characterized in that Any prompt is obtained after multiple iterations of tuning. For any iteration: Through the content extraction model, based on the input data carrying the example document sample and the prompt of the current iteration, the expected special extraction result corresponding to the input data is obtained; According to the expected special extraction result and the standard special extraction result corresponding to the example document sample, the prompt of the current iteration is tuned to obtain a tuned prompt.

4. The method according to claim 1, characterized in that After obtaining any document to be processed in the rail transit industry, and before obtaining input data based on the document to be processed and each pre-tuned instruction prompt, the method further includes: The document to be processed is structured.

5. The method according to claim 1, characterized in that After obtaining any document to be processed in the rail transit industry, and before obtaining input data based on the document to be processed and each pre-tuned instruction prompt, the method further includes: If it is determined that the format of the document to be processed does not conform to the preset standard document format, the document to be processed is converted into any standard document format.

6. The method according to claim 1, characterized in that The method further comprises: Based on the special extraction results corresponding to each of the input data, a structured report is generated and output.

7. A device for extracting document content in rail transit industry, characterized in that: The device comprises: An acquisition unit is used to acquire any document to be processed in the rail transit industry; A processing unit is used to obtain input data based on the document to be processed and the pre-tuned instruction prompts; wherein the prompts include the following types: text name extraction instructions, table content extraction instructions, text quantity extraction instructions, content extraction instructions with differences in the text and no relevant analysis, compliance clause extraction instructions, integrity judgment instructions, and instructions for judging whether fields and values ​​are related; the text name extraction instructions are used to extract name information in the document to be processed; the table content extraction instructions are used to extract the content of the document to be processed. The table content; the text quantity class instruction is used to extract the words and corresponding contents representing quantity in the document to be processed; the content extraction instruction that has differences in the text and no relevant analysis is used to extract the content that has differences but no detailed analysis in the document to be processed; the compliance clause extraction instruction is used to extract the clauses related to compliance and their specific contents in the document to be processed; the basis integrity judgment instruction is used to judge whether the basis given in the document to be processed is complete; the instruction for judging whether fields and values ​​are related is used to judge the correlation between fields and values ​​in the document to be processed; The content extraction unit is used to obtain, for each of the input data, a special document extraction result corresponding to the input data based on the input data through a pre-trained content extraction model supporting rail transit industry documents; wherein the document extraction result includes a document extraction content key-value pair list and document extraction content.

8. The device according to claim 7, characterized in that The device also includes: a prompt tuning unit; The prompt tuning unit is used to obtain the tuned prompt; wherein any prompt is obtained after multiple iterations of tuning, and for any iteration: Through the content extraction model, based on the input data carrying the example document sample and the prompt of the current iteration, the expected special extraction result corresponding to the input data is obtained; According to the expected special extraction result and the standard special extraction result corresponding to the example document sample, the prompt of the current iteration is tuned to obtain a tuned prompt.

9. A computer device, characterized in that: The computer device includes a processor, and the processor is used to implement the steps of the rail transit industry document content extraction method as described in any one of claims 1-6 when executing a computer program stored in a memory.

10. A computer-readable storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements the steps of the method for extracting content from rail transit industry documents as described in any one of claims 1-6 above.