Knowledge grading extraction method for scientific and technical literature in coal industry
Through the knowledge grading and extraction method for scientific and technological literature in the coal industry, PDF document processing, title grading model and identifier rule library are used to solve the problems of low extraction efficiency and high error rate in traditional methods, achieving more efficient and accurate knowledge extraction, and providing a reliable information foundation for intelligent coal mines.
Patent Information
- Application Number
- CN202510696052.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-28
AI Technical Summary
The information extraction method of traditional scientific and technological literature is difficult to efficiently process massive and unstructured scientific and technological literature data in the coal industry, resulting in deviations or abnormalities in the response generated by the corresponding models of intelligent coal mines.
A knowledge grading and extraction method for scientific and technological literature in the coal industry is proposed, and the accuracy and efficiency of knowledge grading and extraction are improved through PDF document processing, title grading model and title-oriented identifier rule library. The specific steps include converting the PDF document to a plain text MD format, defining an identifier rule library, training a title grading model, identifying and identifying a title, generating an MD text file, and performing targeted knowledge grading extraction.
This method significantly improves the accuracy and efficiency of knowledge grading extraction for scientific and technological literature in the coal industry, solves problems such as low extraction efficiency and high error rate caused by the confusion of unstructured text information, and lays the foundation for building a high-precision industry knowledge vector database.
Smart Images

Figure CN120216699A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a method for hierarchical knowledge extraction for scientific and technological literature in the coal industry. Background Art
[0002] Due to the professionalism and complexity of the coal industry, a large amount of industry knowledge and practical experience are contained in scientific and technological literature, and these knowledge and experiences are important foundations for building an intelligent coal mine. However, traditional information extraction methods for scientific and technological literature are difficult to efficiently process these massive and unstructured scientific and technological literature data, resulting in possible deviations or anomalies in the corresponding models of intelligent coal mines, that is, generating information that does not conform to the actual situation or lacks accuracy. Therefore, a more reliable method for hierarchical knowledge extraction for scientific and technological literature in the coal industry is urgently needed. Summary of the Invention
[0003] The present invention aims to solve at least one of the technical problems in the related art to some extent.
[0004] To this end, the first object of the present invention is to propose a method for hierarchical knowledge extraction for scientific and technological literature in the coal industry, which improves the accuracy and efficiency of hierarchical knowledge extraction for scientific and technological literature in the coal industry through PDF document processing, a title classification model, and an identifier rule library for titles.
[0005] The second object of the present invention is to propose a device for hierarchical knowledge extraction for scientific and technological literature in the coal industry.
[0006] The third object of the present invention is to propose an electronic device.
[0007] The fourth object of the present invention is to propose a non-transitory computer-readable storage medium storing computer instructions.
[0008] To achieve the above object, the first aspect embodiment of the present invention proposes a method for hierarchical knowledge extraction for scientific and technological literature in the coal industry, and the method includes: Convert the scientific and technological literature in the coal industry in PDF format into a coal industry document in plain text MD format, and delete the non-text identifiers contained at the beginning of each line in the coal industry document to obtain a target coal industry document; Define an identifier rule library for titles, where the identifier rule library includes language identifiers defined according to the language types of titles at all levels, and level identifiers corresponding to titles at all levels; Use a large model to synthesize multiple training titles at different levels and training texts for the training titles at all levels to form a title classification data set, and then combine with a pre-trained language model to extract the semantic features of the training titles and training texts, and train a decision tree classification algorithm to obtain a title classification model; Identify multiple target-level headings and the body text of each target-level heading in the target coal industry document through a heading classification model; Add a target language identifier and its corresponding target-level identifier to the beginning of each line of each target-level heading according to the identifier rule library, and combine the target-level headings and body text after adding the target language identifier and target-level identifier to form a standard MD text file; Generate a corresponding regularized matching identifier according to the user's question request information to match the target-level heading in the MD text file, and perform directional knowledge classification extraction on the body text under the target-level heading to obtain the extraction text of the question request information.
[0009] To achieve the above object, a second aspect embodiment of the present invention proposes a knowledge classification extraction device for coal industry scientific and technological literature, the device includes: A conversion module, configured to convert a coal industry scientific and technological literature in PDF format into a coal industry document in pure text MD format, and delete non-text identifiers contained at the beginning of each line in the coal industry document to obtain a target coal industry document; A definition module, configured to define an identifier rule library for headings, the identifier rule library includes a language identifier defined according to the language type of each level of heading, and a level identifier corresponding to each level of heading; A training module, configured to use a large model to synthesize multiple training headings at different levels and the training body text of each level of training heading to form a heading classification data set, and then combine a pre-trained language model to extract the semantic features of the training headings and training body text, and train a decision tree classification algorithm to obtain a heading classification model; An identification module, configured to identify multiple target-level headings and the body text of each target-level heading in the target coal industry document through a heading classification model; A formation module, configured to add a target language identifier and its corresponding target-level identifier to the beginning of each line of each target-level heading according to the identifier rule library, and combine the target-level headings and body text after adding the target language identifier and target-level identifier to form a standard MD text file; An extraction module, configured to generate a corresponding regularized matching identifier according to the user's question request information to match the target-level heading in the MD text file, and perform directional knowledge classification extraction on the body text under the target-level heading to obtain the extraction text of the question request information.
[0010] To achieve the above object, an embodiment of the third aspect of the present invention provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect.
[0011] To achieve the above object, an embodiment of the fourth aspect of the present invention provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in the first aspect.
[0012] The knowledge hierarchical extraction method, device, electronic device, and storage medium for coal industry scientific and technological literature provided by the embodiments of the present invention convert PDF-format coal industry scientific and technological literature into pure text MD format and then delete non-text identifiers at the beginning of each line to obtain a target coal industry document; define an identifier rule library composed of language identifiers and level identifiers for each level of headings; train a heading classification model; the heading classification model identifies multiple target-level headings and their corresponding texts in the target coal industry document; the multiple target-level headings are added with identifiers through the identifier rule library and combined with the text to generate an MD text file; regular expression matching identifiers are used to perform directional knowledge hierarchical extraction on the MD text file to obtain the extracted text. Thus, through PDF document processing, the heading classification model, and the heading-oriented identifier rule library, the accuracy and efficiency of knowledge hierarchical extraction for coal industry scientific and technological literature are improved.
[0013] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, wherein: Figure 1 is a schematic flowchart of a method for knowledge hierarchical extraction of coal industry scientific and technological literature provided by an embodiment of the present invention; Figure 2 is a schematic flowchart of another method for knowledge hierarchical extraction of coal industry scientific and technological literature provided by an embodiment of the present invention; Figure 3 is a schematic structural diagram of a device for knowledge hierarchical extraction of coal industry scientific and technological literature provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.
[0016] It should be noted that in the technical solution of the present invention, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of relevant laws and regulations.
[0017] The knowledge hierarchical extraction method, device, electronic device and storage medium for scientific and technological literature in the coal industry according to embodiments of the present invention will be described below with reference to the accompanying drawings.
[0018] Figure 1 It is a schematic flow chart of a knowledge hierarchical extraction method for scientific and technological literature in the coal industry provided by an embodiment of the present invention.
[0019] As Figure 1 shown, the method includes the following steps: Step 101, convert the scientific and technological literature in the coal industry in PDF format into a coal industry document in plain text MD format, and delete the non-text identifiers contained at the beginning of each line in the coal industry document to obtain a target coal industry document.
[0020] In some possible implementation manners, converting the scientific and technological literature in the coal industry in PDF format into a coal industry document in plain text MD format, and deleting the non-text identifiers contained at the beginning of each line in the coal industry document to obtain a target coal industry document includes: using the optical character recognition (OCR) method to convert the scientific and technological literature in the coal industry in PDF format into a coal industry document in plain text MD format (Markdown format); applying a matching algorithm to search each line in the coal industry document, and when a non-text identifier is found at the beginning of each line searched, using a deletion method to delete the non-text identifier to obtain a target coal industry document.
[0021] Specifically, the coal industry document can be: 4.7 Car stop and buffer 4.7.1 Steel wire rope When the steel wire rope is used as a car stop and a buffer element, round-strand, cross-laid steel wire ropes should be selected, and there should be no broken wires or rust on the steel wire rope. When the structure and diameter of the steel wire rope change, the resistance value of the buffer should be re-calibrated.
[0022] 4.7.2 Resistance value of the buffer The resistance values of the buffers shall be calibrated and within the designed resistance value range. The difference in the resistance values of the two buffers shall not be greater than 20%. After calibration, the parts shall have no permanent deformation or damage.
[0023] 4.7.3 Fluorescent signs for car stop barriers The car stop barriers shall have red and white alternating fluorescent signs.
[0024] Step 102: Define a title-oriented identifier rule library. The identifier rule library includes language identifiers defined according to the language types of titles at all levels, and level identifiers corresponding to titles at all levels respectively.
[0025] In some possible implementation manners, among them, when there are four levels of titles at all levels, the level identifier of the first-level title is a preset identifier, the level identifier of the second-level title is two preset identifiers, the level identifier of the third-level title is three preset identifiers, and the level identifier of the fourth-level title is four preset identifiers. This is conducive to the rapid and accurate distinction and segmentation of target coal industry documents. To achieve the precise classification and accurate extraction of target coal industry documents, specific identifiers are added in front of the titles to facilitate rapid search and extraction.
[0026] Specifically, "*" can be added at the beginning of the Chinese title line as the language identifier, and "&" can be added at the beginning of the English title line as the language identifier; the preset identifier of the first-level title can be set as "#", the level identifier of the second-level title can be set as "##", the level identifier of the third-level title can be set as "", and the level identifier of the fourth-level title can be set as "#".
[0027] Among them, Chinese titles generally appear on the first line of target coal industry documents. Add an asterisk "*" at the beginning of the first line as the language identifier for identifying the language type of the title; search for the first line where English appears in the target coal industry document, and it can be confirmed as an English title, and add an "&" in front of the English title. The results after addition are as follows: * Safety technical requirements for coal mine-used car catcher &Safety technical requirements for coal mine-used car catcher Step 103: Use the large model to synthesize multiple training titles at different levels and the training text of each level of training title respectively to form a title classification data set, and then combine the pre-trained language model to extract the semantic features of the training title and the training text, and train the decision tree classification algorithm to obtain the title classification model.
[0028] In some possible implementation manners, a large model is used to separately synthesize multiple training titles at different levels and the training texts of each level of training titles to form a title classification data set, and then combined with a pre-trained language model, the semantic features of the training titles and the training texts are extracted, and a decision tree classification algorithm is trained to obtain a title grading model, including: using the large model to separately synthesize multiple titles at different levels and the training texts of each level of titles to form a title classification data set; based on the title classification data set, combined with the pre-trained language model (Bidirectional Encoder Representations from Transformers, Bert), the semantic features of the training titles and the training texts are extracted, and the decision tree classification algorithm (GBDT) is trained to generate an initial title grading model; adopting a regularization matching method to check whether there are numbers and the symbol "dot" at the beginning of each line in the title classification data set, and calculate the number of the symbol "dot" and the number of the numbers in the middle of the symbol "dot"; if the number of the symbol "dot" is n more than the number of the numbers, it is used as a training title line, where n is a preset threshold; through the training title line, the initial title grading model is optimized to obtain a title grading model (Bert-GBDT), which improves the accuracy of title level recognition.
[0029] Among them, each level of training title can be set to no less than 5,000, with a total of no less than 25,000; the training text contains data similar to different levels of training titles, such as "For a ventilation duct length of 2.5 meters, it needs to be arranged at the heading face", this kind of data is similar to a secondary title but not a secondary title.
[0030] Step 104, identify multiple target level titles and the texts of each target level title in the target coal industry document through the title grading model.
[0031] In some possible implementation manners, specifically, if it is judged as text by the Bert-GBDT model, no level identifier is added; if it is judged as a first-level title, a "#" sign is added at the beginning of the line; if it is a second-level title, two "#" signs are added at the beginning of the line, and the result is "##"; if it is a third-level title, three "#" signs are added at the beginning of the line, and the result is "". If it is a fourth-level title, four "#" signs are added at the beginning of the line, and the result is "#"; the target level titles with added level identifiers and the texts of each target level title are as follows: # 1 Achievements in the Development of Coal Mining Theory ## 1.1 Theoretical Development Achievements "Practical Mine Pressure Control Theory", in terms of the systematicness, depth, integrity of the system of research, and the achievements obtained in guiding mining practice 1.1.1 Clear Guiding Ideology and System (1)Strictly distinguished and defined two basic concepts "mine pressure" and "manifestation of mine pressure", used "existing due to mining" to indicate the absoluteness of the existence of the former, and controlled the "manifestation of mine pressure" within the range that is safe, technically possible, and economically reasonable as mine pressure control.
[0032] Step 105: Add the target language identifier and its corresponding target level identifier to the beginning of each line of the target level headings according to the identifier rule library, and combine the target level headings and the text after adding the target language identifier and the target level identifier to form a standard MD text file.
[0033] In some possible implementation manners, after forming the standard MD text file, it further includes: searching each line of the MD text file, recording multiple target lines with target level identifiers, and simultaneously recording the line numbers and the content of each target line of the target lines; extracting the content of the target lines in the order of the line numbers, saving them in order, and removing all target level identifiers to generate the table of contents of the scientific and technological literature in the coal industry.
[0034] Specifically, search each line of the MD text file. If it is found that the beginning of the line contains "#", add the same number of spaces as the number of "#" signs at the beginning of the line, and then extract all the content of the line (the content of the target line), and simultaneously record the line number of the line (the target line).
[0035] Search each line of the MD text file. If it is found that the beginning of the line contains "", add the same number of spaces as the number of "#" signs at the beginning of the line, and then extract all the content of the line. Simultaneously record the line number of the line.
[0036] Search each line of the MD text file. If it is found that the beginning of the line contains "##", add the same number of spaces as the number of "#" signs at the beginning of the line, and then extract all the content of the line, and simultaneously record the line number of the line.
[0037] Search each line of the MD text file. If it is found that the beginning of the line contains "#", add the same number of spaces as the number of "#" signs at the beginning of the line, and then extract all the content of the line, and simultaneously record the line number of the line.
[0038] Extract the above content strictly in the order of the line numbers, save them in order, and remove all special identifiers "#", then the table of contents of the scientific and technological literature in the coal industry can be generated.
[0039] For the successfully recognized target-level headings, add level identifiers according to the level identifier rules library. For subsequent recognition of level headings, only regular matching and classification of headings with level identifiers need to be used, which improves the response speed by an order of magnitude compared to using classification algorithms for hierarchical classification matching.
[0040] Step 106: Generate corresponding regular matching identifiers according to the user's question request information to match the target-level headings in the MD text file, and perform directional knowledge classification extraction of the text under the target-level headings to obtain the extracted text of the question request information.
[0041] In some possible implementation manners, generating corresponding regular matching identifiers according to the user's question request information to match the target-level headings in the MD text file, and performing directional knowledge classification extraction of the text under the target-level headings to obtain the extracted text of the question request information, including: when the user's question request information is to extract the text under each first-level heading, generating a corresponding regular matching identifier as a preset identifier to match multiple first-level heading lines in the MD text file where each line starts with only one preset identifier, and using a program (Python) script to extract the text between two first-level heading lines from the multiple first-level heading lines as the extracted text under each first-level heading; when the user's question request information is to extract the text under any target first-level heading, generating a corresponding regular matching identifier as a preset identifier to match multiple first-level heading lines in the MD text file where each line starts with only one preset identifier, and using a program script to extract the text between the target first-level heading line and its next first-level heading line from the multiple first-level heading lines as the extracted text under the target first-level heading; when the user's question request information is to extract the text under any target second-level heading, generating a corresponding regular matching identifier as two preset identifiers to match multiple second-level heading lines in the MD text file where each line starts with only two preset identifiers, and using a program script to extract the text between the target second-level heading line and its next second-level heading line from the multiple second-level heading lines as the extracted text under the target second-level heading.
[0042] Optionally, according to the written MD text file, traverse each line to find all the first-level headings. Search each line to check if there is only one '#' at the beginning of the line. If so, it can be determined that this line is the title of the first-level heading. Then record the line number and continue to search the full text to record each first-level heading line with only one '#' at the beginning of the line. Use a Python script to copy and save the text between two first-level heading lines in sequence, and the title-based segmentation of the entire MD text file can be achieved. For example, assume that the MD text file has 7 first-level headings, then 7 line numbers will be located, denoted as na, nb, nc, nd, ne, nf, ng respectively. Use a Python script to extract the text between na and nb, between nb and nc, between nc and nd, between nd and ne, between ne and nf, and between nf and ng respectively, and integrate the text under the same first-level heading into one paragraph to achieve semantic coherence, thus completing the hierarchical segmentation of the text of the entire MD text file.
[0043] Optionally, if only the text under a certain target first-level heading needs to be extracted. For example, if it is necessary to extract the text under the first-level heading '2 Coal Mine Fire Prevention and Extinguishment Specifications', then use a Python script in combination with regularization. Regularization should not only match the level identifier '#' at the beginning of the line but also match the target first-level heading. Then search each line to locate the line where the target first-level heading is located and obtain the line number, denoted as a. At the same time, using the regularization matching principle, continue to match downward until the nearest first-level heading line containing the '#' is matched, locate this line, and obtain the line number denoted as b. Use a Python script to copy the text between the two line numbers (a, b) as the extracted text.
[0044] Optionally, if it is necessary to extract the text under a certain target second-level heading. For example, if it is necessary to extract the content under '2.6 Coal Mine Mining Specifications', then use a Python script in combination with regularization. Regularization should match the level identifier '##' at the beginning of the line and also match the target second-level heading, locate the line where the target second-level heading is located, and obtain the line number, denoted as c. At the same time, using the regularization matching principle, continue to match downward until the nearest second-level heading line containing '##' is matched, locate this line, and obtain the line number denoted as d. Use a Python script to copy the text between the two line numbers (c, d) as the extracted text.
[0045] The knowledge hierarchical extraction method for coal industry scientific and technological literature in the embodiments of the present invention converts PDF-format coal industry scientific and technological literature into pure text MD format and then deletes non-text identifiers at the beginning of each line to obtain a target coal industry document; defines an identifier rule library composed of language identifiers and level identifiers for each level of headings; trains a heading classification model; the heading classification model identifies multiple target-level headings and their corresponding main texts in the target coal industry document; the multiple target-level headings are added with identifiers through the identifier rule library and combined with the main texts to generate an MD text file; regularized matching identifiers perform directional knowledge hierarchical extraction on the MD text file to obtain the extracted text. Thus, through PDF document processing, a heading classification model, and a heading-oriented identifier rule library, the accuracy and efficiency of knowledge hierarchical extraction for coal industry scientific and technological literature are improved.
[0046] Through the addition of heading level identifiers, the structural parsing and modular classification of coal science and technology literature are accurately realized. Combining the regular expression matching and line number positioning methods, the core knowledge such as first-level headings and their main texts, second-level headings and their main texts, third-level headings and their main texts, etc. can be accurately extracted, supporting the merging and splitting of content at the same level, etc. It solves the problems of low extraction efficiency and high error rate caused by the mixing of unstructured text information, and lays a foundation for building a high-precision industry knowledge vector database. At the same time, by extracting the text knowledge in coal science and technology literature, it can provide support for the high-speed construction of a highly credible industry knowledge base, and help the large model enhance retrieval and generation.
[0047] To clearly illustrate the previous embodiment, Figure 2 The following is a schematic flowchart of another knowledge hierarchical extraction method for coal industry scientific and technological literature provided by the embodiments of the present invention, including: converting PDF-format coal industry scientific and technological literature into a target coal industry document in pure text MD format; at the same time, using a large model to synthesize multiple training headings at different levels and training main texts for each level of training headings to form a heading classification dataset, and then combining with a pre-trained language model to extract the semantic features of the training headings and training main texts, and training a decision tree classification algorithm to obtain an initial heading classification model; and using a regularized matching method to match and check the headings in the heading classification dataset to optimize the initial heading classification model to obtain a heading classification model; defining a heading-oriented identifier rule library; adding target-level identifiers corresponding to each target-level heading in the target coal industry document through the heading classification model and the identifier rule library; using regularized matching identifiers to match the target-level identifiers to perform directional knowledge hierarchical extraction of the main text under the target-level headings to obtain the extracted text. Thus, by modularly extracting the key information of coal industry scientific and technological literature, an industry knowledge vector database is constructed. The industry knowledge vector database will provide rich and structured industry knowledge resources for the generative large language model, thereby effectively improving the accuracy and reliability of the generative large language model in the application of the coal industry.
[0048] To implement the above embodiments, the present invention also provides a knowledge hierarchical extraction device for scientific and technological literature in the coal industry.
[0049] Figure 3 It is a schematic structural diagram of a knowledge hierarchical extraction device for scientific and technological literature in the coal industry provided by an embodiment of the present invention.
[0050] As Figure 3 shown, the knowledge hierarchical extraction device 30 for scientific and technological literature in the coal industry includes: a conversion module 31, a definition module 32, a training module 33, an identification module 34, a composition module 35, and an extraction module 36.
[0051] The conversion module 31 is used to convert scientific and technological literature in the coal industry in PDF format into a coal industry document in pure text MD format, and delete non-text identifiers at the beginning of each line in the coal industry document to obtain a target coal industry document; The definition module 32 is used to define an identifier rule library for titles. The identifier rule library includes language identifiers defined according to the language types of titles at all levels, and level identifiers corresponding to titles at all levels; The training module 33 is used to use a large model to synthesize multiple training titles at different levels and training texts of the training titles at all levels respectively to form a title classification data set, and then combine a pre-trained language model to extract semantic features of the training titles and training texts, and train a decision tree classification algorithm to obtain a title grading model; The identification module 34 is used to identify multiple target level titles and the texts of each target level title in the target coal industry document through the title grading model; The composition module 35 is used to add target language identifiers and their corresponding target level identifiers to the beginning of each line of each target level title according to the identifier rule library, and combine the target level titles and texts after adding the target language identifiers and target level identifiers to form a standard MD text file; The extraction module 36 is used to generate a corresponding regularized matching identifier according to the user's question request information, match the target level title in the MD text file, and perform directional knowledge hierarchical extraction of the text under the target level title to obtain the extraction text of the question request information.
[0052] Further, in a possible implementation manner of the embodiment of the present invention, the conversion module 31 is specifically used for: Using an optical character recognition method to convert scientific and technological literature in the coal industry in PDF format into a coal industry document in pure text MD format; Using a matching algorithm, search each line in the coal industry documents. When a non-text identifier is found at the beginning of each line during the search, use the deletion method to delete the non-text identifier to obtain the target coal industry document.
[0053] Further, in a possible implementation manner of the embodiment of the present invention, where there are four levels of headings at each level, the level identifier of the first-level heading is a preset identifier, the level identifier of the second-level heading is two preset identifiers, the level identifier of the third-level heading is three preset identifiers, and the level identifier of the fourth-level heading is four preset identifiers.
[0054] Further, in a possible implementation manner of the embodiment of the present invention, the training module 33 is specifically configured to: Use a large model to synthesize multiple different levels of headings and training texts for each level of heading respectively to form a heading classification data set; Based on the heading classification data set, combined with a pre-trained language model, extract the semantic features of the training headings and training texts, train a decision tree classification algorithm, and generate an initial heading classification model; Adopt a regularization matching method to check whether there are numbers and the symbol "dot" at the beginning of each line in the heading classification data set, and calculate the number of the symbol "dot" and the number of the numbers between the symbol "dot"; If the number of the symbol "dot" is n more than the number of the numbers, it is used as a training heading line, where n is a preset threshold; Optimize the initial heading classification model through the training heading line to obtain the heading classification model.
[0055] Further, in a possible implementation manner of the embodiment of the present invention, the device further includes: A search module for searching each line of the MD text file, recording multiple target lines with target level identifiers, and simultaneously recording the line numbers and target line contents of each target line; A generation module for extracting the target line contents in the order of the line numbers, saving them in order, and removing all target level identifiers to generate a table of contents for coal industry scientific and technological literature.
[0056] Further, in a possible implementation manner of the embodiment of the present invention, the extraction module 36 is specifically configured to: When the user's question request information is to extract the text under each first-level heading, generate a corresponding regularization matching identifier as a preset identifier to match multiple first-level heading lines with only one preset identifier at the beginning of each line in the MD text file, and use a program script to extract the text between two first-level heading lines as the extraction text under each first-level heading; When the user's question request information is to extract the text under any target first-level heading, a corresponding regularization matching identifier is generated as a preset identifier to match multiple first-level heading lines in the MD text file that only contain one preset identifier at the beginning of each line, and a program script is used to extract the text between the target first-level heading line and its next first-level heading line from the multiple first-level heading lines as the extracted text under the target first-level heading; When the user's question request information is to extract the text under any target second-level heading, a corresponding regularization matching identifier is generated as two preset identifiers to match multiple second-level heading lines in the MD text file that only contain two preset identifiers at the beginning of each line, and a program script is used to extract the text between the target second-level heading line and its next second-level heading line from the multiple second-level heading lines as the extracted text under the target second-level heading.
[0057] It should be noted that the foregoing explanation of the method embodiment also applies to the device of this embodiment, and will not be repeated here.
[0058] The knowledge hierarchical extraction device for coal industry scientific and technological literature according to the embodiment of the present invention converts the PDF format coal industry scientific and technological literature into a pure text MD format and then deletes the non-text identifiers at the beginning of each line to obtain the target coal industry document; defines an identifier rule library composed of language identifiers and level identifiers for each level of headings; trains a heading classification model; the heading classification model identifies multiple target level headings and their respective corresponding texts in the target coal industry document; the multiple target level headings are added with identifiers through the identifier rule library and combined with the text to generate an MD text file; the regularization matching identifier performs directional knowledge hierarchical extraction on the MD text file to obtain the extracted text. Thus, through PDF document processing, the heading classification model, and the heading-oriented identifier rule library, the accuracy and efficiency of knowledge hierarchical extraction for coal industry scientific and technological literature are improved.
[0059] To implement the above embodiment, the present invention also proposes an electronic device, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the foregoing method.
[0060] To implement the above embodiment, the present invention also proposes a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to make the computer execute the foregoing method.
[0061] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0062] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0063] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present invention includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0064] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.
[0065] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0066] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0067] In addition, each functional unit in various embodiments of the present invention may be integrated into a processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0068] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A knowledge hierarchical extraction method for scientific and technological literature in the coal industry, characterized in that, The method includes: Converting the scientific and technological literature of the coal industry in PDF format into a coal industry document in pure text MD format, and deleting the non-text identifiers at the beginning of each line in the coal industry document to obtain the target coal industry document; Defining an identifier rule library for headings, where the identifier rule library includes language identifiers defined according to the language types of headings at all levels, and level identifiers corresponding to headings at all levels respectively; Using a large model to synthesize multiple training headings at different levels and training texts for the training headings at all levels respectively to form a heading classification data set, and then combining with a pre-trained language model, extracting the semantic features of the training headings and training texts, and training a decision tree classification algorithm to obtain a heading grading model; Identifying multiple target level headings and the texts of each target level heading in the target coal industry document through the heading grading model; Adding target language identifiers and their corresponding target level identifiers respectively at the beginning of each target level heading according to the identifier rule library, and combining the target level headings and texts after adding the target language identifiers and target level identifiers to form a standard MD text file; Generating a corresponding regularized matching identifier according to the user's question request information to match the target level heading in the MD text file, and performing directional knowledge grading extraction of the text under the target level heading to obtain the extraction text of the question request information.
2. The method according to claim 1, wherein The step of converting the scientific and technological literature of the coal industry in PDF format into a coal industry document in pure text MD format, and deleting the non-text identifiers at the beginning of each line in the coal industry document to obtain the target coal industry document includes: Using an optical character recognition method to convert the scientific and technological literature of the coal industry in PDF format into a coal industry document in pure text MD format; Applying a matching algorithm to search each line in the coal industry document. When a non-text identifier is found at the beginning of each line, using a deletion method to delete the non-text identifier to obtain the target coal industry document.
3. The method according to claim 1, wherein Wherein, When there are four levels of headings at all levels, the level identifier of the first-level heading is a preset identifier, the level identifier of the second-level heading is two preset identifiers, the level identifier of the third-level heading is three preset identifiers, and the level identifier of the fourth-level heading is four preset identifiers.
4. The method according to claim 1, wherein The step of using a large model to synthesize multiple training headings at different levels and training texts for the training headings at all levels respectively to form a heading classification data set, and then combining with a pre-trained language model, extracting the semantic features of the training headings and training texts, and training a decision tree classification algorithm to obtain a heading grading model includes: Using a large model to synthesize multiple headings at different levels and training texts for the headings at all levels respectively to form a heading classification data set; Based on the heading classification data set, combining with a pre-trained language model, extracting the semantic features of the training headings and training texts, and training a decision tree classification algorithm to generate an initial heading grading model; Adopting a regularized matching method to check whether there are numbers and the symbol "dot" at the beginning of each line in the heading classification data set, and calculating the number of the symbol "dot" and the number of the numbers between the symbol "dot". If the number of "dot" symbols is n more than the number of digits, it serves as the training title line, where n is a preset threshold; Optimize the initial title classification model through the training title line to obtain the title classification model.
5. The method according to claim 1, wherein After forming the standard MD text file, it also includes: Search each line of the MD text file, record multiple target lines with target level identifiers, and at the same time record the line numbers and target line contents of each target line; Extract the target line contents in the order of line numbers, save them in order, and remove all target level identifiers to generate the table of contents of the scientific and technological literature in the coal industry.
6. The method according to claim 3, wherein Generating a corresponding regularization matching identifier according to the user's question request information to match the target level title in the MD text file and perform directional knowledge classification extraction of the text under the target level title to obtain the extraction text of the question request information, including: When the user's question request information is to extract the text under each first-level title, generate a corresponding regularization matching identifier as a preset identifier to match multiple first-level title lines in the MD text file where each line starts with only one preset identifier, and use a program script to extract the text between two first-level title lines from the multiple first-level title lines as the extraction text under each first-level title; When the user's question request information is to extract the text under any target first-level title, generate a corresponding regularization matching identifier as a preset identifier to match multiple first-level title lines in the MD text file where each line starts with only one preset identifier, and use a program script to extract the text between the target first-level title line and its next first-level title line from the multiple first-level title lines as the extraction text under the target first-level title; When the user's question request information is to extract the text under any target second-level title, generate a corresponding regularization matching identifier as two preset identifiers to match multiple second-level title lines in the MD text file where each line starts with only two preset identifiers, and use a program script to extract the text between the target second-level title line and its next second-level title line from the multiple second-level title lines as the extraction text under the target second-level title.
7. A knowledge hierarchical extraction device for scientific and technological literature in the coal industry, characterized in that, The device includes: A conversion module for converting the scientific and technological literature in the coal industry in PDF format into a pure text MD format coal industry document, and deleting the non-text identifiers at the beginning of each line in the coal industry document to obtain the target coal industry document; A definition module for defining an identifier rule library for titles. The identifier rule library includes language identifiers defined according to the language types of titles at all levels, and level identifiers corresponding to titles at all levels; A training module for using a large model to synthesize multiple training titles at different levels and training texts for each level of training titles to form a title classification dataset, and then combining with a pre-trained language model to extract the semantic features of the training titles and training texts, and training a decision tree classification algorithm to obtain a title classification model; An identification module, used to identify multiple target-level titles and the text of each target-level title in the target coal industry document through a title classification model; A building module, used for adding target language identifiers and corresponding target level identifiers at the beginning of each target level title according to the identifier rule base, and combining each target level title and text after adding the target language identifier and target level identifier to form a standard MD text file; The extraction module is used to generate a corresponding regularized matching identifier according to the user's question request information, so as to match the target level title in the MD text file, and perform directional knowledge hierarchical extraction of the text under the target level title to obtain the extracted text of the question request information.
8. The device according to claim 7, characterized in that, The conversion module is specifically used for: Use optical character recognition to convert coal industry scientific and technological documents in PDF format into coal industry documents in plain text MD format; A matching algorithm is used to search each line in the coal industry document. When a non-text identifier is found at the beginning of each line, the non-text identifier is deleted using a deletion method to obtain the target coal industry document.
9. The device according to claim 7, characterized in that, in, When each level of headings includes four level headings, the level identifier of the first-level heading is one preset identifier, the level identifier of the second-level heading is two preset identifiers, the level identifier of the third-level heading is three preset identifiers, and the level identifier of the fourth-level heading is four preset identifiers.
10. The device according to claim 7, characterized in that, The training module is specifically used for: Use the big model to synthesize multiple titles of different levels and the training texts of titles of each level to form a title classification dataset; Based on the title classification data set, combined with the pre-trained language model, the semantic features of the training titles and training texts are extracted, the decision tree classification algorithm is trained, and an initial title classification model is generated; Using the regularized matching method, check whether there are numbers and "dots" at the beginning of each line in the title classification data set, and calculate the number of "dots" and the number of numbers between the "dots"; If the number of symbol "dots" is greater than the number of numbers by n, it is used as the training title row, where n is the preset threshold; By training the title lines, the initial title classification model is optimized to obtain a title classification model.
11. The device according to claim 7, characterized in that, The device further comprises: A search module, used for searching each line of the MD text file, recording multiple target lines with target level identifiers, and recording the line number and content of each target line; The generation module is used to extract the target line content in order of line number size, save it in order, and remove all target level identifiers to generate a directory of scientific and technological literature in the coal industry.
12. The device according to claim 9, wherein The extraction module is specifically used for: In the case where the user's question request information is to extract the text under each first-level title, a corresponding regularized matching identifier is generated as a preset identifier, so as to match multiple first-level title lines in the MD text file, each of which has only one preset identifier at the beginning of the line, and use a program script to extract the text between two first-level title lines from the multiple first-level title lines as the extracted text under each first-level title; When the user's question request information is to extract the text under any target first-level heading, a corresponding regularization matching identifier is generated as a preset identifier to match multiple first-level heading lines in the MD text file where each line starts with only one preset identifier, and a program script is used to extract the text between the target first-level heading line and its next first-level heading line from the multiple first-level heading lines as the extracted text under the target first-level heading; When the user's question request information is to extract the text under any target second-level heading, a corresponding regularization matching identifier is generated as two preset identifiers to match multiple second-level heading lines in the MD text file where each line starts with only two preset identifiers, and a program script is used to extract the text between the target second-level heading line and its next second-level heading line from the multiple second-level heading lines as the extracted text under the target second-level heading.
13. An electronic device, characterized in that, Including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-6.
14. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Knowledge processing method and device based on large model, knowledge question and answer method and device based on large model, and medium
CN117743558A
Knowledge base automatic rapid construction method, system and device based on large language model technology
CN118152520A
Data slicing method, system and equipment and storage medium
CN119557427A
Knowledge question-answering method and device based on Markdown document knowledge base
CN120045517A
Systems and methods for identifying a design template matching a search query
US20240311422A1