Method, system and equipment for intelligent analysis and information extraction of feasibility research report

Through natural language processing technology and machine learning methods, key information in feasibility study reports are automatically identified and extracted, solving the problems of inefficiency and accuracy of traditional review methods, and achieving efficient and accurate information extraction and reporting structure.

CN120373306AActive Publication Date: 2025-07-25YONGDAO ENG CONSULTING CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510865171.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

The traditional feasibility study report review method is inefficient, highly subjective, and is prone to omit key information, which cannot meet the needs of modern project approval.

Method used

Natural language processing technology is used to identify the title of the report chapter, and combined with Word2Vec model, named entity recognition and semantic matching technology, key business information is extracted from the report text to realize automated analysis and information extraction.

Benefits of technology

It realizes efficient structure of reports, improves the accuracy and intelligence of information extraction, adapts to documents of different formats and contents, is scalable, and is easy to review quickly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373306A_ABST
    Figure CN120373306A_ABST
Patent Text Reader

Abstract

The invention discloses a feasibility research report intelligent analysis and information extraction method, system and device, and relates to the technical field of natural language processing and machine learning. The method comprises the steps that after a feasibility research report document is read to obtain a report text, each chapter title in the report text is recognized through a regular expression or a syntactic analysis mode based on a natural language processing technology, and the title category of each chapter title is determined according to a title hierarchical structure; then, for any chapter title of which the title category belongs to the report review key object, extracting corresponding chapter content from the report text, and extracting key business information from the chapter content in combination with a Word2Vec model and a named entity recognition and semantic matching technology; and finally, summarizing, outputting and displaying each chapter title and the corresponding title category / and key business information. Therefore, the method has the advantages of high information extraction efficiency, high accuracy, universality, high intelligence, expandability and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing and machine learning, and particularly relates to a method, system and device for intelligent parsing and information extraction of feasibility study reports. Background Art

[0002] A feasibility study report is a document submitted to the decision-makers and the approving authorities for review before engaging in an economic activity (investment). Both parties need to conduct specific investigations, research and analysis on various factors such as economy, technology, production, supply and marketing, as well as various social environments and laws, to determine the favorable and unfavorable factors, whether the project is feasible, estimate the success rate, economic benefits and social effects.

[0003] Currently, with the increase in the number of investment projects and the improvement of project complexity, the need for review of project feasibility study reports by the approving parties is becoming more urgent. However, the traditional review method mainly relies on manual review, which has problems such as low efficiency, strong subjectivity and easy omission, and can no longer meet the needs of modern project approval.

[0004] Therefore, how to perform document analysis and information extraction on feasibility study reports based on automation technology, especially natural language processing (NLP) technology, in order to obtain information extraction results that are beneficial for rapid review has become an urgent research topic for those skilled in the art. Summary of the Invention

[0005] The object of the present invention is to provide a method, system, computer device, computer-readable storage medium and computer program product for intelligent parsing and information extraction of feasibility study reports, so as to solve the problems of low efficiency, strong subjectivity and easy omission existing in the traditional report review method.

[0006] To achieve the above object, the present invention adopts the following technical solutions: In the first aspect, a method for intelligent parsing and information extraction of feasibility study reports is provided, including: Obtain a feasibility study report document; Use an open-source library to identify the document format of the feasibility study report document, and select an appropriate text reading method according to the document format identification result, and read the feasibility study report document to obtain a report text; Identify each chapter title in the report text through regular expressions or syntactic analysis methods based on natural language processing technology; Determine the title category of each chapter title according to the title hierarchy structure; For any chapter title whose title category belongs to the key object of report review, extract the corresponding chapter content from the report text; Combined with the Word2Vec model, named entity recognition technology, and semantic matching technology, key business information is extracted from the chapter content, where the key business information includes project name, construction unit, investment scale, source of funds, and / or implementation plan; Summarize and output to display each chapter title and its corresponding title category / and the key business information.

[0007] Based on the above inventive content, a new solution for automatically analyzing and extracting information from a feasibility study report based on natural language processing technology is provided. That is, after reading the feasibility study report document to obtain the report text, first, through regular expressions or syntactic analysis methods based on natural language processing technology, each chapter title in the report text is identified, and the title category of each chapter title is determined according to the title hierarchy structure. Then, for any chapter title whose title category belongs to the key object of report review, the corresponding chapter content is extracted from the report text, and combined with the Word2Vec model, named entity recognition technology, and semantic matching technology, key business information is extracted from the chapter content. Finally, summarize and output to display each chapter title and its corresponding title category / and key business information. In this way, not only can an information extraction result beneficial to rapid review be obtained, achieving the purpose of efficient structuring of the report and providing data support for automated review, but also has the advantages of high efficiency, high accuracy, generality, high intelligence, and scalability of information extraction, facilitating practical application and promotion.

[0008] In a possible design, after reading the feasibility study report document to obtain the report text, the method further includes: Using regular expressions to remove redundant data, whitespace characters, and redundant blank lines between adjacent paragraphs in the report text to obtain a report text after data cleaning.

[0009] In a possible design, after reading the feasibility study report document to obtain the report text, the method further includes: Performing part-of-speech tagging processing and syntactic analysis processing on the report text based on natural language processing technology to obtain a report text after grammar correction with consistent grammar structure and semantic relationship.

[0010] In a possible design, determining the title category of each chapter title according to the title hierarchy structure includes: For each chapter title, according to the title hierarchy structure, extract the corresponding chapter level information from the report text; For any chapter title in the report text, determine the corresponding title category according to the known structural relationship and corresponding chapter level information of the feasibility study report, where the title category includes a non-terminal preface title, a terminal preface title, a non-terminal body title, a terminal body title, a non-terminal conclusion title, and a terminal conclusion title.

[0011] In a possible design, after determining the title category of each chapter title according to the title hierarchy, the method further includes: For each chapter title, extract the corresponding chapter content from the report text, perform word segmentation on the chapter content to obtain a corresponding number of segmented words, then arrange the number of segmented words in descending order of word frequency to obtain a corresponding segmented word queue, and finally extract the first N segmented words from the segmented word queue, and generate word vectors of the first N segmented words through the Word2Vec model to obtain a corresponding keyword vector group, where N represents an integer greater than or equal to 10; For each chapter title, import the corresponding keyword vector group into the content theme recognition model pre-trained based on the second machine learning algorithm to obtain the corresponding content theme recognition result and the confidence of the content theme recognition result; According to the content theme recognition results of each chapter title, adjust the display order of each chapter title according to the preset report content review order, and for at least two chapter titles with the same content theme recognition result, adjust the display order of the at least two chapter titles in descending order of confidence, where the report content review order includes multiple content themes arranged in sequence.

[0012] In a possible design, combine the Word2Vec model, named entity recognition technology, and semantic matching technology to extract key business information from the chapter content, including: Perform word segmentation on the chapter content to obtain multiple segmented words; Generate word vectors of the multiple segmented words through the Word2Vec model; According to the word vectors of the multiple segmented words, apply the K-means clustering algorithm to cluster and obtain multiple cluster centers; For each segmented word in the multiple segmented words, calculate the corresponding Euler distance to the multiple cluster centers according to the multiple cluster centers and the corresponding word vectors, and use the corresponding word as the belonging word of a certain cluster center corresponding to the shortest Euler distance among the multiple cluster centers; For each of the multiple clustering centers, all the affiliated words corresponding thereto are arranged in descending order of word frequency to obtain a corresponding affiliated word queue, and then the first M affiliated words are extracted from the affiliated word queue. Finally, the word vectors of the first M affiliated words are imported into a semantic recognition model pre-trained based on a first machine learning algorithm to obtain a corresponding semantic recognition result, where M represents an integer greater than or equal to 10; For each of the affiliated words of the first clustering center whose semantic recognition result is a named entity class semantics, named entity information is extracted from the sentence to which the corresponding word belongs by using named entity recognition technology, where the named entity information includes a project name and / or a construction unit; For each of the affiliated words of the second clustering center whose semantic recognition result is a review key point class semantics, review key point information is extracted from the sentence to which the corresponding word belongs by using semantic matching technology, where the review key point information includes an investment scale, a funding source, and / or an implementation plan; The named entity information and the review key point information are summarized to obtain key business information extracted from the chapter content, where the key business information includes a project name, a construction unit, an investment scale, a funding source, and / or an implementation plan.

[0013] In a second aspect, a feasibility study report intelligent parsing and information extraction system is provided, including a report document acquisition unit, a report text reading unit, a chapter title recognition unit, a title category determination unit, a chapter content extraction unit, a business information extraction unit, and an information summary and display unit; The report document acquisition unit is used to acquire a feasibility study report document; The report text reading unit is communicatively connected to the report document acquisition unit and is used to identify the document format of the feasibility study report document by using an open source library, and select an adapted text reading method according to the document format recognition result to read the feasibility study report document to obtain a report text; The chapter title recognition unit is communicatively connected to the report text reading unit and is used to identify each chapter title in the report text by using a regular expression or a syntactic analysis method based on natural language processing technology; The title category determination unit is communicatively connected to the chapter title recognition unit and is used to respectively determine the title category of each chapter title according to the title hierarchy structure; The chapter content extraction unit is communicatively connected to the report text reading unit and the title category determination unit respectively, and is used to extract the corresponding chapter content from the report text for any chapter title whose title category belongs to the report review key point object; The business information extraction unit is communicatively connected to the chapter content extraction unit, and is used to extract key business information from the chapter content by combining the Word2Vec model, named entity recognition technology and semantic matching technology, wherein the key business information includes project name, construction unit, investment scale, source of funds and / or implementation plan; The information summary and display unit is communicatively connected to the title category determination unit and the business information extraction unit respectively, and is used to summarize and output the display of each chapter title and the corresponding title category / and the key business information.

[0014] In a third aspect, the present invention provides a computer device, including a memory, a processor and a transceiver that are communicatively connected in sequence, wherein the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer programs and execute the feasibility study report intelligent parsing and information extraction method as described in the first aspect or any possible design in the first aspect.

[0015] In a fourth aspect, the present invention provides a computer-readable storage medium, on which instructions are stored, and when the instructions are run on a computer, the feasibility study report intelligent parsing and information extraction method as described in the first aspect or any possible design in the first aspect is executed.

[0016] In a fifth aspect, the present invention provides a computer program product, including a computer program or instructions, and when the computer program or the instructions are executed by a computer, the feasibility study report intelligent parsing and information extraction method as described in the first aspect or any possible design in the first aspect is realized.

[0017] Advantages of the above solution: (1)The present invention creatively provides a new solution for automatically analyzing and extracting information from a feasibility study report based on natural language processing technology. That is, after reading the feasibility study report document to obtain the report text, first identify each chapter title in the report text through regular expressions or syntactic analysis methods based on natural language processing technology, and determine the title category of each chapter title according to the title hierarchy structure. Then, for any chapter title whose title category belongs to the key objects of report review, extract the corresponding chapter content from the report text, and combine the Word2Vec model, named entity recognition technology, and semantic matching technology to extract key business information from the chapter content. Finally, summarize and output to display each chapter title and its corresponding title category / and key business information. In this way, not only can an information extraction result beneficial to rapid review be obtained, the report can be efficiently structured and the purpose of providing data support for automatic review can be achieved, but also it has the advantages of high efficiency, high accuracy, generality, high intelligence, and scalability in information extraction, which is convenient for practical application and promotion; (2)It has high efficiency, that is, this solution can efficiently extract information from documents in different formats, reduce manual intervention, and shorten the document analysis time; (3)It has high accuracy, that is, through technologies such as natural language processing and semantic matching, this solution can accurately extract key business information in the report to ensure the accuracy of the analysis results; (4)It has generality, that is, this solution is applicable to various types of investment project reports and can automatically adapt to documents with different formats and contents; (5)It has high intelligence, that is, this solution adopts machine learning / deep learning technology, can continuously optimize the extraction model, improve the extraction quality, and reduce errors; (6)It has scalability, that is, this solution has good scalability, can adapt to changes in different report templates, standards, and policy orientations, and supports document processing requirements in different fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0019] Figure 1 It is a schematic flowchart of the intelligent parsing and information extraction method for the feasibility study report provided by the embodiment of the present application.

[0020] Figure 2It is a schematic structural diagram of the feasibility study report intelligent parsing and information extraction system provided by the embodiments of the present application.

[0021] Figure 3 It is a schematic structural diagram of the computer device provided by the embodiments of the present application. Detailed implementation manners

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the present invention in combination with the accompanying drawings and the descriptions of the embodiments or the prior art. Obviously, the following descriptions of the structures of the accompanying drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other embodiments can be obtained based on these embodiments without creative efforts. It should be noted here that the descriptions of these embodiment modes are used to help understand the present invention, but do not constitute a limitation to the present invention.

[0023] It should be understood that although terms such as first and second etc. may be used herein to describe various objects, these objects should not be limited by these terms. These terms are only used to distinguish one object from another. For example, the first object can be called the second object, and similarly, the second object can be called the first object, without departing from the scope of the exemplary embodiments of the present invention.

[0024] It should be understood that for the term "and / or" that may appear in this article, it is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, B exists alone, or A and B exist simultaneously, etc.; another example, A, B and / or C can represent any one of A, B and C or any combination of them; for the term " / and" that may appear in this article, it is a description of another association object relationship, indicating that two relationships can exist. For example, A / and B can represent: A exists alone or A and B exist simultaneously, etc.; in addition, for the character " / " that may appear in this article, generally it represents that the associated objects before and after are an "or" relationship.

[0025] Embodiment As Figure 1As shown in the figure, the intelligent parsing and information extraction method for the feasibility study report provided in the first aspect of this embodiment can be, but is not limited to, executed by a computer device with certain computing resources, such as a cloud server, a personal computer (Personal Computer, PC, referring to a multi-purpose computer with a size, price, and performance suitable for personal use; desktop computers, laptops, small laptops, tablets, and ultrabooks all belong to personal computers), a smart phone, a personal digital assistant (Personal Digital Assistant, PDA), or a wearable device, etc. As Figure 1 As shown in the figure, the intelligent parsing and information extraction method for the feasibility study report can be, but is not limited to, including the following steps S1 to S7.

[0026] S1. Obtain the feasibility study report document.

[0027] In step S1, the feasibility study report document is the object to be reviewed, which can be, but is not limited to, uploaded by the report approval party or the report production party. In addition, the feasibility study report document supports multiple file formats, such as PDF (Portable Document Format), Word format, TXT format, and / or XML (eXtensible Markup Language) format, etc.

[0028] S2. Use an open-source library to identify the document format of the feasibility study report document, and select an appropriate text reading method according to the document format identification result, and read the feasibility study report document to obtain the report text.

[0029] In step S2, the open-source library specifically, but not limited to, adopts existing Apache POI (which is a free and open-source cross-platform Java API written in Java. Apache POI provides an API for Java programs to read and write Microsoft Office format files. Among them, POI is the abbreviation of "Poor Obfuscation Implementation", meaning "simple version of obfuscation implementation") or PDFBox (which is a pure Java class library prepared for developers to read and create PDF documents), etc. In addition, the text reading method is also an existing reading method.

[0030] After the step S2, in order to ensure that the report text has the characteristics of being clean and tidy, so as to facilitate subsequent content parsing and information extraction, preferably, after reading the feasibility study report document to obtain the report text, the method further includes but is not limited to the following steps: using regular expressions to remove redundant data, white space characters, and redundant blank lines between adjacent paragraphs, etc. from the report text, to obtain the report text after data cleaning. The regular expression (Regular Expression, often abbreviated as regex or regexp) is a pattern tool for matching, searching, and operating on text, and realizes efficient processing of strings through specific syntax rules; the core functions of the regular expression include: (1) pattern matching, that is, identifying strings in a specific format (such as email addresses, phone numbers, etc.); (2) text search and replacement, that is, quickly locating or batch modifying text fragments that meet the rules; (3) input validation, that is, ensuring that the user input conforms to a preset format (such as password complexity). Therefore, the regular expression can be conventionally applied to implement the cleaning process of redundant data, white space characters, and redundant blank lines, etc. In addition, in order to ensure the syntactic structure and semantic consistency of the report text, so as to further facilitate subsequent content parsing and information extraction, preferably, after reading the feasibility study report document to obtain the report text, the method further includes but is not limited to the following steps: performing part-of-speech tagging processing and syntactic analysis processing on the report text based on natural language processing technology, to obtain the report text after grammar correction with consistent syntactic structure and semantic relationship. The aforementioned part-of-speech tagging processing is a basic task in the NLP field, and its main goal is to determine the part of speech of each word in a sentence (such as nouns, pronouns, adjectives, adverbs, verbs, numerals, articles, prepositions, conjunctions, and interjections), so as to help the computer understand the structure and meaning of the sentence, and thus better perform subsequent processing and analysis. The aforementioned syntactic analysis processing is another important task in the NLP field, and its main goal is to determine the syntactic structure of a sentence (that is, the grammatical relationship between words in the sentence, generally represented by a tree structure, also called a syntactic tree or grammar tree), so as to help the computer understand the structure and meaning of the sentence, and thus better perform subsequent processing and analysis. Thus, the part-of-speech tagging result and the syntactic analysis result can be conventionally obtained, and based on these results, it can be conventionally judged whether the syntactic structure and semantic relationship of the report text are consistent. If not, it is modified in a conventional manner to ensure that the finally obtained report text after grammar correction has a consistent syntactic structure and semantic relationship.

[0031] S3. Identify each chapter title in the report text by means of regular expressions or syntactic analysis based on natural language processing technology.

[0032] In the step S3, since the regular expression has text matching and searching functions, various chapter titles in the report text that contain these characters or are located after these characters can be regularly recognized based on special characters such as "1.1", "1.2", or "1.1.3". In addition, since the syntactic analysis method can help the computer understand the structure and meaning of sentences, the chapter titles in the report text can also be regularly recognized according to the sentence structure and the understanding result of the meaning.

[0033] S4. Determine the title categories of the respective chapter titles according to the title hierarchy structure.

[0034] In the step S4, considering the structural relationship of the feasibility study report mainly includes three parts: the theme, the main body, and the appendix. Among them, the main body part usually includes a preface (which is used to briefly introduce the project background, basis, purpose, and its economic benefits, and explain the scope and requirements of the feasibility study), a main body (which is used to describe in detail the market survey, scale and plan analysis, technical strength and level description, source of funds analysis, economic benefit analysis, etc.), and a conclusion (which is used to summarize and generalize the entire report content). The main body part is the core area for report review, especially key aspects such as the project name, construction unit, investment scale, source of funds, and implementation plan need to be reviewed; at the same time, it is also considered that there are multiple sub-chapters under a large chapter, and the detailed content of the main text will only appear in the last-level chapter. Therefore, in order to quickly obtain the information extraction result beneficial to rapid review, it is necessary to first classify the chapter titles (that is, identify whether the corresponding chapter title is a non-last-level preface title, a last-level preface title, a non-last-level main body title, a last-level main body title, a non-last-level conclusion title, or a last-level conclusion title, etc.), so as to perform targeted information extraction according to the classification result in the follow-up. For example, extract key business information from the chapter content of the last-level main body title. To achieve the purpose of accurately classifying the chapter titles, preferably, determine the title categories of the respective chapter titles according to the title hierarchy structure, including but not limited to the following steps S41 to S42.

[0035] S41. For each of the chapter titles, extract the corresponding chapter level information from the report text according to the title hierarchy structure.

[0036] In the step S41, the title hierarchy structure can be but is not limited to being composed of serial numbers such as "one", "two", and "three", or being composed of prefixes such as "Chapter 1", "Chapter 2", and "Chapter 3". Therefore, the chapter level information of each chapter title can be regularly extracted according to these structures.

[0037] S42. For any chapter title in the said report text, determine the corresponding title category according to the known structural relationship of the feasibility study report and the corresponding chapter level information, where the title category includes but is not limited to non-final-level preface titles, final-level preface titles, non-final-level main body titles, final-level main body titles, non-final-level conclusion titles, final-level conclusion titles, etc.

[0038] In the said step S42, the known structural relationship includes three parts: theme, text, and annex, etc. Among them, the text part includes three parts: preface, main body, and conclusion, etc.; therefore, based on this structural relationship and the corresponding chapter level information, it is possible to routinely determine whether any chapter title is a non-final-level preface title, a final-level preface title, a non-final-level main body title, a final-level main body title, a non-final-level conclusion title, or a final-level conclusion title, etc.

[0039] S5. For any chapter title whose title category belongs to the key objects of report review, extract the corresponding chapter content from the said report text.

[0040] In the said step S5, the key objects of report review can be pre-specified by the report approving party according to actual review requirements. For example, the final-level preface title and the final-level main body title are specified as the key objects of report review respectively.

[0041] S6. Combine the Word2Vec model, named entity recognition technology, and semantic matching technology to extract key business information from the said chapter content, where the key business information includes but is not limited to project name, construction unit, investment scale, source of funds, and / or implementation plan, etc.

[0042] In step S6, the Word2Vec model is an existing model that maps vocabulary to high-dimensional vectors through a neural network, including two architectures, CBOW and Skip-gram, for capturing the semantic relationship between words; the named entity recognition (NER) technology is an important technology in NLP, which aims to identify name entities in texts, such as names of people, places, organization names or dates, and can therefore be used to extract key business information such as project names and / or construction units; the semantic matching technology is an important task in NLP, which aims to determine the semantic similarity or matching degree between two texts, and is widely used in various NLP tasks, such as information retrieval, question-answering systems, paraphrasing questions, dialogue systems and machine translation, and can therefore be used to extract key business information such as investment scale, source of funds and / or implementation plan. Considering that the chapter content may contain a large number of sentences, and each sentence does not necessarily contain key business information such as project name, construction unit, investment scale, funding source and / or implementation plan, in order to improve the accuracy and speed of extracting key business information, preferably, the Word2Vec model, named entity recognition technology and semantic matching technology are combined to extract key business information from the chapter content, including but not limited to the following steps S61 to S68.

[0043] S61. Perform word segmentation processing on the chapter content to obtain multiple word segments.

[0044] In the step S61, the specific process of the word segmentation processing can be implemented by but is not limited to using the jieba word segmentation tool.

[0045] S62. Generate word vectors for the multiple word segments using a Word2Vec model.

[0046] S63. Based on the word vectors of the multiple word segments, a K-means clustering algorithm is applied to obtain multiple cluster centers.

[0047] In the step S63, the K-means clustering algorithm is an iterative clustering analysis algorithm, and its steps are generally as follows: The data is pre-divided into K groups, and K objects are randomly selected as the initial cluster centers. Then, the distance between each object and each seed cluster center is calculated, and each object is assigned to the cluster center closest to it, and the cluster centers are continuously updated and iterated until the sum of squared errors (i.e., the criterion function) converges to obtain the clustering result. Considering that the existing K-means clustering algorithm will change the compactness and discreteness of the clusters due to the influence of extreme values during the clustering process, reducing the accuracy of the entire clustering result. Preferably, according to the word vectors of the multiple word segmentations, applying the K-means clustering algorithm to cluster to obtain multiple cluster centers, including but not limited to the following steps S631 to S635.

[0048] S631. For each word segmentation in the multiple word segmentations, according to the corresponding word vector and the word vectors of each other word segmentation in the multiple word segmentations, calculate the corresponding Euler distance from each other word segmentation, and count the number of distances corresponding to and less than a preset distance threshold.

[0049] In the step S631, the specific calculation formula of the Euler distance is an existing formula. For example, for a certain word segmentation, if there are 628 other word segmentations, 628 Euler distances will be calculated. If 376 of these 628 Euler distances are less than the distance threshold, then the number of distances corresponding to the certain word segmentation and less than the preset distance threshold is 376. The aforementioned number of distances positively reflects the number of other word segmentations around the semantics of the certain word segmentation in the multi-dimensional space, that is, it positively reflects the size of the neighbor density corresponding to the certain word segmentation.

[0050] S632. According to the number of distances of each word segmentation and the distribution positions of each word segmentation in the multi-dimensional space, determine multiple word segmentations in the multiple word segmentations that have a maximum value of the number of distances, where each dimension in the multi-dimensional space corresponds one-to-one with each dimension data value in the word vector.

[0051] In the step S632, the number of distances of each word segmentation in the multiple word segmentations is greater than the number of distances of the adjacent word segmentations in each dimension. That is, similar to the way of finding the extreme value through the partial derivative of a multi-variable function, according to the number of distances of each word segmentation and the distribution positions of each word segmentation in the multi-dimensional space, determine the multiple word segmentations in the multiple word segmentations that have a maximum value of the number of distances.

[0052] S633. Arrange the multiple word segmentations in descending order according to the maximum value of the number of distances to obtain a word segmentation sequence.

[0053] S634. Select the first K word segments from the word segmentation sequence, and use the word vectors of the first K word segments as K initial cluster centers one by one, where K represents the total number of preset cluster centers.

[0054] S635. Based on the K initial cluster centers, apply other steps of the K-means clustering algorithm after the initial cluster center selection step to cluster and obtain multiple cluster centers.

[0055] Thus, based on the above steps S631 - S635, the distribution information of word segmentation can be used to select multiple points with the highest density as the initial cluster centers, so as to effectively solve the problem that clustering falls into local optimum due to human factor interference or extreme value influence. Furthermore, the final clustering result can meet the condition that the similarity degree within the same cluster is the highest and the similarity degree between different clusters is the lowest, ensuring the stability of the clustering result.

[0056] S64. For each word segment in the multiple word segments, calculate the corresponding Euler distance to the multiple cluster centers according to the multiple cluster centers and the corresponding word vectors, and use the corresponding word as the belonging word of a certain cluster center corresponding to the shortest Euler distance among the multiple cluster centers.

[0057] In the step S64, for example, if the Euler distances of a certain word segment to the multiple cluster centers are calculated respectively according to the multiple cluster centers and the word vector of the certain word segment as follows: the Euler distance to cluster center A is 12 unit distances, the Euler distance to cluster center B is 34 unit distances, the Euler distance to cluster center C is 8 unit distances, the Euler distance to cluster center D is 21 unit distances, and the Euler distance to cluster center E is 17 unit distances, then the certain word segment can be used as the belonging word of cluster center C.

[0058] S65. For each cluster center in the multiple cluster centers, arrange all the corresponding belonging words in descending order of word frequency to obtain the corresponding belonging word queue, then extract the first M belonging words from the belonging word queue, and finally import the word vectors of the first M belonging words into the semantic recognition model pre-trained based on the first machine learning algorithm to obtain the corresponding semantic recognition result, where M represents an integer greater than or equal to 10.

[0059] In step S65, M is exemplified but not limited to 12. The machine learning algorithm is a core artificial intelligence algorithm that specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. It is the fundamental way to make computers intelligent. Specifically, the first machine learning algorithm preferably but not limited to uses machine learning algorithms based on graph neural networks, support vector machines, K-nearest neighbor methods, stochastic gradient descent methods, multivariable linear regression, multi-layer perceptrons, decision trees, backpropagation neural networks, or radial basis function networks, etc., so as to quickly and accurately find the patterns in the data. Thus, based on a certain amount of sample data (that is, the model input items are the word vectors of M word segments, and the model output items are the semantic labels corresponding to the M word segments), the verified semantic recognition model can be trained through a conventional calibration and verification modeling method (the specific process includes the calibration process and the verification process of the model, that is, first comparing the model simulation results with the measured data, and then adjusting the model parameters according to the comparison results to make the simulation results coincide with the actual situation). In addition, the semantic recognition results can be but not limited to named entity class semantics, review key point class semantics, or other semantics, etc., and different semantic labels can be specified in advance according to the review requirements.

[0060] S66. For each belonging word of the first clustering center with the semantic recognition result of named entity class semantics, use the named entity recognition technology to extract the named entity information from the sentence to which the corresponding word belongs. Among them, the named entity information includes but is not limited to project name and / or construction unit, etc.

[0061] In step S66, since the named entity recognition technology is an existing technology, the specific extraction process of the named entity information can be obtained through conventional derivation based on existing technical means.

[0062] S67. For each belonging word of the second clustering center with the semantic recognition result of review key point class semantics, use the semantic matching technology to extract the review key point information from the sentence to which the corresponding word belongs. Among them, the review key point information includes but is not limited to investment scale, source of funds, and / or implementation plan, etc.

[0063] In step S67, since the semantic matching technology is an existing technology, the specific extraction process of the review key point information can be obtained through conventional derivation based on existing technical means.

[0064] S68. Summarize the named entity information and the review key point information to obtain the key business information extracted from the chapter content. Among them, the key business information includes but is not limited to project name, construction unit, investment scale, source of funds, and / or implementation plan, etc.

[0065] Based on the foregoing steps S61 to S68, key business information extraction can be performed only on the sentences containing the semantic words of the review focus category and the sentences of the review focus category semantic words, without the need to extract information from each sentence. Therefore, the purpose of extraction can be targeted, which is beneficial to improving the accuracy and speed of key business information extraction.

[0066] S7. Summarize and output to display each of the chapter titles and their corresponding title categories / and the key business information.

[0067] In step S7, since this embodiment does not perform key business information extraction on the chapter titles whose title categories do not belong to the objects of report review focus, for such chapter titles, only the corresponding title categories need to be displayed. For the chapter titles whose title categories belong to the objects of report review focus, the corresponding title categories and the key business information will be displayed. In addition, considering that the listing order of the chapter titles in the feasibility study report may be inconsistent with the review order of the report content, in order to make the display order of each chapter title consistent with the review order of the report content, so as to further facilitate the rapid review of the information extraction results by the report approval party. Preferably, after respectively determining the title categories of each chapter title according to the title hierarchy structure, the method further includes but is not limited to the following steps S701 to S703.

[0068] S701. For each of the chapter titles, extract the corresponding chapter content from the report text, perform word segmentation on the chapter content to obtain a corresponding number of word segments, then arrange the number of word segments in descending order of word frequency to obtain a corresponding word segment queue, and finally extract the first N word segments from the word segment queue, and generate word vectors of the first N word segments through a Word2Vec model to obtain a corresponding keyword vector group, where N represents an integer greater than or equal to 10.

[0069] In step S701, N can also be exemplified but not limited to 12. In addition, the keyword vector group contains the word vectors of the first N word segments.

[0070] S702. For each of the chapter titles, import the corresponding keyword vector group into a content theme recognition model pre-trained based on a second machine learning algorithm to obtain a corresponding content theme recognition result and the confidence of the content theme recognition result.

[0071] In step S702, specifically, the second machine learning algorithm preferably but not limited to machine learning algorithms based on graph neural network, support vector machine, K-nearest neighbor method, stochastic gradient descent method, multivariable linear regression, multi-layer perceptron, decision tree, backpropagation neural network, or radial basis function network, etc., so as to quickly and accurately find out the rules in the data. Similarly, based on a certain amount of sample data (i.e., the input item of the model is the word vectors of N word segments, and the output item of the model is the content theme label corresponding to the N word segments), the verified content theme recognition model can be obtained through the conventional calibration and verification modeling method. In addition, the content theme recognition result can be but not limited to project background and objective theme, technical feasibility analysis theme, market feasibility analysis theme, economic feasibility analysis theme, environmental and social impact assessment theme, or risk assessment and response measure theme, etc., and different content theme labels can be specified in advance according to the review requirements.

[0072] S703. According to the content theme recognition results of the respective chapter titles, adjust the display order of the respective chapter titles according to the preset report content review order, and for at least two chapter titles with the same content theme recognition result, adjust the display order of the at least two chapter titles in descending order of confidence, where the report content review order includes a plurality of content themes arranged in sequence.

[0073] In step S703, the report content review order can be exemplified as follows: successively there are technical feasibility analysis theme, market feasibility analysis theme, economic feasibility analysis theme, environmental and social impact assessment theme, risk assessment and response measure theme, and project background and objective theme, etc., so that the display order of the respective chapter titles can be adjusted and determined according to this order. In addition, if two chapter titles both belong to the economic feasibility analysis theme: that is, the confidence of chapter title A belonging to the economic feasibility analysis theme is 70%, while the confidence of chapter title B belonging to the economic feasibility analysis theme is 80%, then the display order of these two chapter titles can be determined as follows: chapter title B and chapter title A.

[0074] Based on the feasibility study report intelligent parsing and information extraction method described in the foregoing steps S1 to S7, a new solution for automatically performing document analysis and information extraction on a feasibility study report based on natural language processing technology is provided. That is, after reading the feasibility study report document to obtain the report text, first, through regular expressions or syntactic analysis methods based on natural language processing technology, each chapter title in the report text is identified, and the title category of each chapter title is determined according to the title hierarchy structure. Then, for any chapter title whose title category belongs to the key object of report review, the corresponding chapter content is extracted from the report text, and combined with the Word2Vec model, named entity recognition technology and semantic matching technology, the key business information is extracted from the chapter content. Finally, all chapter titles and their corresponding title categories / and key business information are summarized and output for display. In this way, not only can an information extraction result beneficial to rapid review be obtained, the report can be efficiently structured and the purpose of providing data support for automatic review can be achieved, but also it has the advantages of high efficiency, high accuracy, generality, high intelligence and scalability of information extraction, which is convenient for practical application and promotion.

[0075] As Figure 2 shown, in the second aspect of this embodiment, a virtual system for implementing the feasibility study report intelligent parsing and information extraction method described in the first aspect is provided, including a report document acquisition unit, a report text reading unit, a chapter title recognition unit, a title category determination unit, a chapter content extraction unit, a business information extraction unit and an information summary and display unit; The report document acquisition unit is used to acquire a feasibility study report document; The report text reading unit is communicatively connected to the report document acquisition unit and is used to identify the document format of the feasibility study report document using an open source library, and select an adapted text reading method according to the document format recognition result, and read the feasibility study report document to obtain the report text; The chapter title recognition unit is communicatively connected to the report text reading unit and is used to identify each chapter title in the report text through regular expressions or syntactic analysis methods based on natural language processing technology; The title category determination unit is communicatively connected to the chapter title recognition unit and is used to determine the title category of each chapter title according to the title hierarchy structure; The chapter content extraction unit is communicatively connected to the report text reading unit and the title category determination unit respectively, and is used to extract the corresponding chapter content from the report text for any chapter title whose title category belongs to the key object of report review; The business information extraction unit is communicatively connected to the chapter content extraction unit, and is used to extract key business information from the chapter content by combining the Word2Vec model, named entity recognition technology, and semantic matching technology, where the key business information includes project name, construction unit, investment scale, funding source, and / or implementation plan; The information summarization and display unit is communicatively connected to the title category determination unit and the business information extraction unit respectively, and is used to summarize and output the display of each chapter title and the corresponding title category / and the key business information.

[0076] For the working process, working details, and technical effects of the foregoing system provided in the second aspect of this embodiment, reference may be made to the feasibility study report intelligent parsing and information extraction method described in the first aspect, which will not be elaborated here.

[0077] As Figure 3 shown, the third aspect of this embodiment provides a computer device that executes the feasibility study report intelligent parsing and information extraction method described in the first aspect, including a memory, a processor, and a transceiver that are communicatively connected in sequence, where the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer programs and execute the feasibility study report intelligent parsing and information extraction method described in the first aspect. Specifically, by way of example, the memory may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, first input first output (FIFO), and / or first input last output (FILO), etc.; the processor may include, but is not limited to, a microprocessor of the STM32F105 series. In addition, the computer device may include, but is not limited to, a power supply module, a display screen, and other necessary components.

[0078] For the working process, working details, and technical effects of the foregoing computer device provided in the third aspect of this embodiment, reference may be made to the feasibility study report intelligent parsing and information extraction method described in the first aspect, which will not be elaborated here.

[0079] The fourth aspect of this embodiment provides a computer-readable storage medium storing instructions for the intelligent parsing and information extraction method of the feasibility study report as described in the first aspect, that is, instructions are stored on the computer-readable storage medium, and when the instructions run on a computer, they execute the intelligent parsing and information extraction method of the feasibility study report as described in the first aspect. Among them, the computer-readable storage medium refers to a carrier for storing data, which may include, but is not limited to, computer-readable storage media such as floppy disks, optical discs, hard disks, flash memories, USB flash drives, and / or Memory Sticks. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0080] For the working process, working details, and technical effects of the aforementioned computer-readable storage medium provided in the fourth aspect of this embodiment, reference may be made to the intelligent parsing and information extraction method of the feasibility study report as described in the first aspect, which will not be elaborated here.

[0081] The fifth aspect of this embodiment provides a computer program product, including a computer program or instructions, and when the computer program or the instructions are executed by a computer, they implement the intelligent parsing and information extraction method of the feasibility study report as described in the first aspect. Among them, the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0082] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for intelligent parsing and information extraction of feasibility study reports, characterized in that, Including: Obtain a feasibility study report document; Use an open-source library to identify the document format of the feasibility study report document, and select an appropriate text reading method according to the document format recognition result, and read the feasibility study report document to obtain a report text; Identify each chapter title in the report text by using regular expressions or syntactic analysis methods based on natural language processing techniques; Determine the title category of each chapter title according to the title hierarchy structure; For any chapter title whose title category belongs to the key object of report review, extract the corresponding chapter content from the report text; Combine the Word2Vec model, named entity recognition technology and semantic matching technology to extract key business information from the chapter content, where the key business information includes project name, construction unit, investment scale, source of funds and / or implementation plan; Summarize and output to display each chapter title and its corresponding title category / and the key business information.

2. The intelligent parsing and information extraction method for the feasibility study report according to claim 1, characterized in that, After reading the feasibility study report document to obtain a report text, the method further includes: Use regular expressions to remove redundant data, whitespace characters and redundant blank lines between adjacent paragraphs in the report text to obtain a report text after data cleaning.

3. The intelligent parsing and information extraction method for the feasibility study report according to claim 1, characterized in that, After reading the feasibility study report document to obtain a report text, the method further includes: Perform part-of-speech tagging processing and syntactic analysis processing on the report text based on natural language processing techniques to obtain a report text after grammar correction with consistent grammatical structure and semantic relationship.

4. The intelligent parsing and information extraction method for feasibility study reports according to claim 1, characterized in that Determine the title category of each chapter title according to the title hierarchy structure, including: For each chapter title, extract the corresponding chapter level information from the report text according to the title hierarchy structure; For any chapter title in the report text, determine the corresponding title category according to the known structural relationship of the feasibility study report and the corresponding chapter level information, where the title category includes non-final-level preface titles, final-level preface titles, non-final-level main body titles, final-level main body titles, non-final-level conclusion titles and final-level conclusion titles.

5. The intelligent parsing and information extraction method for the feasibility study report according to claim 1, characterized in that, After determining the title category of each chapter title according to the title hierarchy structure, the method further includes: For each chapter title, extract the corresponding chapter content from the report text, perform word segmentation processing on the chapter content to obtain a corresponding number of word segments, then arrange the number of word segments in descending order of word frequency to obtain a corresponding word segment queue, and finally extract the first N word segments from the word segment queue, and generate word vectors of the first N word segments through the Word2Vec model to obtain a corresponding keyword vector group, where N represents an integer greater than or equal to 10; For each chapter title, import the corresponding keyword vector group into a content theme recognition model pre-trained based on a second machine learning algorithm to obtain a corresponding content theme recognition result and the confidence level of the content theme recognition result. According to the content theme recognition results of each chapter title, adjust the display order of each chapter title according to the preset report content review order, and for at least two chapter titles with the same content theme recognition results, adjust the display order of the at least two chapter titles in descending order of confidence, where the report content review order includes multiple content themes arranged in sequence.

6. The intelligent parsing and information extraction method for feasibility study reports according to claim 1, wherein Combining the Word2Vec model, named entity recognition technology, and semantic matching technology, extract key business information from the chapter content, including: Perform word segmentation processing on the chapter content to obtain multiple word segments; Generate word vectors of the multiple word segments through the Word2Vec model; According to the word vectors of the multiple word segments, apply the K-means clustering algorithm to cluster and obtain multiple cluster centers; For each word segment in the multiple word segments, calculate the corresponding Euler distance to the multiple cluster centers according to the multiple cluster centers and the corresponding word vectors, and use the corresponding word as the belonging word of a certain cluster center corresponding to the shortest Euler distance among the multiple cluster centers; For each cluster center in the multiple cluster centers, arrange all the corresponding belonging words in descending order of word frequency to obtain a corresponding belonging word queue, then extract the first M belonging words from the belonging word queue, and finally import the word vectors of the first M belonging words into a semantic recognition model pre-trained based on the first machine learning algorithm to obtain corresponding semantic recognition results, where M represents an integer greater than or equal to 10; For each belonging word of the first cluster center with a semantic recognition result of named entity class semantics, use named entity recognition technology to extract named entity information from the sentence to which the corresponding word belongs, where the named entity information includes a project name and / or a construction unit; For each belonging word of the second cluster center with a semantic recognition result of review key point class semantics, use semantic matching technology to extract review key point information from the sentence to which the corresponding word belongs, where the review key point information includes an investment scale, a funding source, and / or an implementation plan; Summarize the named entity information and the review key point information to obtain the key business information extracted from the chapter content, where the key business information includes a project name, a construction unit, an investment scale, a funding source, and / or an implementation plan.

7. An intelligent parsing and information extraction system for feasibility study reports, characterized in that, Including a report document acquisition unit, a report text reading unit, a chapter title recognition unit, a title category determination unit, a chapter content extraction unit, a business information extraction unit, and an information summary and display unit; The report document acquisition unit is used to acquire a feasibility study report document; The report text reading unit is communicatively connected to the report document acquisition unit and is used to use an open-source library to identify the document format of the feasibility study report document, and select an adapted text reading method according to the document format recognition result to read the feasibility study report document to obtain a report text; The chapter title recognition unit is communicatively connected to the report text reading unit and is configured to recognize each chapter title in the report text by means of regular expressions or syntactic analysis based on natural language processing techniques; The title category determination unit is communicatively connected to the chapter title recognition unit and is configured to determine the title category of each chapter title according to the title hierarchy structure; The chapter content extraction unit is communicatively connected to the report text reading unit and the title category determination unit respectively, and is configured to extract the corresponding chapter content from the report text for any chapter title whose title category belongs to the key object of report review; The business information extraction unit is communicatively connected to the chapter content extraction unit and is configured to extract key business information from the chapter content by combining the Word2Vec model, named entity recognition technology and semantic matching technology, wherein the key business information includes project name, construction unit, investment scale, source of funds and / or implementation plan; The information summary and display unit is communicatively connected to the title category determination unit and the business information extraction unit respectively, and is configured to summarize and output for display each chapter title and the corresponding title category / and the key business information; 8. A computer device, characterized in that, It includes a memory, a processor and a transceiver that are communicatively connected in sequence, wherein the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer programs and execute the feasibility study report intelligent parsing and information extraction method according to any one of claims 1 to 6; 9. A computer-readable storage medium, characterized in that ,Instructions are stored on the computer-readable storage medium, and when the instructions are run on a computer, the feasibility study report intelligent parsing and information extraction method according to any one of claims 1 to 6 is executed; 10. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or the instructions, when executed by a computer, implement the feasibility study report intelligent parsing and information extraction method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Project feasibility research report generation method and device, equipment and storage medium

    CN114357961A

  • Document information extraction method and device, electronic equipment and medium

    CN119783658A