Intelligent document identification method and computer equipment

By using a trained text detection model and neural network model to determine the meanings of polysemous words and proper nouns, and combining them with an entity detection model, the problem of insufficient understanding of complex graphic and icon structure data and inaccurate understanding of polysemous proper nouns in document intelligent recognition is solved, thus achieving accurate parsing and extraction of key information in documents.

CN120853207APending Publication Date: 2025-10-28ZHENGZHOU TOBACCO RES INST OF CNTC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510956752.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-28

Smart Images

  • Figure CN120853207A_ABST
    Figure CN120853207A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of document identification, and particularly relates to an intelligent document identification method and computer equipment. When a body text is processed, after a text detection model is used for recognizing the content of the body text, the step of semantic disambiguation processing is added, specifically, polysemy words and / or proper nouns in the content of the body text are / is recognized firstly, then the most possible word meanings of the polysemy words and / or proper nouns are determined through a parameter estimation method, and the word meanings of the polysemy words and / or proper nouns are determined. Further determining the word meaning of the polysemous word and / or the proper noun by using the trained neural network after a better result cannot be matched by the parameter estimation method (that is, the probability, estimated by the parameter estimation method, of the polysemous word and / or the proper noun belonging to the most probable word meaning is lower than a set probability threshold value); according to the method, the understanding deviation of polysemy words in different contexts is eliminated, the clear meaning of proper nouns is clear, analysis and extraction of key information of the document are achieved, a user can understand the obtained knowledge conveniently, and the document recognition effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of document recognition technology, specifically relating to a document intelligent recognition method and computer device. Background Technology

[0002] Document intelligent recognition, as a cutting-edge technology, emerged in response to the urgent need of modern industry for processing massive amounts of unstructured data. In today's information-explosive era, multimodal documents have become a primary medium for carrying knowledge and information, covering a wide range from daily office work to professional fields. These documents not only include traditional digital text files but also complex scanned documents, web pages, and multimedia resources, forming a multidimensional information matrix integrating images, text, and typography. The unique feature of multimodal documents lies in their ability to integrate data from multiple modalities, such as clear illustrations, detailed indicator charts, vivid background images, intuitive user interface elements, and rich text content, including titles, body paragraphs, and annotations. Furthermore, the document's typographical characteristics, such as font size, color, paragraph structure, UI layout, dividing lines, and grid systems, are all important components of its visual language, providing readers with a deeper understanding and contextual environment beyond pure text.

[0003] However, the richness and diversity of this information also present challenges for document recognition. Traditional manual processing methods are no longer sufficient to handle such massive amounts of data and complexity, especially in enterprise environments where quickly and accurately extracting key information from visually rich documents has become a bottleneck for improving work efficiency and decision-making quality. Therefore, document intelligence technology has emerged to automate the understanding and analysis of such documents. The core capability of multimodal document intelligent recognition lies in its ability to parse the complex structure of multimodal documents, identify and understand the multimodal information within them, and ultimately transform it into structured data. This process typically involves key technologies such as image recognition, natural language processing, and layout analysis. Through deep learning models and algorithms, it automatically identifies text content, image features, and layout styles, thereby summarizing and organizing the document's key information and logical structure. For enterprises, this means significantly improving the speed and accuracy of document processing, reducing labor costs, accelerating business processes, and thus gaining a competitive edge in the wave of digital transformation.

[0004] While existing document intelligent recognition methods have achieved some success, they still suffer from the following problems: First, there is a lack of understanding of complex graphics and icons. Parsing complex visual elements such as flowcharts, charts, maps, and background images in documents remains a challenge. Existing technologies may not be able to accurately extract the structure and data of these graphics, affecting the integrity of the information. Second, there is the problem of semantic understanding of proper nouns and polysemous words. Existing document intelligent recognition methods simply recognize the textual information without understanding the precise meaning of proper nouns and polysemous words. This affects the user's further understanding of the acquired knowledge, resulting in poor document recognition performance. Summary of the Invention

[0005] The purpose of this invention is to provide a document intelligent recognition method and computer device to solve the problem that the existing technology does not understand the precise meaning of proper nouns and polysemous words, resulting in poor document recognition performance.

[0006] To address the aforementioned technical problems, this invention provides a document intelligent recognition method, comprising:

[0007] 1) Perform layout recognition on the document to be identified to identify each block in the document. The types of each block include text images, and text images are text images in image format.

[0008] 2) Parse each block separately to convert the document's content into structured data;

[0009] The method for parsing text images includes: inputting the text image into a trained text detection model to obtain text detection results; finding polysemous words and / or proper nouns from the text detection results; determining the most likely meaning of the polysemous word and / or proper noun and the probability of the polysemous word and / or proper noun belonging to the most likely meaning based on the context of the polysemous word and / or proper noun in the text detection results; if the probability is lower than a set probability threshold, inputting the context into a trained neural network model to obtain the meaning of the polysemous word and / or proper noun.

[0010] 3) Integrate the parsing results of each block.

[0011] Furthermore, step 2) also includes: after obtaining the meanings of polysemous words and / or proper nouns, inputting the text detection results after obtaining the meanings of polysemous words and / or proper nouns into the trained entity detection model to identify the entities, where entities refer to words with clear referential meanings; and then analyzing the relationships between entities.

[0012] Furthermore, the types of each block also include an image and an image title in a matching image format, referred to as an image title image; the methods for parsing the image and image title include: inputting the image title image into a trained text detection model to obtain the text detection result of the image title image; inputting an image, the text detection result of the image title image matching the image, the context of the image in the text, and the set image parsing prompts into a trained large model to obtain the image parsing result guided by the image parsing prompts.

[0013] Furthermore, the types of each block also include tables in image format, referred to as table images; the method for parsing table images includes: inputting the table image into a trained text detection model to identify the text content and the position of the text content in the table image; inputting the table image into a trained table structure prediction model to obtain the structured information of the table, wherein the structured information of the table includes at least the cell positions; aggregating the outputs of the text detection model and the table structure prediction model to output table data with table structured information and text content.

[0014] Furthermore, the structured information of the table also includes the row and column relationships, merged cell information, row headers, and column headers.

[0015] Furthermore, if the document to be identified is not a Word document, the method for identifying the document layout includes: determining whether the document is in image format; if not, converting the document to image format and referring to it as a document image; inputting the document image into a trained layout recognition model to identify each block in the document; and sorting each block according to the reading order.

[0016] Furthermore, the method of sorting the blocks according to the reading order includes: first, sorting the blocks according to the order of the document images, with the blocks appearing earlier in the document image; then, sorting the blocks within a document image in the following way: first, sorting the blocks according to the ascending order of the vertical coordinate of the top left corner; then, for blocks with the same vertical coordinate of the top left corner, sorting them according to the ascending order of the horizontal coordinate of the top left corner.

[0017] Furthermore, the types of each block also include a reference section. The method for parsing the reference section includes: first, splitting the reference section into multiple separate entries based on the format characteristics of the reference section, with each entry including one reference information; then, using a rule matching method or a trained field recognition model to identify the fields of each entry; the rule matching method refers to using a designed regular expression for field matching; finally, storing the identified fields in a structured data format.

[0018] Furthermore, the parameter estimation method is the maximum likelihood estimation method.

[0019] To address the aforementioned technical problems, the present invention also provides a computer device, including a processor, which executes a computer program to implement the steps of the document intelligent recognition method described above.

[0020] Its beneficial effects are as follows: This invention is an improved invention. When processing the main text, after identifying the main text content using a text detection model, this invention adds a semantic disambiguation step. Specifically, it first identifies polysemous words and / or proper nouns, then uses parameter estimation to determine the most likely meaning of the polysemous word and / or proper noun. Furthermore, if the parameter estimation method cannot match a better result (i.e., the probability that the polysemous word and / or proper noun belongs to the most likely meaning estimated by the parameter estimation method is lower than a set probability threshold), it further uses a trained neural network to determine the meaning of the polysemous word and / or proper noun, eliminating the misunderstanding bias of polysemous words in different contexts, clarifying the explicit meaning of proper nouns, realizing the parsing and extraction of key information in the document, facilitating the user's understanding of the acquired knowledge, and improving the document recognition effect. Attached Figure Description

[0021] Figure 1 This is a flowchart of the document intelligent recognition method of the present invention;

[0022] Figure 2 This is an example of a document related to the present invention. Detailed Implementation

[0023] This invention provides a document intelligent recognition method. The concept involves first performing text detection using a trained text detection model during the parsing of text images in a document, followed by semantic disambiguation. Specifically, this semantic disambiguation process first uses parameter estimation to determine the most likely meaning of polysemous words and / or proper nouns. If parameter estimation fails to yield a satisfactory result, a trained neural network is then used to determine the meaning of the polysemous word and / or proper noun, thereby improving the parsing and extraction of key document information. Based on this, a document intelligent recognition method and a computer device of this invention can be implemented.

[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0025] An embodiment of a document intelligent recognition method:

[0026] This invention provides a document intelligent recognition method. Its purpose is to convert content recognizable by the human eye into content recognizable by a computer. It integrates multimodal data such as text, images, and tables within a document, transforming the document data into clear and easily searchable structured data for output, thereby enhancing document recognition performance and accuracy. Structured data, a fixed term, refers to data with a fixed format or organization, typically stored in a database, and can be queried and analyzed using predefined patterns. Unstructured data refers to data without a fixed format or organization, usually existing in the form of free text, images, audio, video, etc. In this embodiment, the document is specifically converted into HTML format. HTML is a markup language that can represent data in a structured way. The specific implementation process of this method is as follows: Figure 1 As shown below, a detailed introduction will follow.

[0027] Step 1: Obtain the document to be identified and determine its file type.

[0028] Document types mainly include Word, PDF, PPT, and images, which can be determined by the file extension. This step is primarily to distinguish between Word documents and non-Word documents.

[0029] Step two: Use a layout recognition method that matches the document's file type to perform layout recognition on the document in order to identify the various blocks in the document.

[0030] Page layout recognition refers to analyzing a document by examining its page layout, element positioning, and block division, classifying and extracting the blocks contained in different sections. If the document is a Word document, its storage format is a standardized tree structure, allowing direct extraction of its layout and content using computer languages; that is, directly extracting text, tables, and images based on its XML tree structure. However, if the document is not a Word document, its storage format may not support direct extraction of its structured content (such as an XML tree structure), thus requiring the use of a page layout recognition model. The specific page layout recognition process is as follows:

[0031] 1) Convert the content of each page of the document into an image format, referred to as a document image. Of course, if the document is already in image format (such as a scanned copy), no conversion is needed. Moreover, as an optimized processing method, the document images can be preprocessed and used as input to the layout recognition model in step 2). This preprocessing may include grayscale conversion or binarization.

[0032] 2) The layout recognition model sequentially inputs document images one by one to identify and distinguish different blocks within the document images. The identification and distinction results include the type and coordinate position of each block. The types of blocks mainly include main text, titles, illustrations, illustration titles, tables, table titles, headers, footers, reference sections, formulas, etc. The specific types can be adjusted according to the actual situation, and all blocks are in image format. The layout recognition model can use deep learning algorithms (such as YOLOv, Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs)) to distinguish the layout of each block in the document images. The training data for training the layout recognition model includes a large number of labeled document images, with annotations including block types and their corresponding coordinate positions. Furthermore, publicly available document datasets (such as PubLayNet and DocBank) can be used for training. In addition, the origin of the coordinate system on which the coordinate positions are based is the top left corner of the document image; and the target box recognized by the model is a rectangular box, and the coordinates of each block can be represented as (xmin, ymin, xmax, ymax), where (xmin, ymin) represents the coordinates of the top left corner of the block, and (xmax, ymax) represents the coordinates of the bottom right corner of the block.

[0033] 3) Sort the results identified by the layout recognition model according to the reading order to ensure that the semantics of the document are not confused. The specific process is as follows:

[0034] The blocks within each document image are sorted; that is, the blocks within a single document image are sorted. Humans generally write from top to bottom, then from left to right. Based on the coordinates of each block within a document image, the identified blocks are sorted according to this writing habit. Specifically, the y-coordinates of the top-left corner of each block can be sorted from smallest to largest; when y-coordinates are the same, the x-coordinates of the top-left corner of each block can be sorted from smallest to largest.

[0035] Step 3: Based on the layout recognition results obtained in Step 2, parse each block sequentially to ensure the contextual relevance of the content. Different parsing methods are used for different types of blocks. Specifically:

[0036] 1) For blocks containing text (including body text, titles, references, tables, etc.), OCR (Optical Character Recognition) technology is used to extract the text content. This step is fundamental to all parsing processes; it is required for all content, including body text and titles.

[0037] 2) For table-type blocks, the specific processing procedure is as follows:

[0038] ① First, a trained table structure recognition model is used to identify the structured information of the table. This structured information includes cell coordinates, row and column relationships, and merged cell information. The model also distinguishes between columns, rows, column headers, and row headers. The table structure recognition model can be based on a Transformer architecture (such as LayoutLM), which can process both text and layout information simultaneously. The training data for this model includes labeled table images and their corresponding structured information. For example, the dataset from the ICDAR (International Conference on Document Analysis and Recongnition) table recognition competition can be used as training data. The purpose of further distinguishing between rows, columns, column headers, and row headers in this step is to more accurately structure the table content and convert it into semantic HTML code. This semantic HTML structure not only faithfully preserves the original table's visual and logical structure but also improves the table's usability in downstream tasks such as data extraction, data analysis, and webpage display.

[0039] ② Use a trained text detection model to identify the text content in the table (including the locations where text content is present).

[0040] ③ Aggregate the content obtained from steps ① and ②, and export the table data in HTML format with complete table format and content to better preserve the original structure and format of the table, making the table knowledge clearer and easier to retrieve. This aggregation includes coordinate aggregation and text aggregation. Coordinate aggregation is used to solve the problem of how to reassemble text spanning multiple rows into a single cell, i.e., determining which cell the text content belongs to. Specifically, it uses the intersection and union size (the ratio of the intersection area of ​​the text detection box and the cell area to the union area (IoU)) and vertex distance between the text box coordinates obtained through the text detection model and the cell coordinates obtained through the table structure recognition model to perform single-row to multi-row aggregation. The intersection and union size is used to determine which text content belongs to the same cell; that is, when the intersection area of ​​the text detection box and the cell area equals the total area of ​​the text detection box, the text is considered to belong to that cell. The text is then sorted based on the coordinate position of the text box (e.g., the top-left corner coordinates). Text aggregation concatenates the text recognition results according to the existing text box order, from top to bottom and left to right, and combines the contents of multiple lines of text into a single string. During the concatenation process, if a cell contains multiple lines of text, a newline character (\n) or other separator is inserted between each line of text to preserve the original layout information. If the table contains multiple languages, the concatenation rules need to be adjusted according to the language characteristics (e.g., spaces need to be inserted between English words, while no extra spaces are needed for Chinese words).

[0041] 3) For blocks with images and image titles, use a large model and combine it with the set image parsing prompts to parse the images and image titles.

[0042] In this embodiment, the large model can use existing open-source models. The model input includes an image, the text detection result corresponding to the image title that matches the image, the context of the image, and the set image parsing prompts. The model output is the image parsing result that meets the requirements of the image parsing prompts. The matching of the image and its title is achieved through their coordinate positions. For example, when the image title is below the image, for an image, its bottom left corner coordinates are compared with the bottom left corner coordinates of all image titles; the image title closest to the bottom left corner coordinates of the image is the matching title. Similarly, when the image title is above the image, for an image, its top left corner coordinates are compared with the bottom left corner coordinates of all image titles; the image title closest to the bottom left corner coordinates of the image is the matching title. Furthermore, the text detection result corresponding to the matching image title is obtained in step 1) of step three.

[0043] The following example illustrates how to use a large model. For instance, the image parsing prompt could be "Please describe the specific content in the image in detail," and the output would be the parsing result from the large model. An example of the large model's output is shown below:

[0044] This image illustrates a mathematical expression and an inequality. Specifically, it contains two formulas:

[0045] 1. The formula above means:

[0046] \[

[0047] \sigma_w=\frac{3hP_1}{\pi dZb^2}

[0048] \]

[0049] This formula may be related to some physical quantity or engineering parameter.

[0050] 2. The formula below is an inequality:

[0051] \[

[0052] \sigma_w\leq[\sigma_w]

[0053] \]

[0054] This symbol represents a rounding operation on a variable.

[0055] As can be seen from the image, this expression is likely part of some physical or engineering calculation, involving the proportionality constant \(h\), the coefficient \(P_1\), and some length and angle-related variables \(d\) and \(b\). The inequality at the bottom may be restricting the range of values ​​for \(\sigma_w\).

[0056] 4) For the header / footer blocks, which usually contain information such as page numbers and chapter titles, they can be directly extracted and marked as metadata for subsequent retrieval or analysis.

[0057] 5) For the reference block, extract the entries from the reference list and convert them into structured data (such as fields for author, title, publication year, etc.). The specific steps are as follows:

[0058] ① Based on the format characteristics of the reference list, the entire reference block is split into individual entries. The splitting criteria are: each entry begins with a number (such as [1] or 1.); and entries are separated by newlines or blank lines.

[0059] Example:

[0060] [1] Zhang San, "Deep Learning Research", Computer Science, Vol. 10, No. 3, pp. 45-50, 2021.

[0061] [2] Li Si, "Advances in Natural Language Processing", Journal of Artificial Intelligence, Vol. 5, No. 2, pp. 10-20, 2022.

[0062] ② Field identification: Use rule matching or machine learning models to identify fields for each entry and extract key fields.

[0063] Rule-based approach: For common reference formats, design regular expressions (Regex) to match each field.

[0064] Machine learning-based methods: If the format of references is diverse and not fixed, a sequence labeling model (such as a BERT-based named entity recognition model) can be trained to identify fields. Input: Text of reference entries; Output: Labels for each word (such as "author", "title", "journal name", etc.).

[0065] ③ Structured storage: Store the extracted fields in a structured data format, such as JSON.

[0066] Example:

[0067]

[0068] 6) For formula blocks, formulas are usually treated as independent blocks, which can be parsed by OCR tools (such as Mathpix) and stored in LaTeX format for easy retrieval and display of mathematical formulas later.

[0069] Step four involves mining the deeper meaning of the main text, building upon step three. This step includes two aspects: semantic disambiguation and extracting entities and relationships between them.

[0070] 1) Semantic disambiguation.

[0071] ① Collect and organize sentences or texts containing polysemous words and proper nouns, and build a list containing all possible meanings for each polysemous word and proper noun.

[0072] ② Compare the results obtained from step 1) of step 3 with the results collected and organized in step ① to determine whether there are polysemous words and proper nouns in the main text. If polysemous words or proper nouns exist, determine the most likely meaning of the polysemous word / proper noun and the probability that the polysemous word / proper noun belongs to the most likely meaning based on the context of the polysemous word / proper noun in the main text.

[0073] The formula for calculating the maximum likelihood estimation method is:

[0074]

[0075] sense * =argmax sense P(sense|context)

[0076] In the formula, *sense* represents a specific meaning of a word, *context* represents the context in which the word appears, *P(sense|context)* represents the probability of the word given the context, *(context|sense)* represents the probability of the context appearing given the word meaning, *P(sense)* is the prior probability of the word meaning, and *P(context)* is the probability of the context appearing. * It is the optimal meaning of the target word.

[0077] For maximum likelihood estimation, we first need to collect training data with context, and calculate the prior probability of each word meaning and the conditional probability of its occurrence in different contexts. Then, we estimate these probability parameters by maximizing the likelihood function, which is the probability of the observed data given the model parameters. Finally, for a new context, we calculate the posterior probability of each word meaning and select the word meaning with the highest probability as the matching result.

[0078] ③ When the probability of the word meaning matched by the maximum likelihood estimation method is low (specifically, below a set probability threshold), the context of the polysemous word / proper noun is input into the trained neural network model to obtain the word meaning of the polysemous word / proper noun. The neural network model here can specifically be a long short-term memory network, including networks with attention mechanisms, or other existing network models. Specifically, the text data with context can be converted into word embedding vectors as input to the neural network. During training, a dataset with correct word meaning labels is used, and the loss function (such as cross-entropy loss) is minimized through backpropagation and optimization algorithms (such as Adam). The model parameters are gradually adjusted until the model can accurately predict the correct meaning of the polysemous word. A probability threshold is set, and the calculated probability value is compared with the threshold to achieve optimized training of the model.

[0079] It should be noted that semantic disambiguation is not required for table text because the entire text content is preserved during the HTML parsing process, without disrupting the context, thus eliminating the need for disambiguation. Furthermore, after determining the correct meaning of each polysemous word and proper noun in step ②, these meanings will be used as explanatory descriptions of that word, for example, word A (the meaning of word A). The algorithm used to determine the correct meaning of each polysemous word and proper noun in step ② is the maximum likelihood estimation method. Of course, other methods in existing technology can also be used, such as the Naive Bayes classifier; these types of parameter estimation methods are all feasible and can achieve the same effect.

[0080] 2) Mining deeper semantic information.

[0081] ① Input the semantically disambiguated text into the trained entity detection model to identify the entities within it.

[0082] In natural language processing, an entity refers to a named object or concept in text that has a specific meaning. It usually refers to words with clear referential meaning, such as names of people, places, organizations, dates, times, quantities, and monetary amounts.

[0083] The entity detection model in this step can employ existing probabilistic models, classifiers, or deep learning models to identify different types of entities in the text content. Probabilistic models can include Hidden Markov Models (HMMs) and Conditional Random Fields (CRFs). HMMs are suitable for sequence labeling tasks, such as Named Entity Recognition (NER), and can predict the entity type corresponding to each word by modeling the transition probabilities between entity labels and the emission probabilities between words and labels. Conditional Random Fields are discriminative models commonly used in sequence labeling tasks. Compared to HMMs, CRFs can better utilize contextual features (such as part-of-speech tags and syntactic structure), thereby improving the accuracy of entity recognition. Classifiers can include Support Vector Machines (SVMs) and Random Forests. SVMs map text features to a high-dimensional space, enabling classification on complex boundaries, and are suitable for entity classification tasks with small datasets. Random Forests, based on decision trees, are ensemble learning methods that can handle non-linear feature relationships and are suitable for classifying entities in text. Deep learning models can include Long Short-Term Memory networks and Transformer architecture networks (such as BERT). Long Short-Term Memory (LSTM) networks are suitable for processing sequential data and can capture long-distance dependencies, making them widely used in named entity recognition tasks. The Transformer architecture captures global contextual information through a self-attention mechanism, significantly improving entity recognition performance. It should be noted that the input to the model is preprocessed, complete text content, not individual isolated words or phrases.

[0084] The training data required to train the above models typically includes the following: First, a labeled dataset is needed, requiring a large amount of labeled text data, where each word is labeled with a specific entity type (such as person names, place names, organization names, dates, etc.). Examples include open-source datasets such as CoNLL-2003, OntoNotes, and MSRA-NER. A custom dataset is also required: collecting and labeling domain-specific text data according to the specific application scenario. For traditional machine learning models, part-of-speech tagging (POSTagging), syntactic dependency relations, context windows (several words before and after), and character-level features (such as prefixes and suffixes) are necessary. For deep learning models (such as BERT and LSTM), pre-trained language models can be used directly, and fine-tuned on labeled data.

[0085] Suppose we have the following text: "Zhang San joined a certain university on January 1, 2023, as a professor in the Department of Computer Science." After processing in step 3, step 1), the text is fed into a trained entity recognition model (such as BERT-CRF), and the model output is as follows: Zhang San → Person Name (PER), January 1, 2023 → Date (DATE), A certain university → Organization Name (ORG), Department of Computer Science → Organization Name (ORG).

[0086] ② By analyzing the grammatical, semantic, and contextual information in the text, relationships between entities are extracted, further enriching the semantic information of structured data and providing deeper insights and knowledge. Relationships between entities are diverse, and the relationships between entities differ across different domains of text content. For example, relationships can include attribution, temporal, kinship, transactional, and medical relationships. For instance, the relationship between Zhang San and XX University extracted from the previous example, "Zhang San joined XX University on January 1, 2023, as a professor in the Department of Computer Science," falls under the category of "attribution." Steps ① and ② here essentially constitute the process of constructing entities and extracting relationships between them from a knowledge graph.

[0087] Step five involves populating the content according to the predetermined storage template and document order, integrating all the parsing results to obtain a complete document parsing result containing complex relationships and semantic structures. This provides strong support for subsequent applications such as text understanding, information retrieval, and knowledge graph construction. During the final integration, for the entities obtained in step four, not only the entities themselves are integrated, but also their corresponding text content. This ensures that the structured representation of the entire document includes both entity information and the corresponding text content, enabling a better understanding of the document's overall meaning and logical relationships. The integrated document parsing result is typically output in a structured format (such as HTML, JSON, XML, etc.) for easy querying, analysis, and utilization by subsequent applications.

[0088] Example (using HTML format as an example): Suppose there is a... Figure 2 The document content shown is as follows:

[0089] The final result of integrating the above content is as follows:

[0090]

[0091]

[0092]

[0093]

[0094] An embodiment of a computer device:

[0095] An embodiment of a computer device according to the present invention includes a memory, a processor, an internal bus, and a computer program stored in the memory. The processor and the memory communicate and interact with each other via the internal bus. The processor executes the computer program to implement the steps of the method described in an embodiment of the document intelligent recognition method of the present invention. The processor can be a microprocessor (MCU), a programmable logic device (FPGA), or other processing devices; the memory can be various types of memory that store information using electrical energy, such as RAM or ROM, or other types of memory.

[0096] In summary, the present invention has the following characteristics:

[0097] 1) Multimodal Content Extraction. Multimodal content extraction involves using layout recognition models to identify different blocks in a document and distinguish between different modalities, extracting document data information from different modalities (tables, images, text) into an easily storable representation format. Specifically, for blocks containing text, OCR technology is used to extract the text content; for tables, the OCR model recognizes the text content, then a table structure prediction module predicts the table structure, and finally, the table structure and text content are aggregated to generate HTML table data with complete format and content. Furthermore, for image content in the document, a large model can be used for parsing, with specific prompts to guide the model in understanding the image content.

[0098] 2) Deep Semantic Understanding of Text. Deep semantic understanding of text mainly includes semantic disambiguation, entity recognition, and entity relation extraction. First, a corpus of sentences or texts containing polysemous words and proper nouns is constructed, and a dictionary containing all possible meanings is created. Contextual information in sentences is extracted through context analysis to determine word meanings; if the matching probability is low, a neural network model is used for further processing. Regarding entity recognition and relation extraction, entity recognition models are used to identify entities in the text, dependency parsing is used to extract dependency relations between entities, and semantic role labeling is used to label semantic roles in sentences, helping to understand the functional relationships between entities and their complex relationships.

[0099] 3) High-quality data output. The extracted information is converted into structured data, such as converting table content into HTML format, for easier further processing and display. Detailed metadata is provided for each part of the document, including entities, relationships, and contextual information, to enrich the document's content description.

[0100] Specific implementation methods have been given above, but the present invention is not limited to the described implementation methods. The basic idea of ​​the present invention lies in the above basic scheme. For those skilled in the art, designing various modified models, formulas, and parameters based on the teachings of the present invention does not require creative effort. Changes, modifications, substitutions, and variations made to the implementation methods without departing from the principles and spirit of the present invention still fall within the protection scope of the present invention.

Claims

1. A document intelligent recognition method, characterized in that, include: 1) Perform layout recognition on the document to be identified to identify each block in the document. The types of each block include text images, and text images are text images in image format. 2) Parse each block separately to convert the document's content into structured data; The method for parsing text images includes: inputting the text image into a trained text detection model to obtain text detection results; finding polysemous words and / or proper nouns from the text detection results; determining the most likely meaning of the polysemous word and / or proper noun and the probability of the polysemous word and / or proper noun belonging to the most likely meaning based on the context of the polysemous word and / or proper noun in the text detection results; if the probability is lower than a set probability threshold, inputting the context into a trained neural network model to obtain the meaning of the polysemous word and / or proper noun. 3) Integrate the parsing results of each block.

2. The document intelligent recognition method according to claim 1, characterized in that, Step 2) also includes: after obtaining the meanings of polysemous words and / or proper nouns, inputting the text detection results after obtaining the meanings of polysemous words and / or proper nouns into the trained entity detection model to identify the entities, where entities refer to words with clear referential meanings; and then analyzing the relationships between entities.

3. The document intelligent recognition method according to claim 1, characterized in that, The types of each block also include accompanying images and image titles in image formats that match the accompanying images, referred to as accompanying image title images; The methods for parsing images and captions include: inputting the image and caption image into a trained text detection model to obtain the text detection results of the image and caption image; The text detection results of an accompanying image, the title image of the accompanying image that matches the accompanying image, the context of the accompanying image in the text, and the set image parsing prompts are input into a trained large model to obtain the image parsing results guided by the image parsing prompts.

4. The document intelligent recognition method according to claim 1, characterized in that, The types of each block also include tables in image format, referred to as table images; Methods for parsing table images include: inputting the table image into a trained text detection model to identify the text content and its location in the table image; The table image is input into the trained table structure prediction model to obtain the structured information of the table, which includes at least the cell positions. The outputs of the text detection model and the table structure prediction model are aggregated to output table data with table structured information and text content.

5. The document intelligent recognition method according to claim 4, characterized in that, The structured information of the table also includes the row and column relationships, merged cell information, row headers, and column headers.

6. The document intelligent recognition method according to any one of claims 1 to 5, characterized in that, If the document to be identified is not a Word document, the layout recognition method includes: determining whether the document is in image format; if not, converting the document to image format and calling it a document image; inputting the document image into a trained layout recognition model to identify each block in the document; and sorting each block according to the reading order.

7. The document intelligent recognition method according to claim 6, characterized in that, Methods for sorting blocks according to reading order include: First, sort the blocks according to the order of the document images. The earlier the document image appears in the order, the earlier the block in that document image appears. Then, the blocks in a document image are sorted in the following way: first, the blocks are sorted in ascending order of their top-left corner ordinates; then, for blocks with the same top-left corner ordinates, they are sorted in ascending order of their top-left corner abscissas.

8. The document intelligent recognition method according to any one of claims 1 to 5, characterized in that, Each block type also includes a reference section, and the methods for parsing the reference section include: First, based on the format characteristics of the reference section, the reference section is split into multiple separate entries, each of which includes one reference entry. Then, rule matching or a trained field recognition model is used to identify the field for each entry; rule matching refers to using a well-designed regular expression for field matching. Finally, the identified fields are stored in a structured data format.

9. The document intelligent recognition method according to any one of claims 1 to 5, characterized in that, The parameter estimation method is the maximum likelihood estimation method.

10. A computer device, comprising a processor, characterized in that, The processor is used to execute a computer program to implement the steps of the method according to any one of claims 1 to 9.