Document structure identification method, target identification model training method, question and answer processing method, legal document query method, equipment, medium and program product
Through the target recognition model training method, the deep learning model with large-scale model parameters is used, and the document hierarchical structure is identified in combination with the text block attribute information, which solves the problem of time-consuming large-scale document recognition and achieves efficient and accurate document structure recognition.
Patent Information
- Application Number
- CN202410383967.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-10-03
AI Technical Summary
In the prior art, the process of identifying the hierarchical structure of large documents is time-consuming and inefficient, especially for long documents.
The target recognition model training method is adopted to obtain the text block information and hierarchical information of the text block to identify the hierarchical structure in the document. The deep learning model with large-scale model parameters is used to recognize the document structure. The position, size, font, color and other attribute information of the text block are combined to improve the recognition accuracy and efficiency.
It improves the accuracy and efficiency of document structure recognition, reduces the multiple predictions of the relationship between adjacent text blocks, and improves the recognition speed of the entire document.
Smart Images

Figure CN120747993A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and in particular to a method, device, medium, and program product for document structure recognition, target recognition model training, question-answering processing, and legal document query. Background Art
[0002] With the development of computer technology, the complexity of document reading and analysis tasks continues to increase, and the demand for document level recognition is growing.
[0003] Related technologies predict the relationships between two adjacent paragraphs in a document, and then use these relationships to predict the hierarchical structure of large documents. However, this approach requires predicting the relationships between two adjacent paragraphs, which is time-consuming for long documents and leaves much to be desired.
[0004] Therefore, there is an urgent need for a solution that can quickly and accurately identify the hierarchical structure of large documents. Summary of the Invention
[0005] In view of this, embodiments of this specification provide a document structure recognition method. One or more embodiments of this specification also involve a target recognition model training method, a question-and-answer processing method, a legal document query method, a document structure recognition device, a target recognition model training method, a question-and-answer processing method, a legal document query device, a computing device, a computer-readable storage medium, and a computer program product to address the technical shortcomings of the prior art in document hierarchical recognition, which is time-consuming and inefficient.
[0006] According to a first aspect of an embodiment of this specification, a document structure recognition method is provided, comprising:
[0007] Get the document to be identified;
[0008] Analyze the document to be recognized and determine the text block information of each text block in the document to be recognized;
[0009] The document to be identified is input into the target recognition model, and based on the text block information of each text block, the hierarchical structure of each text block in the document to be identified is identified, wherein the target recognition model is trained based on the sample document, the text block information of each sample text block in the sample document, and the hierarchical information of each sample text block, and the hierarchical information represents the hierarchical structure of each sample text block in the sample document.
[0010] According to a second aspect of an embodiment of this specification, a target recognition model training method is provided, which is applied to a cloud-side device, including:
[0011] Acquire a sample set, wherein the sample set includes a plurality of sample documents, the sample documents include a plurality of sample text blocks, and the sample documents carry hierarchical information of each sample text block;
[0012] determining text block information of each sample text block in the sample document;
[0013] Input the sample document into the initial recognition model, and obtain the predicted hierarchical structure of each sample text block based on the text block information of each sample text block;
[0014] The initial recognition model is trained based on the predicted hierarchical structure and hierarchical information to obtain a trained target recognition model;
[0015] The model parameters of the target recognition model are sent to the terminal device.
[0016] According to a third aspect of the embodiments of this specification, a question-answering processing method is provided, including:
[0017] Receive question information sent by front-end users and obtain target documents based on the question information;
[0018] Analyze the target document to determine the text block information of each text block in the target document;
[0019] Inputting the target document into a target recognition model, and identifying the hierarchical structure of each text block in the target document based on the text block information of each text block, wherein the target recognition model is trained based on sample documents, the text block information of each sample text block in the sample document, and the hierarchical information of each sample text block, wherein the hierarchical information represents the hierarchical structure of each sample text block in the sample document;
[0020] Determine the target text block and the target hierarchy of the target text block based on the question information and the hierarchical structure of each text block in the target document;
[0021] Based on the target text block and target level, output the answer information to the front-end user.
[0022] According to a fourth aspect of the embodiments of this specification, a legal document query method is provided, including:
[0023] Receive query information input by front-end users and obtain target legal documents based on the query information;
[0024] Analyze the target legal document to determine text block information of each text block in the target legal document;
[0025] inputting the target legal document into a target recognition model, and identifying the hierarchical structure of each text block in the target legal document based on text block information of each text block, wherein the target recognition model is trained based on sample documents, text block information of each sample text block in the sample documents, and hierarchical information of each sample text block, wherein the hierarchical information represents the hierarchical structure of each sample text block in the sample document;
[0026] Determine a target text block based on the query information and the hierarchical structure of each text block in the target document;
[0027] Based on the target text blocks and the hierarchical structure of each text block in the target legal document, the query results are output to the front-end user.
[0028] According to a fifth aspect of the embodiments of this specification, there is provided a computing device, including:
[0029] memory and processor;
[0030] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above-mentioned first aspect, second aspect, third aspect or fourth aspect method are implemented.
[0031] According to the sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned first aspect, second aspect, third aspect, or fourth aspect method.
[0032] According to the seventh aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned first aspect, second aspect, third aspect, or fourth aspect method.
[0033] In one embodiment of the present specification, the document to be identified is input into the target recognition model. Since the target recognition model is trained with the text block information of sample documents and sample text blocks, and the hierarchical information of sample text blocks, it has the ability to identify hierarchical structures based on text block information. Therefore, the target recognition model can better understand the hierarchical relationship of each text block in the document to be identified based on the text block information of each text block in the document to be identified, thereby improving the accuracy of document structure recognition. At the same time, the entire document to be identified is input into the target recognition model. At the same time, the target recognition model does not need to predict the relationship between adjacent text blocks in the document to be identified multiple times, thereby improving the efficiency of document structure recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 This is an architectural diagram of a document structure recognition system provided by one embodiment of this specification;
[0035] Figure 2 This is an architectural diagram of another document structure recognition system provided by one embodiment of this specification;
[0036] Figure 3 This is a flowchart of a document structure recognition method provided by one embodiment of this specification;
[0037] Figure 4 This is a flowchart of a target recognition model training method provided by one embodiment of this specification;
[0038] Figure 5 is a flowchart of another target recognition model training method provided by one embodiment of this specification;
[0039] Figure 6 This is a flowchart of a question-and-answer processing method provided by one embodiment of this specification;
[0040] Figure 7 This is a flowchart of a legal document query method provided by one embodiment of this specification;
[0041] FIG8( a ) is a schematic diagram of an interactive interface of a legal document query method provided in one embodiment of this specification;
[0042] FIG8( b ) is a schematic diagram of actual input and output of a target recognition model of a legal document query method provided by one embodiment of this specification;
[0043] Figure 9 This is a schematic diagram of the structure of a document structure recognition device provided by an embodiment of this specification;
[0044] Figure 10 This is a schematic diagram of the structure of a target recognition model training device provided by one embodiment of this specification;
[0045] Figure 11 This is a schematic diagram of the structure of a question-answering processing device provided by one embodiment of this specification;
[0046] Figure 12 This is a schematic diagram of the structure of a legal document query device provided by one embodiment of this specification;
[0047] Figure 13 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION
[0048] The following description sets forth numerous specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art may make similar generalizations without departing from the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0049] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit the one or more embodiments of this specification. The singular forms "a," "an," "the," and "the" used in one or more embodiments of this specification and the appended claims are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any and all possible combinations of one or more of the associated listed items.
[0050] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0051] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0052] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model, also known as a foundation model, is pre-trained on a large-scale, unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model is adaptable to a wide range of downstream tasks and has good generalization capabilities, such as large language models (LLMs) and multi-modal pre-training models.
[0053] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as natural language processing (NLP) and computer vision. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image description (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0054] First, the terms involved in one or more embodiments of this specification are explained.
[0055] Hierarchy: The way document content is organized and arranged. The purpose of the hierarchy is to help readers better understand and navigate the document content.
[0056] Documents: including but not limited to Portable Document Format (PDF), Word (a document format), Excel (a document format), pictures and other files.
[0057] RGBA value: A color representation method where RGBA represents the values of the red, green, blue, and transparency (alpha) channels. By using RGBA values, you can determine the visual style and transparency of a color.
[0058] In recent years, with the continued development of artificial intelligence and natural language processing, document reading and question-answering systems based on large models have received increasing attention and are being widely used. In various large-model products, to efficiently process and utilize long documents, it is often necessary to first identify the hierarchical structure of the document to be identified using small models, thereby improving the large model's understanding of the document.
[0059] However, there are the following shortcomings in using small models to identify the hierarchical structure of documents: First, in the process of identifying the hierarchical structure of the document to be identified based on the small model, text blocks are usually extracted based on visual information, and then semantic analysis is performed on adjacent text blocks. Relationships are classified according to the results of the semantic analysis. Relationships usually include parent-child relationships, sibling relationships, and other relationships. For example, the title and paragraph are parent-child relationships, paragraphs are sibling relationships, and paragraphs and page numbers are other relationships. Finally, based on the relationship between two adjacent text blocks, the hierarchical structure of the entire document to be identified is restored. It can be seen that the small model focuses on the local information between two adjacent paragraphs and lacks global information of the entire document, resulting in inaccurate recognition of the hierarchical structure of the entire document to be identified. In addition, for long documents, it is necessary to calculate the relationship between a large number of adjacent text blocks, which takes a long time. For example, N text blocks require at least N-1 calculations. In addition, the small model usually has limited understanding of information such as the semantics of the text blocks, so the effect is often less than people expect. In addition, the position, font, color and other information of the text block also plays an important role in the recognition of hierarchical structure, but the small model has limited effect in utilizing this information, and it is difficult to fully explore the structural characteristics of the document. In response to the above problems, this specification provides a document structure recognition method. One or more embodiments of this specification also involve a target recognition model training method, a question and answer processing method, a legal document query method, a document structure recognition device, a target recognition model training, a question and answer processing device, a legal document query device, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments. See Figure 1 , Figure 1An architectural diagram of a document structure recognition system provided by an embodiment of the present specification is shown. The document structure recognition system may include a client 100 and a server 200; the client 100 is used to send a document to be recognized to the server 200; the server 200 is used to analyze the document to be recognized and determine the text block information of each text block in the document to be recognized; the document to be recognized is input into a target recognition model, and based on the text block information of each text block, the hierarchical structure of each text block in the document to be recognized is recognized, wherein the target recognition model is trained based on a sample document, the text block information of each sample text block in the sample document, and the hierarchical information of each sample text block, and the hierarchical information represents the hierarchical structure of each sample text block in the sample document; the hierarchical structure is sent to the client 100; the client 100 is also used to receive and display the hierarchical structure sent by the server 200. Through this system, the document to be identified is input into the target recognition model. Since the target recognition model is trained with the text block information of sample documents and sample text blocks, as well as the hierarchical information of sample text blocks, it has the ability to identify hierarchical structures based on text block information. Therefore, the target recognition model can better understand the hierarchical relationship between each text block in the document to be identified based on the text block information of each text block in the document to be identified, thereby improving the accuracy of document structure identification. At the same time, when the entire document to be identified is input into the target recognition model, the target recognition model does not need to repeatedly predict the relationship between adjacent paragraphs in the document to be identified, thereby improving the efficiency of document structure identification. Figure 2 , Figure 2The following is an architecture diagram of another document structure recognition system provided by an embodiment of the present specification. The document structure recognition system may include multiple clients 100 and a server 200, wherein the client 100 may include an end-side device and the server 200 may include a cloud-side device. A communication connection can be established between multiple clients 100 through the server 200. In the document structure recognition scenario, the server 200 is used to provide document structure recognition services between multiple clients 100. Multiple clients 100 can act as senders or receivers respectively and realize communication through the server 200. Users can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In the document structure recognition scenario, the user can send a document to be recognized to the server 200 through the client 100, and the server 200 generates a hierarchical structure based on the document to be recognized, and pushes the hierarchical structure to other clients that have established communication. The client 100 and the server 200 are connected through a network. The network provides a medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. Data transmitted by the client 100 may require encoding, transcoding, compression, and other processing before being released to the server 200. The client 100 can be a browser, an application (APP), a web application such as Hypertext Markup Language 5 (H5), a lightweight application (also known as a mini-program, a type of lightweight application), or a cloud application. The client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by the server 200, such as the Real-Time Communication (RTC) SDK. The client 100 can be deployed in an electronic device and rely on the device or certain applications in the device to operate. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, tablet computer, or personal computer. Electronic devices can also typically be configured with various other types of applications, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc. The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers that support background training for models used on clients, and servers that process data sent by clients.It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with a blockchain. The server can also be a cloud server for basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host equipped with artificial intelligence technology. It is worth noting that the document structure recognition method provided in the embodiments of this specification is generally performed by the server. However, in other embodiments of this specification, the client can also have similar functions to the server to perform the document structure recognition method provided in the embodiments of this specification. In other embodiments, the document structure recognition method provided in the embodiments of this specification can also be performed jointly by the client and the server. See. Figure 3 , Figure 3This is a flowchart of a document structure recognition method provided by one embodiment of this specification, specifically including the following steps. Step 302: Obtaining the document to be recognized. In actual applications, obtaining the document to be recognized can involve the client searching based on user requirements and document fragments to obtain documents that meet the user's requirements. Alternatively, the client can obtain documents directly input by the user, or the user can enter a document link to the document to be recognized, and the client can use resources such as a search engine or document library to search and obtain relevant documents to be recognized. Documents to be recognized include, but are not limited to, text, tables, and images. Step 304: Analyzing the document to be recognized to determine text block information for each text block in the document to be recognized. A text block is a continuous section of content in the document to be recognized, and can take the form of a paragraph, sentence, title, list, etc. Text block information indicates the hierarchical structure of the text block in the document to be recognized, and can include text content, location information (e.g., start and end locations in the document), and format information (e.g., font, color, size). Specifically, analyzing the document to be recognized is to obtain text block information for each text block. Determining text block information can provide richer semantic and structural information, facilitating subsequent document structure recognition. The analysis process may include dividing text blocks, recognizing text, extracting location information, etc. The specific analysis method can be performed using natural language processing (NLP) technology, computer vision technology, text segmentation algorithms, text positioning algorithms, color recognition algorithms, or a combination of the above technologies, and this specification does not specifically limit this. In actual applications, after obtaining the document to be recognized, the content in the document to be recognized can be divided into different text blocks according to preset rules, such as segmentation; text recognition technology, such as optical character recognition (OCR), is used to extract the text in the document to be recognized as editable text, and the recognized text content is used as text block information; the text block information can be preprocessed, such as removing noise, correcting recognition errors, and unifying the text format. In one or more embodiments of the present specification, taking the document to be identified "Example Document.pdf" containing three paragraphs a, b, and c as an example, after obtaining the "Example Document.pdf", the document structure identification system divides the "Example Document.pdf" into three text blocks A, B, and C according to the paragraphs, wherein text block A includes paragraph a, text block B includes paragraph b, and text block C includes paragraph c; OCR is used to extract the text in text blocks A, text blocks B, and text blocks C as editable text, and the identified text content is used as text block information; the text block information can be preprocessed, such as removing noise, correcting recognition errors, and unifying text formats.Furthermore, the text block information includes text content; analyzing the document to be identified and determining the text block information of each text block in the document to be identified includes: determining each text block in the document to be identified; performing text recognition on each text block to obtain the text content of each text block.
[0060] Specifically, the determination of text blocks and text recognition are performed to extract and understand the text content in the document. The text content can be extracted through OCR, deep learning, random forest algorithm, etc. The text content can provide a basis for subsequent semantic analysis.
[0061] In practical applications, for images or scanned documents, OCR technology can be used to convert the text in the image into readable text content. For documents to be recognized that can be directly edited, such as Word documents, the text content in the document to be recognized is directly read, and the recognized text content is further cleaned and processed, including removing irrelevant characters, incorrect punctuation caused by noise, correcting errors, etc., and the processed text content is saved as text block information.
[0062] Based on this, when the document to be recognized is an image, since images cannot be directly edited or extracted, methods such as optical character recognition (OCR), deep learning, and random forest algorithms can be used to convert the text in the document to be recognized into editable text. The processed text is then saved as text blocks for subsequent processing and analysis. This process can extract text content from the image-based document as text blocks, providing a foundation for subsequent document structure recognition.
[0063] Furthermore, the text block information also includes block attribute information; after determining each text block in the document to be identified, it also includes: performing layout analysis on each text block to obtain the block attribute information of each text block; based on the text block information of each text block, identifying the hierarchical structure of each text block in the document to be identified, including: based on the text content of each text block and the block attribute information of each text block, identifying the hierarchical structure of each text block in the document to be identified.
[0064] The block attribute information refers to other attributes or features related to the text block, such as the position, size, font, color, style, etc. of the text block.
[0065] Specifically, layout analysis and obtaining block attribute information for text blocks are intended to provide a more comprehensive understanding of the document's structure. This helps the target recognition model understand the arrangement, size, font style, and other characteristics of each text block in the document to be recognized, further improving the depth and breadth of its understanding of the document to be recognized, thereby providing support for subsequent document structure recognition. Layout analysis can include processes such as region detection, layout analysis, and font analysis of the document to be recognized.
[0066] In practical applications, after identifying each text block in the document to be recognized, layout analysis is performed to obtain block attribute information for each text block. Specifically, layout analysis analyzes the document's structure, layout, and design to extract information such as the text block's location, size, font, color, and style. The object recognition model can identify the hierarchical structure of each text block in the document to be recognized based on the text content, location, size, font, color, and style of the text block.
[0067] In one or more embodiments of the present specification, the position block attribute information can be determined by identifying the coordinates of the text block; for the size block information, the size of each text block, including the width and height, can be obtained through layout analysis; for the font block attribute information, the layout analysis can extract the font information of the text block, including the font name, size, thickness, etc.; for the color block attribute information, the layout analysis can identify the pixel values of the pixel points in the text block and determine the color of the text block; for the style block attribute information, it is detected whether there are bold, italic, underline and other styles in the text content. Taking the document to be identified as an annual report as an example, the text content of the first text block is: Annual Report, position coordinates: (x=50, y=20), size: width=300px, height=50px, font information: Ar ia l, size=24pt, thickness=bold, style: no italics, underline, etc. The target recognition model can identify the first text block as the title level based on the above information; the content of the second text block is the specific content of the annual report, position coordinates: (x=50, y=100), size: width=400px, height=200px, font information: Times New Roman, size=12pt, thickness=normal, color: RGB(0,0,0) (black), style: no italics, underline, etc. The target recognition model can identify the second text block as the body level under the title level based on the above information.
[0068] Based on this, obtaining block attribute information such as position, size, font, color, and style can enable the target recognition model to more comprehensively understand the structure of the document, understand the arrangement, size, font style and other features of each text block in the document to be identified, and further improve the depth and breadth of understanding of the document to be identified, thereby providing support for subsequent document structure recognition.
[0069] Furthermore, the block attribute information includes at least one of layout type, position information and text attribute information; performing layout analysis on each text block to obtain the block attribute information of each text block includes at least one of the following steps: identifying the layout type of each text block; determining the position information of each text block based on the position of each text block in the document to be identified; performing text recognition on each text block to determine the text attribute information of each text block.
[0070] The layout types include text, tables, and pictures, and the text attribute information includes font, font size, and font color.
[0071] Specifically, by identifying the layout type, positional information, and text attributes of text blocks, the target recognition model can gain a deep understanding of the document's text block type, layout structure, and text attribute characteristics, thereby providing fundamental support for document analysis and document-level recognition. The specific process of layout analysis includes: distinguishing the layout type of each text block, such as title, body, notes, and quotes; determining the exact position of each text block within the document based on the document's layout, including positional information such as coordinates within the page and position within the header / footer, thereby understanding the spatial relationship of the text blocks; and utilizing text recognition technology to extract attribute information of the text content within the text blocks, such as font, size, and style.
[0072] In practical applications, layout analysis is performed on the document to be identified. By analyzing the structure, layout and design of the document, the layout type of each text block can be identified. The layout type of each text block can be identified as text, table and picture by combining image processing, OCR and line segment detection algorithms; by detecting the position of the text block in the document, its coordinates, boundaries, width, height and other information in the document can be determined, which can be achieved using computer vision and image processing technologies such as bounding box detection algorithms or contour detection algorithms; for text blocks containing text, text recognition can be performed to obtain its text content and text attribute information such as font, font size and font color, which can be achieved by combining OCR, font and color analysis algorithms.
[0073] In one or more embodiments of the present specification, taking the document format to be identified as PDF as an example, with respect to the layout type, the document is first structurally analyzed by utilizing a line segment detection algorithm or other layout analysis algorithm to detect whether there is a table in the text block, and the layout type of the text block with a table is determined to be a table; for non-table text blocks, image processing and OCR technology are used for analysis, and the layout type of the text block with a text ratio less than a preset threshold is determined to be a picture; and the layout type of the text block with a text ratio greater than a preset threshold is determined to be text.
[0074] Based on the position information, the coordinates of each text block in the page of the document to be recognized are determined, such as one or more of the center coordinates, the upper left corner coordinates, or the lower right corner coordinates.
[0075] For text attribute information, OCR technology is used, combined with font and color analysis algorithms, to obtain text attribute information such as font, font size and font color of each text block.
[0076] Based on this, by identifying the layout type, position information and text attribute information of each text block, the target recognition model can deeply understand the structure and content organization of the document to be identified, and more accurately identify the hierarchical structure of the document to be identified, the relationship between each element and the layout characteristics of the document to be identified. It fully combines the hints of various block attribute information on the document structure, and improves the accuracy of the target recognition model in identifying the document structure.
[0077] Furthermore, based on the position of each text block in the document to be identified, the position information of each text block is determined, including: determining the target document page number of the target text block in the document to be identified, wherein the target text block is any text block; and determining the position information of the target text block in the target document page corresponding to the target document page number.
[0078] The position information may be coordinates indicating the position of the target text block, such as one or more of center point coordinates, upper left corner coordinates, and lower right corner coordinates.
[0079] Specifically, to determine the specific location of the target text block, the document page number on which the target text block is located is determined based on the length and area of the text block. The relative position and size of the target text block on the target document page are calculated to determine the specific location of the text block on the page, such as its coordinates or the page layout area it is located in. Computer vision and image processing techniques can be used to analyze the layout structure and layout of the document to determine the document page number on which the target text block is located. Based on the page number on which the target text block is located, its relative position and size on the target document page are calculated to determine its specific location on the page.
[0080] In practical applications, for any target text block, the page range of the target text block in the document to be identified is determined, and based on the paging rules carried by the document to be identified, such as the length of the document as a page, the document page number where the target text block is located is determined in combination with the page range of the target text block in the document to be identified; on the target document page number where the target text block is located, computer vision and image processing techniques, such as a bounding box detection algorithm or a contour detection algorithm, are used to determine the coordinates of the target text block position, or, by analyzing the target text block, its relative position ratio in the target document page is determined, and the specific position of the target text block is calculated using the ratio of the width and height of the page to the position of the target text block.
[0081] In a specific embodiment of the present specification, taking the position information as the upper left corner coordinates and the lower right corner coordinates of the target text block as an example, for the target text block, the page range of the target text block in the document to be identified is determined to be the vertical range of 104 cm to 116 cm, and the horizontal range of 2 cm to 12 cm. According to the paging rules of the document to be identified, a vertical distance of 25 cm is one document page. It can be determined that the target text block is on page 4. Further calculation shows that the upper left corner coordinates of the target text block on page 4 are [2, 4], and the lower right corner coordinates are [12, 16].
[0082] Step 306: Input the document to be identified into the target recognition model, and based on the text block information of each text block, identify the hierarchical structure of each text block in the document to be identified, wherein the target recognition model is trained based on the sample document, the text block information of each sample text block in the sample document, and the hierarchical information of each sample text block, and the hierarchical information represents the hierarchical structure of each sample text block in the sample document.
[0083] Among them, the hierarchical information can be a structural label indicating the hierarchical structure of each sample text block during the training process of the target recognition model. The hierarchical structure represents the hierarchical relationship and logical order between the various parts of the document (such as titles, chapters, paragraphs, etc.). This representation of the hierarchical structure is very important for understanding the structure and content of the document, and is helpful for further text analysis and processing. The hierarchical structure generally includes parent-child relationships, sibling relationships, and other relationships. The parent-child relationship refers to one paragraph being the parent of another paragraph, the sibling relationship refers to two paragraphs having a common parent, and other relationships refer to two paragraphs each having a different parent. For example, in the following document part:
[0084] "
A
[0085]
B
[0086]
C
[0087]
D
[0088] [E] Request submission tomorrow."
[0089] [A] and [B] are father and son, [A] and [C] are father and son, [A] and [D] are father and son, [B], [C] and [D] are brothers, and [B], [C], [D] and [E] are other relationships.
[0090] Specifically, the trained target recognition model is used to predict the hierarchical structure of the input document to be recognized based on the text block information of each text block in the document to be recognized and to understand the hints of the hierarchical structure provided by the text block information.
[0091] In practical applications, the document to be recognized may contain text block information for each text block. In this case, this information is fed into the target recognition model along with the document to be recognized. Alternatively, the information can be extracted and fed into the target recognition model along with the document to be recognized. Before feeding the document to be recognized into the target recognition model, the target recognition model must be trained using sample documents, the text block information for each text block within the sample documents, and the hierarchical information for each sample text block. This allows the target recognition model to identify the hierarchical structure of the input document based on the text block information for each text block. After training the target recognition model, the entire document to be recognized is fed into the target recognition model. The target recognition model uses the knowledge gained from training and the text block information for each text block within the document to be recognized to identify the hierarchical structure of each text block within the document to be recognized. Based on the text block information, the target recognition model infers the logical order and hierarchical relationships of the text blocks within the document to be recognized and assigns an appropriate hierarchical structure to each text block, such as title, subtitle, body, paragraph, etc.
[0092] In one or more embodiments of this specification, using legal documents as an example, during the training of a target recognition model, the system obtains various legally relevant sample documents, along with text block information and hierarchical information for each text block within the sample documents, to train the target recognition model. Once trained, the model is capable of identifying the hierarchical structure of legal documents based on the text information within each text block within the legally relevant document.
[0093] During the object recognition model's inference process, the document structure recognition system receives a legal document to be identified, which includes different types of text blocks, such as chapter titles, paragraphs, and clauses. The object recognition model leverages the knowledge gained through training to infer the logical order and hierarchical relationships of each text block within the legal document based on its underlying information. For example, it marks chapter titles as first-level titles, marks paragraphs within each chapter as body text, and identifies and labels other hierarchical information, such as clauses.
[0094] In one or more embodiments of the present specification, the target recognition model can identify the hierarchical structure of the document to be identified by analyzing the semantic information of the text content in each text block. For example, the semantic information of the first text block of the document to be identified indicates that the first paragraph mentions content A, content B and content C, the semantic information of the second text block indicates that the second paragraph explains content A in detail, the semantic information of the third text block indicates that the third paragraph explains content B in detail, and the semantic information of the fourth text block indicates that the fourth paragraph explains content C in detail. Then, the target recognition model can determine that the first text block and the second text block, the third text block, and the fourth text block are in a parent-child relationship, and the second text block, the third text block, and the fourth text block are in a brother relationship.
[0095] In one or more embodiments of the present specification, the target recognition model can identify the hierarchical structure of each text block in the document to be identified by retrieving preset key text content in each text block, such as punctuation marks with obvious hierarchical cues such as colons and semicolons, and other keywords with obvious hierarchical cues such as "first", "secondly", "first", and "second".
[0096] Furthermore, the document to be identified is input into the target recognition model, and based on the text block information of each text block, the hierarchical structure of each text block in the document to be identified is identified, including: inputting the document to be identified into the target recognition model, and the target recognition model determines the semantic information of each text block based on the text content of each text block, and based on the semantic information of each text block, identifying the hierarchical structure of each text block in the document to be identified.
[0097] Among them, semantic information includes the meaning of words, the logical relationship of sentences, the understanding of context, and the actual things or concepts represented by language expressions.
[0098] Specifically, the object recognition model extracts semantic information from text content to better understand the text. This helps the model more accurately understand the structure and organization of the document to be recognized, providing a foundation for subsequent document hierarchical structure recognition. Specifically, semantic information from text content can be extracted through methods such as NLP and feature extraction.
[0099] In practical applications, the target recognition model will perform semantic analysis and understanding based on the text content of the text block, thereby encoding the meaning and semantics of each text block. By analyzing the similarity, relevance and contextual information between text blocks, the hierarchical relationship between text blocks can be determined, including parent-child relationship, sibling relationship, etc.
[0100] In one or more embodiments of the present specification, taking the example of a document to be identified including four text blocks, the semantic information of the first text block of the document to be identified indicates that the first paragraph mentions content A, content B, and content C, the semantic information of the second text block indicates that the second paragraph explains content A in detail, the semantic information of the third text block indicates that the third paragraph explains content B in detail, and the semantic information of the fourth text block indicates that the fourth paragraph explains content C in detail, then the target recognition model can determine that the first text block and the second text block, the third text block, and the fourth text block are in a parent-child relationship, and the second text block, the third text block, and the fourth text block are in a brother relationship with each other.
[0101] Based on this, the target recognition model uses its semantic understanding and analysis of text blocks to more accurately identify the connection methods between text blocks, thereby understanding and analyzing the intrinsic structure and meaning of the text, helping the target recognition model understand the structure of the text, and providing a basis for subsequently combining semantic information to infer the relationship and hierarchical structure between text blocks.
[0102] Furthermore, before inputting the document to be recognized into the target recognition model and identifying the hierarchical structure of each text block in the document to be recognized based on the text block information of each text block, it also includes: obtaining task indication information, wherein the task indication information is used to indicate the task type and output format of the target recognition model; inputting the document to be recognized into the target recognition model and identifying the hierarchical structure of each text block in the document to be recognized based on the text block information of each text block, including: inputting the document to be recognized and the task indication information into the target recognition model, identifying the hierarchical structure of each text block in the document to be recognized based on the text block information of each text block according to the task type, and outputting the hierarchical structure according to the output format.
[0103] Specifically, the diversity of document structures leads to non-uniform hierarchical structure formats in the output, which makes it difficult for users to read and view. Therefore, task instruction information is used to guide the task type and output format of the target recognition model, so that the model can identify the hierarchical structure in the document according to specific task requirements and output the results in a predetermined format. This helps to make the model more flexible and customizable, and meet the needs of specific application scenarios. The process includes: obtaining task instruction information from a configuration file preset by the system, which includes indicating the task type and output format of the target recognition model. According to the task type, based on the text block information of each text block, the target recognition model identifies the hierarchical structure of each text block in the document to be identified, and the model outputs the identified text block hierarchical structure in the specified output format according to the output format requirements.
[0104] In one or more embodiments of the present specification, taking the prompt information as "a list of paragraph contents of a given file, each line format is: [paragraph number] paragraph text content [paragraph page number (starting from 1), paragraph layout type (text, table, picture, etc.), coordinates of the upper left and lower right corners of the paragraph text block, RGBA value]. Your task is to format the hierarchical structure between output paragraphs (the paragraph structure is represented by a # sign before the paragraph, and only the [paragraph number] is returned). The example return format is: #[2], the number of #s indicates which level the paragraph number is located at" as an example, the task instruction information is obtained from the system preset configuration file, and the task and type indicated by the task prompt information are "formatting the hierarchical structure between output paragraphs (the paragraph structure is represented by a # sign before the paragraph, and only the [paragraph number] is returned". According to the task type, based on the text block information of each text block, the target recognition model identifies the hierarchical structure of each text block in the document to be recognized, and the model outputs the recognized text block hierarchical structure according to the specified output format according to the output format requirements, for example:
[0105] “【2】
[0106] #【3】
[0107] #【4】
[0108] #【5】
[0109] ##【6】
[0110] ###【7】” The number of “#” represents the level of the text block. The fewer the number of “#”, the higher the level. [2] is the highest level, [3], [4], and [5] are sub-levels of [2], [6] is a sub-level of [5], and [7] is a sub-level of [6].
[0111] Based on this, task instruction information plays a guiding and normative role in the target recognition model, which helps to improve the accuracy, efficiency and applicability of the model, and also helps to unify the format and integration of the output results of the target recognition model.
[0112] Furthermore, before inputting the document to be recognized into the target recognition model and identifying the hierarchical structure of each text block in the document to be recognized based on the text block information of each text block, it also includes: obtaining a sample set, wherein the sample set includes multiple sample documents, the sample documents include multiple sample text blocks, and the sample documents carry the hierarchical information of each sample text block; determining the text block information of each sample text block in the sample document; inputting the sample document into the initial recognition model, and obtaining the predicted hierarchical structure of each sample text block based on the text block information of each sample text block; training the initial recognition model based on the predicted hierarchical structure and the hierarchical information to obtain a trained target recognition model.
[0113] The hierarchical information of the sample text block may be manually annotated structural labels.
[0114] Specifically, since the document structure recognition model in the current technology does not fully utilize the text block information, the prediction accuracy needs to be improved. Therefore, this manual will use multiple sample documents, where the sample documents include multiple sample text blocks and the sample documents carry the hierarchical information of each sample text block to train the initial recognition model, so that the initial recognition model has the function of using text block information to predict the hierarchical structure of the document to be recognized.
[0115] In practical applications, before training begins, the training data must first be acquired and preprocessed, and a large number of representative sample documents must be collected. These documents should cover various situations that the model may encounter. For example, if the goal is to identify the hierarchical structure of legal documents, then these sample documents can include various types of legal documents, such as contracts, regulations, treaties, etc.; the collected sample documents are segmented into sample text blocks, such as titles, chapters, paragraphs, clauses, etc.; for each sample text block, it is necessary to annotate it with corresponding hierarchical information to indicate its hierarchical structure in the document, such as title, subtitle, body, etc. Labeling hierarchical information usually requires manual participation. Relevant personnel can add appropriate hierarchical information to each sample text block based on the semantics and logic of each sample text block; the sample documents and the hierarchical information of the sample text blocks in the sample documents are combined into a sample set to ensure that this sample set is as comprehensive and diverse as possible, covering various types of sample documents.
[0116] After obtaining the sample set, the document structure recognition system can divide the content of the sample document into different sample text blocks according to preset rules, such as segment division, and use OCR technology and layout analysis technology to extract the text in the text block into editable text, and determine the block attribute information such as the position, font, and color of the text block. At this point, the preprocessing stage is ready.
[0117] During the training phase, the initial recognition model predicts the hierarchical structure of each text block based on the text block information of each sample text block in the sample document, compares and evaluates the predicted hierarchical structure output by the initial recognition model with the hierarchical information in the sample document, and calculates indicators such as the accuracy and recall rate of the prediction results to measure the performance of the initial recognition model. Based on the evaluation results, the initial recognition model is further optimized, such as adjusting the initial recognition model architecture, adjusting parameters, and increasing training data. The adjustment process can be iterated multiple times until the initial recognition model achieves satisfactory performance and serves as the target recognition model.
[0118] Based on this, the initial recognition model is trained according to the sample document, the text block information of each text block in the sample document, and the hierarchical information of the sample text block, so that the initial recognition model has the function of using the text block information to predict the hierarchical structure of the document to be recognized, so that the initial recognition model can better understand the hierarchical relationship of each text block in the sample document, and then adjust the model parameters of the initial recognition model according to the predicted hierarchical structure and hierarchical information, so that the final target recognition model can be more accurate.
[0119] Before inputting the sample document into the initial recognition model and obtaining the predicted hierarchical structure of each sample text block based on the text block information of each sample text block, it also includes: obtaining preset prompt content; constructing prompt information according to the sample document, the text block information of each sample text block and the hierarchical information according to the example information in the prompt content; inputting the sample document into the initial recognition model and obtaining the predicted hierarchical structure of each sample text block based on the text block information of each sample text block, including: inputting the prompt information into the initial recognition model and obtaining the predicted hierarchical structure of each sample text block based on the text block information of each sample text block.
[0120] In one or more embodiments of this specification, the preset prompt content is "a list of paragraph contents of a given file, each line format is: [paragraph number] paragraph text content [paragraph page number (starting from 1), paragraph layout type (text, table, picture, etc.), paragraph text block upper left corner and lower right corner position coordinates, RGBA value]. Your task is to format the hierarchical structure between the output paragraphs (use the # sign before the paragraph to indicate the paragraph structure, and only return the [paragraph number]). The example return format is: # [2], the number of # refers to the level at which the paragraph number is located." as an example, The example information in the sample content is “[paragraph number] paragraph text content [paragraph page number (starting from 1), paragraph layout type (text, table, picture, etc.), coordinates of the upper left and lower right corners of the paragraph text block, RGBA value]”. For each document block of a sample document, prompt information including text block information and hierarchical information is constructed according to the format indicated by the example information, for example, “[2] Text content of sample text block [5, text, paragraph secondary title, [328, 677], [707, 705], RGBA [204, 153, 51, 1]]”. Finally, the prompt information is input into the initial recognition model, and the predicted hierarchical structure of each sample text block is obtained based on the text block information of each sample text block.
[0121] Based on this, the prompt information plays a guiding and normative role in the initial recognition model, which helps to improve the accuracy, efficiency and applicability of the initial recognition model, helps the initial recognition model to better understand the text block information of the sample text block, and also helps to unify the format and integration of the output results of the initial recognition model.
[0122] Furthermore, after inputting the document to be identified into the target recognition model and identifying the hierarchical structure of each text block in the document to be identified based on the text block information of each text block, it also includes: sending the hierarchical structure of each text block in the document to be identified to the front-end user; receiving an evaluation message sent by the front-end user for the hierarchical structure; generating an inquiry guidance message based on the evaluation message, and sending the inquiry guidance message to the front-end user; collecting feedback information input by the front-end user based on the inquiry guidance message, and obtaining an updated sample set based on the feedback information; and using the updated sample set to train the target recognition model.
[0123] Evaluation messages are user feedback on the hierarchical structure of the document to be identified, provided by the document structure recognition system. Inquiry guidance messages are questions the document structure recognition system asks the user about the document structure recognition results, encouraging them to provide further feedback and suggestions. Feedback can include user confirmation, corrections, and suggestions regarding the document structure recognition system's recognition results.
[0124] Specifically, the recognition results of the target recognition model cannot be guaranteed to be accurate. In order to continuously improve and optimize the model and adapt to user needs, it is necessary to collect the latest training data in a timely manner. This process may include: sending the document hierarchical structure identified by the target recognition model to the front-end user interface for user evaluation and feedback; collecting front-end user evaluation messages on the hierarchical structure, including evaluations on aspects such as correctness; generating corresponding query guidance messages based on user evaluation messages, and initiating targeted inquiries to users to gain a deeper understanding of user needs and improvement points; collecting user feedback information: collecting user feedback on the query guidance messages, including specific modification suggestions for the document hierarchical structure; updating the sample set and retraining the model: updating the sample set based on user feedback, including adding new sample documents, text blocks, and their hierarchical information, and retraining the target recognition model to improve the model's performance and adaptability. Continuously improving and optimizing the target recognition model through user feedback allows the target recognition model to better adapt to user needs and improve recognition accuracy and effectiveness.
[0125] In one or more embodiments of the present specification, taking the document to be identified as a resume document as an example, the document structure recognition system parses the personal information, education experience, work experience, skills and other parts of the resume document into a hierarchical structure based on their positional relationship in the document, and sends this hierarchical structure to the front-end user for display; the front-end user evaluates the document structure, such as pointing out that the hierarchical structure of the education experience is incorrectly labeled, or suggesting that the skills part be subdivided into professional skills and soft skills, and generates an evaluation message which is received and recorded by the document structure recognition system; the document structure recognition system can generate corresponding inquiry guidance messages based on the user's evaluation message, for example, the document structure recognition system may ask the user for Suggestions for the subdivision of education experience and skills are sent to the front-end user to guide the user to provide further feedback and suggestions; the front-end user can answer the guidance inquiry messages raised by the document structure recognition system, correct incorrect annotations or provide more detailed hierarchical structure suggestions, and the document structure recognition system collects and records these feedback information; the document structure recognition system converts the user's feedback information into an updated sample set, such as adding the user's corrected education experience annotations to the sample set, or organizing and adding the user's suggested subdivided skill hierarchical structure, and retraining the target recognition model in the hope of improving the target recognition model's recognition accuracy of the resume document structure.
[0126] Based on this, evaluation information, inquiry guidance information and feedback information play a key role in the interaction process between users and the system. Continuously improving and optimizing the target recognition model enables the target recognition model to better adapt to and understand user needs and expectations, thereby guiding the target recognition model to be improved and optimized, and improving recognition accuracy and user satisfaction.
[0127] Furthermore, after the document to be recognized is input into the target recognition model and the hierarchical structure of each text block in the document to be recognized is identified based on the text block information of each text block, it also includes: sending the hierarchical structure of each text block in the document to be recognized to the front-end user, wherein the hierarchical structure includes the block identifier of the text block; receiving the target text block identifier fed back by the front-end user; based on the target text block identifier, determining the target text block corresponding to the target text block identifier from the document to be recognized; and sending the text content of the target text block to the front-end user.
[0128] Specifically, the front-end user may need to view and confirm the content of a specific text block. After returning the document level to the client, the server needs to determine the target text block based on the user's further instructions and return it to the client. This process specifically includes: the user selects the target text block identifier from the front-end interface, the server receives the user's feedback information, and based on the target text block identifier fed back by the user, determines the target text block corresponding to the target text block identifier from the document to be identified, and sends the text content of the target text block to the front-end user, allowing the user to view and confirm the content of the corresponding text block. It is understandable that the above process can also be performed locally on the client.
[0129] In practical applications, the hierarchical structure of each text block in the document to be recognized is sent to the front-end user. The text block hierarchical structure includes the block identifiers of the text blocks, which are used to locate and identify the hierarchical relationship of each text block in the document. When reading the document, the front-end user can select a specific text block, for example, by clicking or selecting the text block. The block identifier of the selected target text block is then sent back to the document structure recognition system. Based on the block identifier of the target text block, the document structure recognition system determines the target text block corresponding to the block identifier in the document to be recognized, locates and extracts the corresponding target text block, and sends the text content of the target text block to the front-end user.
[0130] In one or more embodiments of this specification, the hierarchical structure of each text block in the document to be identified is
[0131] “【2】
[0132] #【3】
[0133] #【4】
[0134] #【5】
[0135] ##【6】
[0136] ###【7】" as an example, the front-end user selects the block identifier of 【5】, the document structure recognition system determines the text block corresponding to 【5】, and returns the text content of the text block corresponding to 【5】 to the front-end user.
[0137] Based on this, the hierarchical structure of text blocks is sent to the front-end user, enabling a visual representation of the document structure. Users can select specific text blocks based on their needs and provide feedback to the system. Based on the user's feedback, the system locates the corresponding target text block and returns its text content to the front-end user. This process allows users to more easily browse and manipulate specific text blocks in the document.
[0138] One embodiment of the present specification inputs the entire document to be identified and the text block information of each text block in the document to be identified into the target recognition model, so that the target recognition model can better understand the hierarchical relationship between each text block in the document to be identified, and fully combines the text block information of each text block in the document to be identified to perform document structure recognition, thereby improving the accuracy of document structure recognition. At the same time, the entire document to be identified is input into the target recognition model, so that the target model does not need to predict the relationship between adjacent paragraphs in the document to be identified multiple times, thereby improving the recognition efficiency of the document structure.
[0139] Since the document structure recognition model in the current technology does not make full use of text block information, the prediction accuracy needs to be improved. Therefore, this manual will use multiple sample documents, which include multiple sample text blocks. The sample documents carry the hierarchical information of each sample text block to train the initial recognition model, so that the initial recognition model has the function of using text block information to predict the hierarchical structure of the document to be recognized. Figure 4 , Figure 4 This is a flowchart of a target recognition model training method provided by an embodiment of this specification, which specifically includes the following steps.
[0140] Step 402: Acquire a sample set, wherein the sample set includes a plurality of sample documents, the sample documents include a plurality of sample text blocks, and the sample documents carry hierarchical information of each sample text block.
[0141] The hierarchical information of the sample text block may be manually annotated.
[0142] In practical applications, a large number of representative sample documents are collected. These documents should cover various situations that the model may encounter. For example, if the goal is to identify the hierarchical structure of legal documents, then these sample documents can include various types of legal documents, such as contracts, regulations, treaties, etc.; the collected sample documents are segmented into sample text blocks, such as titles, chapters, paragraphs, clauses, etc.; for each sample text block, it is necessary to annotate it with corresponding hierarchical information to indicate its hierarchical structure in the document, such as title, subtitle, body, etc. Labeling hierarchical information usually requires manual participation. Relevant personnel can add appropriate hierarchical information to each sample text block based on the semantics and logic of each sample text block; the sample documents and the hierarchical information of the sample text blocks in the sample documents are combined into a sample set to ensure that this sample set is as comprehensive and diverse as possible, covering various types of sample documents.
[0143] Step 404: Determine the text block information of each sample text block in the sample document.
[0144] In practical applications, after obtaining a sample document, the document structure recognition system can divide the content in the sample document into different sample text blocks according to preset rules, such as segment division; use OCR technology, such as optical character recognition, to extract the text in the sample document as editable text, and use the recognized text content as the text block information of the sample text block; and pre-process the text block information of the sample text block, such as removing noise, correcting recognition errors, and unifying text format.
[0145] Step 406: Input the sample document into the initial recognition model, and obtain the predicted hierarchical structure of each sample text block based on the text block information of each sample text block.
[0146] In practical applications, the initial recognition model predicts the hierarchical structure of each text block based on the text block information of each sample text block in the sample document.
[0147] Step 408: Train the initial recognition model based on the predicted hierarchical structure and hierarchical information to obtain a trained target recognition model.
[0148] In practical applications, the predicted hierarchical structure output by the initial recognition model is compared and evaluated with the hierarchical information in the sample documents. The accuracy, recall rate and other indicators of the prediction results are calculated to measure the performance of the initial recognition model. Based on the evaluation results, the initial recognition model is further optimized, such as adjusting the initial recognition model architecture, adjusting parameters, adding training data, etc. The adjustment process can be iterated multiple times until the initial recognition model achieves satisfactory performance as the target recognition model.
[0149] Step 410: Send the model parameters of the target recognition model to the terminal device.
[0150] Among them, the target recognition model is trained on the cloud-side device, and the model parameters of the target recognition model need to be sent to the terminal device before actual application.
[0151] In practical applications, after the target recognition model training is completed, the model parameters of the target recognition model are exported from the cloud computing environment, which usually involves serializing the model parameters and structure into a file so that it can be loaded and used on the terminal device. Then, in the application of the terminal device, the exported model file is loaded and initialized as an instance of the target recognition model, which usually involves loading the model parameters into the memory and constructing the structure and weights of the entire model.
[0152] In one or more embodiments of this specification, see Figure 5 As shown, Figure 5 This is a flowchart of another target recognition model training method provided by an embodiment of this specification, which specifically includes the following steps:
[0153] Training phase:
[0154] Step 502: Obtaining a sample text block. It is understandable that, in the training phase, the aforementioned step 402 includes the process of step 502, which will not be described in detail here.
[0155] Step 504: Determine the text block information of the sample text block. It is understood that during the training phase, the process of step 504 is the same as that of step 404, and will not be described in detail here.
[0156] Step 506: Obtaining hierarchical information of the sample text block. It is understandable that, during the training phase, the aforementioned step 402 includes the process of step 506, which will not be described in detail here.
[0157] Step 508: Training the initial recognition model. It is understood that in the training phase, the processes of the aforementioned steps 406 and 408 are the same as step 508, and will not be repeated here.
[0158] Step 510: Deploy the trained initial recognition model. The trained initial recognition model is deployed as the target recognition model. It is understood that during the training phase, the process of step 410 is the same as step 510 and will not be repeated here.
[0159] Testing phase:
[0160] Step 502: Obtain sample text blocks. It's important to note that the sample text blocks used in the training phase are different from those used in the testing phase. The testing phase is primarily used to evaluate the generalization ability of the target recognition model on unseen data. If the sample text blocks used in the training phase are used, the target recognition model may overfit and fail to make accurate predictions for new documents. Therefore, the sample text blocks used in the testing phase need to be independent of the sample text blocks used in the training phase.
[0161] Step 504: Determine the text block information of the sample text block. It is understood that during the testing phase, the process of step 504 is the same as that of step 404, and will not be described in detail here.
[0162] Step 512: Input text block information of the sample text block. The sample text block in the test phase is directly input into the target recognition model for testing.
[0163] By applying the solution of the embodiments of this specification, the initial recognition model is trained based on the sample document, the text block information of each text block in the sample document, and the hierarchical information of the sample text block, so that the initial recognition model has the function of predicting the hierarchical structure of the document to be recognized using the text block information, so that the initial recognition model can better understand the hierarchical relationship of each text block in the sample document, and then adjust the model parameters of the initial recognition model based on the predicted hierarchical structure and hierarchical information, so that the final target recognition model can be more accurate.
[0164] See also Figure 6 , Figure 6 This is a flowchart of a question-and-answer processing method provided by an embodiment of this specification, which specifically includes the following steps.
[0165] Step 602: Receive question information sent by the front-end user, and obtain the target document based on the question information.
[0166] The user may send question information through a front-end interface or an application programming interface (API).
[0167] In practical applications, the received question information is preprocessed, including word segmentation, stop word removal, stemming, etc., to extract the key content in the question information and prepare for subsequent information retrieval and processing; secondly, based on the preprocessed question information, information retrieval technology (such as inverted index, text similarity calculation, etc.) is used to utilize search engines or document libraries and other resources to find and obtain the target documents most relevant to the user's question information.
[0168] Step 604: Analyze the target document to determine the text block information of each text block in the target document.
[0169] In practical applications, after acquiring the target document, the document structure recognition system can divide the content of the target document into different text blocks according to preset rules, such as segment division; use text recognition technology, such as OCR, to extract the text in the target document as editable text, and use the recognized text content as text block information; and pre-process the text block information, such as removing noise, correcting recognition errors, and unifying text formats.
[0170] Step 606: Input the target document into the target recognition model, and identify the hierarchical structure of each text block in the target document based on the text block information of each text block, wherein the target recognition model is trained based on the sample document, the text block information of each sample text block in the sample document, and the hierarchical information of each sample text block, and the hierarchical information represents the hierarchical structure of each sample text block in the sample document.
[0171] In practical applications, before inputting the target document into the target recognition model, the target recognition model must be trained using sample documents, the text block information of each text block within the sample document, and the hierarchical information of each sample text block. This allows the target recognition model to identify the hierarchical structure of the input document based on the text block information of each text block. After the target recognition model is trained, the entire target document is input into the target recognition model. The target recognition model uses the knowledge gained from training and the text block information of each text block within the target document to identify the hierarchical structure of each text block within the target document. Based on the text block information, the target recognition model infers the logical order and hierarchical relationships of the text blocks within the target document and assigns an appropriate hierarchical structure to each text block, such as title, subtitle, body, paragraph, etc.
[0172] Step 608: Determine the target text block and the target level of the target text block based on the question information and the hierarchical structure of each text block in the target document.
[0173] The target text block is a text block that includes answer information or is related to question information.
[0174] In practical applications, directly searching for answer information in large or complex answer documents based on question information is inefficient. Therefore, in this specification, the question information is first matched with the text blocks of the target document to find the text block most relevant to the question as the target text block. The target level of the target text block is determined based on the hierarchical structure of the target document.
[0175] In one or more embodiments of this specification, taking the question "How to lose weight quickly?" and the hierarchical structure of the target document as follows: 1.1 Overview of healthy eating; 1.2 Principles of weight loss; 2. Healthy eating plan; 3. Exercise and weight loss; 4. Conclusion, as an example, the question is matched with the text blocks of the target document, and the similarity score between the question and each text block is calculated. Based on the similarity score, the text block "4. Conclusion" that is most relevant to the question is selected as the target text block, and the target level of the target text block is determined to be "4."
[0176] Step 610: Based on the target text block and the target level, output answer information to the front-end user.
[0177] In practical applications, if the target text block contains the answer information, the answer information in the target text block can be directly extracted and returned to the front-end user. If the target text block does not contain the answer information, further analysis is required to generate the answer information and return it to the front-end user.
[0178] One embodiment of the present specification inputs the entire target document and the text block information of each text block in the target document into the target recognition model, so that the target recognition model can better understand the hierarchical relationship of each text block in the target document, and fully combines the text block information of each text block in the target document to perform document structure recognition, thereby improving the accuracy of document structure recognition. At the same time, accurate document structure recognition is conducive to quickly locating the target level where the problem information is located, further improving the efficiency of locating and outputting answer information.
[0179] See also Figure 7 , Figure 7 This is a flowchart of a legal document query method provided by an embodiment of this specification, which specifically includes the following steps.
[0180] Step 702: Receive query information input by the front-end user, and obtain the target legal document based on the query information.
[0181] The query information may be a case description, legal issue information, or a specific legal provision.
[0182] In practical applications, based on query information, information retrieval technology (such as inverted index, text similarity calculation, etc.) is used to utilize search engines or legal document libraries and other resources to find and obtain the target legal documents that are most relevant to the user's query information.
[0183] In one or more embodiments of the present specification, information retrieval techniques, such as inverted indexing and text similarity calculation, can be used based on case description query information input by the user, and resources such as search engines or legal document libraries can be used to find the target legal document most relevant to the case description. Specifically, the keywords involved in the case description (such as traffic accidents, running red lights, legal liability, compensation regulations, etc.) can be searched in the legal document library using inverted indexing technology to find the legal document with the highest relevance to these keywords as the target legal document.
[0184] Step 704: Analyze the target legal document to determine the text block information of each text block in the target legal document.
[0185] In actual applications, the content of the target legal document can be divided into different text blocks according to preset rules, such as segmentation; each divided text block is analyzed to obtain text block information of each text block in the target legal document, which may include text content, position information (such as the start and end positions in the document), format information (such as font, color, size), etc.
[0186] Step 706: Input the target legal document into the target recognition model, and identify the hierarchical structure of each text block in the target legal document based on the text block information of each text block, wherein the target recognition model is trained based on the sample document, the text block information of each sample text block in the sample document, and the hierarchical information of each sample text block, and the hierarchical information represents the hierarchical structure of each sample text block in the sample document.
[0187] In practical applications, before inputting the target legal document into the target recognition model, the target recognition model must be trained using sample documents, the text block information of each text block within the sample document, and the hierarchical information of each sample text block. This allows the target recognition model to identify the hierarchical structure of the input document based on the text block information of each text block. After the target recognition model is trained, the entire target legal document is input into the target recognition model. The target recognition model utilizes the knowledge gained from training and, based on the text block information of each text block within the target legal document, identifies the hierarchical structure of each text block within the target legal document. Based on this text block information, the target recognition model infers the logical order and hierarchical relationships of the text blocks within the target legal document and assigns an appropriate hierarchical structure to each text block, such as title, subtitle, body, paragraph, etc.
[0188] Step 708: Determine the target text block based on the query information and the hierarchical structure of each text block in the target document.
[0189] In practical applications, directly searching for a target text block in a large or complex legal document based on query information is inefficient. Therefore, this specification first matches the query information with the text blocks in the legal document to identify the text block most relevant to the question as the target text block. The target level of the target text block is then determined based on the hierarchical structure of the target document.
[0190] Step 710: Based on the target text blocks and the hierarchical structure of each text block in the target legal document, output the query results to the front-end user.
[0191] In actual applications, the query result can be to output the hierarchical structure of the target legal document, highlight the target level, or directly output the text content of the target level.
[0192] In one or more embodiments of the present specification, refer to Figure 8(a), which is a schematic diagram of an interactive interface of a legal document query method provided by an embodiment of the present specification. In this interactive interface, there are an input box, an output box, a clear control, and a submit control. The input box is used for the front-end user to input query information or the target legal document, and to display the text content and text block information of each text block of the target legal document. The output box is used to display the query results or the hierarchical structure of the target legal document to the front-end user.
[0193] In Figure 8(a), the content in the input box includes the query "Output the hierarchical structure of the following target legal document" and the specific text content of the target legal document. It is understood that due to the limited size of the input box, the specific text content of the target legal document in the input box in Figure 8(a) does not represent the entire content of the target legal document. The user can adjust the portion of the target legal document displayed in the input box using the slider control.
[0194] In Figure 8(a), the content in the output box includes:
[0195] “【2】
[0196] #【3】
[0197] #【4】
[0198] #【5】
[0199] ##【6】
[0200] ###【7】
[0201] ####【8】
[0202] ####【9】
[0203] ####
[10]
[0204] ####
[11]
[0205] ####
[12]
[0206] ####
[13]
[0207] ####
[14]
[0208] #####
[15] ” In which, the text blocks are divided into paragraphs. The number of “#” represents the hierarchy of the text blocks. The fewer the number of “#”, the higher the hierarchy. For example, [2] is the highest hierarchy, [3], [4], and [5] are sub-levels of [2], [6] is a sub-level of [5], [7] is a sub-level of [6], [8], [9],
[10] ,
[11] ,
[12] ,
[13] , and
[14] are sub-levels of [7], and
[15] is a sub-level of
[14] . It can be understood that, due to the size of the output box, the hierarchical structure of the text blocks [2]-
[15] in the output box of Figure 8(a) is not the hierarchical structure of all the text blocks in the target legal document. The user can adjust the displayed part of the hierarchical structure of the target legal document in the output box through the slider control.
[0209] See FIG8( b ), which is a schematic diagram of the actual input and output of a target recognition model of a legal document query method provided in one embodiment of this specification.
[0210] In Figure 8(b), the content in the input box includes:
[0211] "
[49] ***************** [5, text, paragraph level 1 title, [138, 191], [888, 618], rgba [204, 51, 84, 1]]
[0212]
[50] ************************【5, text, paragraph secondary title, [328, 677], [707, 705], rgba[204, 153, 51, 1]】
[0213]
[51] **************【5, text, table, [143, 733], [901, 830], rgba[204, 51, 84, 1]】
[0214]
[52] ********************", where the content of 【】 is the text block information, indicating that the 49th text block is located on the 5th page of the target legal document, the text content includes text, the paragraph type is the first-level paragraph title, the coordinates of the upper left corner of the 49th text block are [138,191], the coordinates of the lower right corner are [888,618], and the color is rgba[204,51,84,1]; indicating that the 50th text block is located on the 5th page of the target legal document, the text content includes The paragraph type is a second-level paragraph title. The coordinates of the upper left corner of the 50th text block are [328,677], the coordinates of the lower right corner are [707,705], and the color is rgba[204,153,51,1]. This means that the 51st text block is located on page 5 of the target legal document. The text content includes text and tables. The coordinates of the upper left corner of the 51st text block are [143,733], the coordinates of the lower right corner are [901,830], and the color is rgba[204,51,84,1].
[0215] One embodiment of the present specification inputs the entire target legal document and the text block information of each text block in the target legal document into the target recognition model, so that the target recognition model can better understand the hierarchical relationship of each text block in the target legal document, and fully combine the text block information of each text block in the target legal document to perform document structure recognition, thereby improving the accuracy of document structure recognition. At the same time, accurate document structure recognition is conducive to quickly locating the target text block corresponding to the query information, thereby improving the efficiency of outputting the target text block.
[0216] Corresponding to the above method embodiment, this specification also provides a document structure recognition device embodiment, Figure 9 This is a schematic diagram of a document structure recognition device provided by an embodiment of this specification. Figure 9 As shown, the device includes:
[0217] The first acquisition module 902 is configured to acquire a document to be identified.
[0218] The first analysis module 904 is configured to determine each text block in the document to be recognized; analyze the document to be recognized, and determine text block information of each text block in the document to be recognized.
[0219] Optionally, the first analysis module 904 is further configured to perform text recognition on each text block to obtain the text content of each text block.
[0220] Optionally, the first analysis module 904 is further configured to input the document to be identified into the target recognition model, determine the semantic information of each text block based on the text content of each text block through the target recognition model, and identify the hierarchical structure of each text block in the document to be identified based on the semantic information of each text block.
[0221] Optionally, the first analysis module 904 is further configured to perform layout analysis on each text block to obtain block attribute information of each text block.
[0222] Optionally, the first analysis module 904 is further configured to identify the layout type of each text block; determine the position information of each text block based on the position of each text block in the document to be identified; perform text recognition on each text block to determine the text attribute information of each text block.
[0223] Optionally, the first analysis module 904 is further configured to determine the target document page number of the target text block in the document to be identified, wherein the target text block is any text block; and determine the position information of the target text block in the target document page corresponding to the target document page number.
[0224] The first recognition module 906 is configured to input the document to be recognized into a target recognition model, and identify the hierarchical structure of each text block in the document to be recognized based on the text block information of each text block, wherein the target recognition model is trained based on the sample document, the text block information of each sample text block in the sample document, and the hierarchical information of each sample text block, and the hierarchical information represents the hierarchical structure of each sample text block in the sample document.
[0225] Optionally, the document structure recognition device also includes a task indication module, which is configured to obtain task indication information, wherein the task indication information is used to indicate the task type and output format of the target recognition model; accordingly, the first recognition module 906 is further configured to input the document to be recognized and the task indication information into the target recognition model, and according to the task type, based on the text block information of each text block, identify the hierarchical structure of each text block in the document to be recognized, and output the hierarchical structure according to the output format.
[0226] Optionally, the document structure recognition device also includes a training module, which is configured to obtain a sample set, wherein the sample set includes multiple sample documents, the sample documents include multiple sample text blocks, and the sample documents carry hierarchical information of each sample text block; determine the text block information of each sample text block in the sample document; input the sample document into the initial recognition model, and obtain the predicted hierarchical structure of each sample text block based on the text block information of each sample text block; train the initial recognition model based on the predicted hierarchical structure and hierarchical information to obtain a target recognition model that has completed the training.
[0227] Optionally, the training module is further configured to obtain preset prompt content; according to the example information in the prompt content, the prompt information is constructed according to the sample document, the text block information of each sample text block and the hierarchical information; accordingly, the first recognition module 906 is further configured to input the prompt information into the initial recognition model, and based on the text block information of each sample text block, obtain the predicted hierarchical structure of each sample text block.
[0228] Optionally, the training module is further configured to receive evaluation messages sent by front-end users regarding the hierarchical structure; generate an inquiry guidance message based on the evaluation message, and send the inquiry guidance message to the front-end user; collect feedback information input by the front-end user based on the inquiry guidance message, and obtain an updated sample set based on the feedback information; and use the updated sample set to train the target recognition model.
[0229] Optionally, the document structure recognition device also includes a sending module, which is configured to send the hierarchical structure of each text block in the document to be recognized to the front-end user, wherein the hierarchical structure includes the block identifier of the text block; receive the target text block identifier feedback from the front-end user; based on the target text block identifier, determine the target text block corresponding to the target text block identifier from the document to be recognized; and send the text content of the target text block to the front-end user.
[0230] Through the document structure recognition device, the entire document to be recognized obtained by the first acquisition module 902 and the text block information of each text block in the document to be recognized obtained by the first analysis module are input into the target recognition model, so that the target recognition model can better understand the hierarchical relationship of each text block in the document to be recognized, and fully combine the text block information of each text block in the document to be recognized to perform document structure recognition, thereby improving the recognition accuracy of the first recognition module 906. At the same time, the entire document to be recognized is input into the target recognition model, so that the target model does not need to predict the relationship between adjacent paragraphs in the document to be recognized multiple times, thereby improving the recognition efficiency of the document structure.
[0231] The above is a schematic diagram of a document structure recognition device according to this embodiment. It should be noted that the technical solution of the document structure recognition device and the technical solution of the document structure recognition method described above are based on the same concept. For details not described in detail in the technical solution of the document structure recognition device, please refer to the description of the technical solution of the document structure recognition method described above.
[0232] Corresponding to the above method embodiment, this specification also provides an embodiment of a target recognition model training device, Figure 10 This is a schematic diagram of the structure of a target recognition model training device provided by an embodiment of this specification. Figure 10 As shown, the device is applied to cloud testing equipment, including:
[0233] The second acquisition module 1002 is configured to acquire a sample set, wherein the sample set includes a plurality of sample documents, the sample documents include a plurality of sample text blocks, and the sample documents carry hierarchical information of each sample text block.
[0234] The second analysis module 1004 is configured to determine text block information of each sample text block in the sample document.
[0235] The prediction module 1006 is configured to input the sample document and the text block information of each sample text block into the initial recognition model, and obtain the predicted hierarchical structure of each sample text block based on the text block information of each sample text block.
[0236] The training module 1008 is configured to train the initial recognition model based on the predicted hierarchical structure and hierarchical information to obtain a trained target recognition model.
[0237] The sending module 1010 is configured to send the model parameters of the target recognition model to the terminal device.
[0238] Through the target recognition model training device, based on the hierarchical information of the sample document and sample text block obtained by the second acquisition module 1002 and the text block information of each text block in the sample document analyzed by the second analysis module 1004, the initial recognition model of the prediction module 1006 can better understand the hierarchical relationship of each text block in the sample document, and then adjust the model parameters of the initial recognition model according to the predicted hierarchical structure and hierarchical information, so that the target recognition model obtained by the final training module 1008 can be more accurate.
[0239] The above is a schematic scheme of a target recognition model training device of this embodiment. It should be noted that the technical scheme of the target recognition model training device and the technical scheme of the target recognition model training method described above are based on the same concept. For details not described in detail in the technical scheme of the target recognition model training device, please refer to the description of the technical scheme of the target recognition model training method described above.
[0240] Corresponding to the above method embodiment, this specification also provides a question-answering processing device embodiment, Figure 11 This is a schematic diagram of the structure of a question-answering processing device provided by an embodiment of this specification. Figure 11 As shown, the device includes:
[0241] The third acquisition module 1102 is configured to receive question information sent by the front-end user and acquire the target document based on the question information.
[0242] The third analysis module 1104 is configured to analyze the target document and determine text block information of each text block in the target document.
[0243] The second recognition module 1106 is configured to input the target document into a target recognition model, and identify the hierarchical structure of each text block in the target document based on the text block information of each text block, wherein the target recognition model is trained based on the sample document, the text block information of each sample text block in the sample document, and the hierarchical information of each sample text block, and the hierarchical information represents the hierarchical structure of each sample text block in the sample document.
[0244] The first determination module 1108 is configured to determine the target text block and the target level of the target text block based on the question information and the hierarchical structure of each text block in the target document.
[0245] The first output module 1110 is configured to output answer information to a front-end user based on the target text block and the target level.
[0246] Through the question-and-answer processing device, the entire target document obtained by the third acquisition module 1102 and the text block information of each text block in the target document obtained by the third analysis module 1104 are input into the target recognition model, so that the target recognition model in the second recognition module 1106 can better understand the hierarchical relationship of each text block in the target document, and fully combine the text block information of each text block in the target document to perform document structure recognition, thereby improving the accuracy of document structure recognition. At the same time, accurate document structure recognition is beneficial for the first determination module 1108 to quickly locate the target level where the question information is located, further improving the efficiency of the first output module 1110 in locating and outputting answer information.
[0247] The above is a schematic diagram of a question-and-answer processing device according to this embodiment. It should be noted that the technical solution of this question-and-answer processing device and the technical solution of the aforementioned question-and-answer processing method are based on the same concept. For details not described in detail in the technical solution of the question-and-answer processing device, please refer to the description of the technical solution of the aforementioned question-and-answer processing method.
[0248] Corresponding to the above method embodiment, this specification also provides a legal document query device embodiment, Figure 12 This is a schematic diagram of the structure of a legal document query device provided by an embodiment of this specification. Figure 12 As shown, the device includes:
[0249] The fourth acquisition module 1202 is configured to receive query information input by a front-end user and acquire a target legal document based on the query information.
[0250] The fourth analysis module 1204 is configured to analyze the target legal document and determine text block information of each text block in the target legal document.
[0251] The third recognition module 1206 is configured to input the target legal document into a target recognition model, and identify the hierarchical structure of each text block in the target legal document based on the text block information of each text block, wherein the target recognition model is trained based on the sample document, the text block information of each sample text block in the sample document, and the hierarchical information of each sample text block, and the hierarchical information represents the hierarchical structure of each sample text block in the sample document.
[0252] The second determining module 1208 is configured to determine the target text block based on the query information and the hierarchical structure of each text block in the target document.
[0253] The second output module 1210 is configured to output the query results to the front-end user based on the target text blocks and the hierarchical structure of each text block in the target legal document.
[0254] The entire target legal document obtained by the fourth acquisition module 1202 and the text block information of each text block in the target legal document obtained by the fourth analysis module 1204 are input into the target recognition model through the legal document query device, so that the target recognition model of the third recognition module 1206 can better understand the hierarchical relationship of each text block in the target legal document, fully combine the text block information of each text block in the target legal document to perform document structure recognition, improve the accuracy of document structure recognition, and at the same time, accurately recognize the document structure, which is conducive to the second determination module 1208 to quickly locate the target text block corresponding to the query information, and improve the efficiency of the second output module 1210 in outputting the target text block.
[0255] The above is a schematic diagram of a legal document query device according to this embodiment. It should be noted that the technical solution of this legal document query device and the technical solution of the aforementioned legal document query method are based on the same concept. For details not described in detail in the technical solution of the legal document query device, please refer to the description of the technical solution of the aforementioned legal document query method.
[0256] Figure 1313 shows a block diagram of a computing device 1300 according to one embodiment of the present disclosure. Components of the computing device 1300 include, but are not limited to, a memory 1310 and a processor 1320. The processor 1320 is connected to the memory 1310 via a bus 1330, and a database 1350 is used to store data.
[0257] Computing device 1300 also includes an access device 1340 that enables computing device 1300 to communicate via one or more networks 1360. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 1340 may include one or more of any type of wired or wireless network interface (for example, a network interface card (NIC)), such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, and a Near Field Communication (NFC).
[0258] In one embodiment of the present specification, the above components of the computing device 1300 and Figure 13 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 13 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0259] Computing device 1300 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1300 can also be a mobile or stationary server.
[0260] Among them, the processor 1320 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above-mentioned document structure recognition or target recognition model training or question and answer processing or legal document query method.
[0261] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of this computing device and the technical scheme of the document structure recognition or target recognition model training or question-and-answer processing or legal document query method described above are based on the same concept. For details not described in detail in the technical scheme of the computing device, please refer to the description of the technical scheme of the document structure recognition or target recognition model training or question-and-answer processing or legal document query method described above.
[0262] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned document structure recognition or target recognition model training or question-answering processing or legal document query method.
[0263] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solution of the aforementioned document structure recognition, target recognition model training, question-and-answer processing, or legal document query method. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the aforementioned document structure recognition, target recognition model training, question-and-answer processing, or legal document query method.
[0264] An embodiment of this specification also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned document structure recognition or target recognition model training or question-answering processing or legal document query method.
[0265] The above is a schematic scheme of a computer program product of this embodiment. It should be noted that the technical scheme of this computer program product and the technical scheme of the aforementioned document structure recognition or target recognition model training or question-and-answer processing or legal document query method are based on the same concept. For details not described in detail in the technical scheme of the computer program product, please refer to the description of the technical scheme of the aforementioned document structure recognition or target recognition model training or question-and-answer processing or legal document query method.
[0266] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0267] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0268] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0269] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0270] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A document structure recognition method, comprising: Get the document to be identified; Analyzing the document to be identified to determine text block information of each text block in the document to be identified; The document to be identified is input into a target recognition model, and based on the text block information of each text block, the hierarchical structure of each text block in the document to be identified, wherein the target recognition model is trained based on sample documents, the text block information of each sample text block in the sample documents, and the hierarchical information of each sample text block, and the hierarchical information represents the hierarchical structure of each sample text block in the sample document.
2. The method according to claim 1, wherein the text block information includes text content; and the step of analyzing the document to be identified to determine the text block information of each text block in the document to be identified comprises: Determining each text block in the document to be recognized; Perform text recognition on each text block to obtain the text content of each text block.
3. The method according to claim 2, wherein inputting the document to be recognized into a target recognition model and identifying the hierarchical structure of each text block in the document to be recognized based on the text block information of each text block comprises: The document to be identified is input into a target recognition model, and the target recognition model determines the semantic information of each text block based on the text content of each text block, and identifies the hierarchical structure of each text block in the document to be identified based on the semantic information of each text block.
4. The method according to claim 2, wherein the text block information further includes block attribute information; after determining each text block in the document to be identified, the method further includes: Performing layout analysis on each text block to obtain block attribute information of each text block; The step of identifying the hierarchical structure of each text block in the document to be identified based on the text block information of each text block includes: Based on the text content of each text block and the block attribute information of each text block, the hierarchical structure of each text block in the document to be identified is identified.
5. The method according to claim 4, wherein the block attribute information includes at least one of layout type, position information, and text attribute information; and performing layout analysis on each text block to obtain the block attribute information of each text block comprises at least one of the following steps: Identifying the layout type of each text block; Determining position information of each text block based on the position of each text block in the document to be recognized; Performing text recognition on each text block to determine text attribute information of each text block.
6. The method according to claim 5, wherein determining the position information of each text block based on the position of each text block in the document to be recognized comprises: Determining a target document page number of a target text block in the document to be recognized, wherein the target text block is any text block; In the target document page corresponding to the target document page number, position information of the target text block is determined.
7. The method according to any one of claims 1 to 6, further comprising: before inputting the document to be recognized into a target recognition model and identifying the hierarchical structure of each text block in the document to be recognized based on the text block information of each text block: Acquiring task indication information, wherein the task indication information is used to indicate the task type and output format of the target recognition model; Inputting the document to be recognized into a target recognition model, and identifying the hierarchical structure of each text block in the document to be recognized based on the text block information of each text block, includes: The document to be identified and the task instruction information are input into the target recognition model, and according to the task type and based on the text block information of each text block, the hierarchical structure of each text block in the document to be identified is identified, and the hierarchical structure is output according to the output format.
8. The method according to any one of claims 1 to 6, further comprising: before inputting the document to be recognized into a target recognition model and identifying the hierarchical structure of each text block in the document to be recognized based on the text block information of each text block: Acquire a sample set, wherein the sample set includes a plurality of sample documents, the sample documents include a plurality of sample text blocks, and the sample documents carry hierarchical information of each sample text block; Determining text block information of each sample text block in the sample document; Inputting the sample document into an initial recognition model, and obtaining a predicted hierarchical structure of each sample text block based on text block information of each sample text block; The initial recognition model is trained based on the predicted hierarchical structure and the hierarchical information to obtain a trained target recognition model.
9. The method according to claim 8, before inputting the sample document into an initial recognition model and obtaining a predicted hierarchical structure of each sample text block based on text block information of each sample text block, further comprising: Get the preset prompt content; According to the example information in the prompt content, construct prompt information according to the sample document, the text block information and the hierarchical information of each sample text block; Inputting the sample document into an initial recognition model and obtaining a predicted hierarchical structure of each sample text block based on text block information of each sample text block includes: The prompt information is input into an initial recognition model, and based on the text block information of each sample text block, a predicted hierarchical structure of each sample text block is obtained.
10. The method according to claim 8, further comprising, after inputting the document to be recognized into a target recognition model and identifying the hierarchical structure of each text block in the document to be recognized based on the text block information of each text block: Sending the hierarchical structure of each text block in the document to be recognized to a front-end user; receiving an evaluation message for the hierarchical structure sent by the front-end user; Generate an inquiry guidance message based on the evaluation message, and send the inquiry guidance message to the front-end user; Collecting feedback information input by the front-end user based on the inquiry guidance message, and obtaining an updated sample set based on the feedback information; The target recognition model is trained using the updated sample set.
11. The method according to claim 1, after inputting the document to be recognized into a target recognition model and identifying the hierarchical structure of each text block in the document to be recognized based on the text block information of each text block, further comprising: Sending the hierarchical structure of each text block in the document to be recognized to a front-end user, wherein the hierarchical structure includes the block identifiers of the text blocks; Receive the target text block identifier fed back by the front-end user; Based on the target text block identifier, determining a target text block corresponding to the target text block identifier from the document to be recognized; The text content of the target text block is sent to the front-end user.
12. A target recognition model training method, applied to a cloud-side device, comprising: Acquire a sample set, wherein the sample set includes a plurality of sample documents, the sample documents include a plurality of sample text blocks, and the sample documents carry hierarchical information of each sample text block; Determining text block information of each sample text block in the sample document; Inputting the sample document into an initial recognition model, and obtaining a predicted hierarchical structure of each sample text block based on text block information of each sample text block; Training the initial recognition model based on the predicted hierarchical structure and the hierarchical information to obtain a trained target recognition model; The model parameters of the target recognition model are sent to the terminal device.
13. A question-answering processing method, comprising: Receive question information sent by the front-end user, and obtain the target document based on the question information; Analyzing the target document to determine text block information of each text block in the target document; Inputting the target document into a target recognition model, and identifying the hierarchical structure of each text block in the target document based on the text block information of each text block, wherein the target recognition model is trained based on sample documents, the text block information of each sample text block in the sample document, and the hierarchical information of each sample text block, wherein the hierarchical information represents the hierarchical structure of each sample text block in the sample document; Determining a target text block and a target hierarchy of the target text block based on the question information and the hierarchical structure of each text block in the target document; Based on the target text block and the target level, answer information is output to the front-end user.
14. A legal document query method, comprising: Receive query information input by a front-end user, and obtain target legal documents based on the query information; Analyzing the target legal document to determine text block information of each text block in the target legal document; Inputting the target legal document into a target recognition model, and identifying, based on the text block information of each text block, a hierarchical structure of each text block in the target legal document, wherein the target recognition model is trained based on sample documents, the text block information of each sample text block in the sample document, and hierarchical information of each sample text block, wherein the hierarchical information represents the hierarchical structure of each sample text block in the sample document; Determining a target text block based on the query information and the hierarchical structure of each text block in the target document; Based on the target text block and the hierarchical structure of each text block in the target legal document, a query result is output to the front-end user.
15. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the document structure recognition method described in any one of claims 1 to 11, the target recognition model training method described in claim 12, the question and answer processing method described in claim 13, or the legal document query method described in claim 14 are implemented.
16. A computer-readable storage medium storing a computer program / instruction, which, when executed by a processor, implements the steps of the document structure recognition method described in any one of claims 1 to 11, the target recognition model training method described in claim 12, the question-answering processing method described in claim 13, or the legal document query method described in claim 14.
17. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the document structure recognition method described in any one of claims 1 to 11, the target recognition model training method described in claim 12, the question-answering processing method described in claim 13, or the legal document query method described in claim 14.