Document Recognition Method, Intelligent Interaction Method, and Related Devices, Equipment, and Media

By identifying and correcting the title hierarchy in the document, the coherence problem of title modeling in the cross-page document recognition is solved, and the effect of accurately distinguishing the title hierarchy is achieved.

CN119990112BActive Publication Date: 2025-07-11ANHUI IFLYTEK INTELLIGENT SYST
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510454667.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-11
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The prior art lacks coherent modeling of cross-page content in document recognition, making it difficult to accurately distinguish the hierarchical relationships of each title in the document, making it difficult to reconstruct the logical structure.

Method used

By identifying layout elements in the document, especially the title, generating a preliminary title sequence, correcting the title level, performing verification, and iterating the correction until the conditions are met, ensuring the consistency of title modeling.

Benefits of technology

In the cross-page scenario, the consistency of title modeling is improved, the hierarchical relationships of each title in the document are accurately distinguished, and the accuracy of document recognition is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990112B_ABST
    Figure CN119990112B_ABST
Patent Text Reader

Abstract

The present application discloses a document recognition method, an intelligent interaction method, and related devices, equipment, and media. Among them, the document recognition method includes: performing recognition based on a document to be recognized to obtain layout elements in the document to be recognized and recognition results of the layout elements; analyzing based on the recognition results of the titles to obtain a first title sequence; correcting the title levels of the first title sequence to obtain a second title sequence; performing verification based on the second title sequence to obtain a verification result; wherein, the verification result indicates whether the second title sequence is correct; in response to the verification result indicating that the second title sequence is incorrect, selecting the second title sequence as the new first title sequence, and returning to the step of correcting the title levels of the first title sequence to obtain the second title sequence for iteration until the end condition is met. The above solution can improve the coherence of title modeling during document recognition to distinguish the hierarchical relationships of each title in the document, especially in the cross-page scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of natural language processing, and particularly to a document recognition method, an intelligent interaction method, and related devices, equipment, and media. Background Art

[0002] Document recognition is of great significance in many scenarios. For example, by recognizing and extracting structured information in a document, it can ensure that in the retrieval task of retrieval-augmented generation in the vertical domain, the document structure is correctly recognized, so that key knowledge can be accurately and completely extracted from it.

[0003] Among them, the correct recognition of the document title is an important part of document recognition. However, the prior art usually analyzes documents on a single-page basis, lacking the modeling of the coherence of cross-page content. As a result, it can only recognize the existence of the title, but it is difficult to distinguish its hierarchical relationship, and thus it is difficult to reconstruct the logical structure of the document. In view of this, how to improve the coherence of title modeling during document recognition to distinguish the hierarchical relationships of each title in the document, especially in the cross-page scenario, has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem to be solved by the present application is to provide a document recognition method, an intelligent interaction method, and related devices, equipment, and media, which can improve the coherence of title modeling during document recognition to distinguish the hierarchical relationships of each title in the document, especially in the cross-page scenario.

[0005] To solve the above technical problem, in the first aspect of the present application, a document recognition method is provided, including: performing recognition based on a document to be recognized to obtain layout elements and recognition results of the layout elements in the document to be recognized; where the layout elements at least include titles; analyzing based on the recognition results of the titles to obtain a first title sequence; correcting the title levels of the first title sequence to obtain a second title sequence; verifying based on the second title sequence to obtain a verification result; where the verification result indicates whether the second title sequence is correct; in response to the verification result indicating that the second title sequence is incorrect, selecting the second title sequence as the new first title sequence, and returning to the step of correcting the title levels of the first title sequence to obtain the second title sequence for iteration until the end condition is met.

[0006] To solve the above technical problems, a second aspect of the present application provides an intelligent interaction method, including: obtaining a first statement to be responded to, and obtaining a reference document required to respond to the first statement; performing identification based on the reference document to obtain the structured information of the reference document; wherein the structured information at least includes the hierarchical title sequence of the reference document, and the structured information is obtained by the document identification method in the first aspect above; performing a search based on the structured information to obtain a reference segment for responding to the first statement; and generating a second statement for responding to the first statement based on the reference segment.

[0007] To solve the above technical problems, a third aspect of the present application provides a document identification device, including: an identification module, an analysis module, a correction module, a verification module, and a loop module. The identification module is configured to perform identification based on a document to be identified to obtain the layout elements and the identification results of the layout elements in the document to be identified; wherein the layout elements at least include titles. The analysis module is configured to perform analysis based on the identification results of the titles to obtain a first title sequence. The correction module is configured to correct the title levels of the first title sequence to obtain a second title sequence. The verification module is configured to perform verification based on the second title sequence to obtain a verification result; wherein the verification result indicates whether the second title sequence is correct. The loop module is configured to, in response to the verification result indicating that the second title sequence is incorrect, select the second title sequence as the new first title sequence, and return to the step of correcting the title levels of the first title sequence to obtain the second title sequence for iteration until the end condition is met.

[0008] To solve the above technical problems, a fourth aspect of the present application provides an intelligent interaction device, including: an acquisition module, an identification module, a search module, and a generation module. The acquisition module is configured to obtain a first statement to be responded to, and obtain a reference document required to respond to the first statement. The identification module is configured to perform identification based on the reference document to obtain the structured information of the reference document; wherein the structured information at least includes the hierarchical title sequence of the reference document, and the structured information is obtained by the document identification device in the third aspect above. The search module is configured to perform a search based on the structured information to obtain a reference segment for responding to the first statement. The generation module is configured to generate a second statement for responding to the first statement based on the reference segment.

[0009] To solve the above technical problems, a fifth aspect of the present application provides an electronic device, at least including a memory and a processor that are coupled to each other. The memory stores at least program instructions, and the processor is configured to execute the program instructions to implement the document identification method in the first aspect above, or implement the intelligent interaction method in the second aspect above.

[0010] To solve the above technical problems, a sixth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor, and the program instructions are used to implement the document recognition method in the first aspect above, or implement the intelligent interaction method in the second aspect above.

[0011] In the above solution, recognition is performed based on the document to be recognized to obtain the layout elements and the recognition results of the layout elements in the document to be recognized, and the layout elements at least include a title. Analysis is performed based on the recognition result of the title to obtain a first title sequence, and the title hierarchy of the first title sequence is corrected to obtain a second title sequence. Verification is performed based on the second title sequence to obtain a verification result, and the verification result indicates whether the second title sequence is correct. Then, in response to the verification result indicating that the second title sequence is incorrect, the second title sequence is selected as the new first title sequence, and the step of returning to correct the title hierarchy of the first title sequence to obtain the second title sequence is iterated until the end condition is met. Therefore, by correcting the title hierarchy, verifying the title sequence, and returning to correct the title hierarchy for iteration in the case of incorrect verification, even in a multi-page scenario, the coherence of title modeling can be ensured as much as possible. Therefore, the coherence of title modeling during document recognition can be improved to distinguish the hierarchical relationships of each title in the document, especially in a multi-page scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a flowchart of an embodiment of the document recognition method of the present application;

[0013] Figure 2a is a schematic diagram of the process of an embodiment of confirming the preliminary hierarchy of the title in the present application;

[0014] Figure 2b is a schematic diagram of the process of an embodiment of resetting the title height in the present application;

[0015] Figure 2c is a schematic diagram of the process of an embodiment of refining the title hierarchy in the present application;

[0016] Figure 2d is a schematic diagram of the process of an embodiment of the document recognition method of the present application;

[0017] Figure 3 is a flowchart of an embodiment of the intelligent interaction method of the present application;

[0018] Figure 4 is a schematic diagram of the framework of an embodiment of the document recognition device of the present application;

[0019] Figure 5 is a schematic diagram of the framework of an embodiment of the intelligent interaction device of the present application;

[0020] Figure 6It is a schematic diagram of the framework of an embodiment of the electronic device of the present application;

[0021] Figure 7 It is a schematic diagram of the framework of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners

[0022] The following combines the description of the drawings of the specification to elaborate on the solutions of the embodiments of the present application in detail.

[0023] In the following description, specific details such as specific system structures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0024] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the fragment " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more than two.

[0025] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the document recognition method of the present application. Specifically, it may include the following steps:

[0026] Step S11: Perform recognition based on the document to be recognized to obtain the layout elements and the recognition results of the layout elements in the document to be recognized.

[0027] In the embodiments of the present disclosure, the layout elements may at least include a title. It should be noted that the layout element "title" may specifically include various levels of titles, such as main titles, first-level titles, second-level titles, third-level titles, etc., and will not be exemplified one by one here. In addition, the recognition result of the layout element "title" may specifically include the title area (such as the rectangular box coordinates surrounding the title text) and the title text (such as "Chapter 1 XXXX", "1.1 XXX", "1.1.1 XXXX", etc.). Of course, in actual application processes, the layout elements may also include other types other than titles, such as paragraphs, tables, pictures, etc., and will not be exemplified one by one here.

[0028] In an implementation scenario, the file format of the document to be recognized may include, but is not limited to: PDF, doc, epub, etc., and the file format of the document to be recognized is not limited here.

[0029] In an implementation scenario, in order to obtain the layout elements and their recognition results in the document to be recognized, the document to be recognized can be first split to obtain each document page in the document to be recognized, and then layout analysis and text recognition are respectively performed on each document page. By combining the two types of information, namely the analysis results of layout analysis and the recognition results of text recognition, the layout elements and their recognition results in the document page can be obtained. Finally, the layout elements and their recognition results in each document page are sequentially formed into a set, which can be used as the layout elements and their recognition results in the document to be recognized.

[0030] In a specific implementation scenario, as a possible implementation example, a layout analysis model such as YOLO can be used to perform layout analysis on the document page to obtain each layout element (such as title, paragraph, table, picture, etc.) and its coordinate information in the document page, which is the analysis result of the layout analysis of the document page.

[0031] In a specific implementation scenario, as a possible implementation example, a text recognition tool such as Tesseract OCR can be used to perform text recognition on the document page to obtain the text string within the text block and the coordinate information of the text block in the document page, which is the text recognition result of the text recognition of the document page.

[0032] In a specific implementation scenario, after obtaining the analysis results of layout analysis and the recognition results of text recognition, the above two types of information can be merged. Specifically, for each layout element in the document page, the coordinate information of the layout element can be matched with the coordinate information of each text block for overlap. If the overlap between the text block and the layout element is higher than the set threshold, the text block is attributed to the layout element. For example, if the overlap between the coordinate information of the text block "Chapter 1 XXXX" and the layout element "Title" is higher than the set threshold, the text block "Chapter 1 XXXX" can be attributed to the layout element "Title". Of course, the above example is only one possible example in actual application, and other possible situations will not be given one by one here. In addition, in the actual application process, there are also situations where the text block cannot be classified into any layout element based on the coordinate information of the layout element and the coordinate information of the text block. In this case, the text block can be temporarily classified into the layout element "text" (for example, the text block "1.1.2" cannot be classified into any layout element based on its coordinate information and the coordinate information of the layout element, so it can be temporarily classified into the layout element "text"), and it can be further classified later (for example, it can be classified into the layout element "title" according to the regular expression later). For details, please refer to the subsequent related descriptions, which will not be repeated here. Of course, in the actual application process, there may also be situations where any text block cannot be classified into a certain layout element based on the coordinate information of the layout element and the coordinate information of the text block. In this case, the layout element and its coordinate information can be temporarily retained for subsequent analysis of possible empty tables or empty titles. At this point, for the document to be identified, each layout element, the coordinate information of the layout element, and the text content of the layout element on each document page can be obtained.

[0033] Step S12: Analyze based on the recognition result of the title to obtain a first title sequence.

[0034] Specifically, as mentioned above, the level of each layout element "title" (i.e., the level of title) cannot be determined through the above identification, so we can first analyze based on the identification result of the layout element "title" to obtain the first title sequence to obtain the hierarchical title sequence. In other words, the first title sequence contains each layout element "title" organized in a hierarchical structure, such as the organization in order: "Chapter 1 XXXX", "1.1 XXXX", "1.2 XXXX", "Chapter 2 XXXX", "2.1XXXX", "2.2 XXXX", etc. The specific content of the first title sequence is not limited here. Please refer to Figure 2a , Figure 2a This is a schematic diagram of the process of confirming the preliminary level of the title in this application. Figure 2aAs shown, for each layout element "title" in the document to be identified, the title height can be reset first, and then the abnormal titles can be identified based on this (the abnormal titles will be reset to the layout element "text"). Then, for the remaining layout elements "title" (that is, the layout elements "title" without abnormalities), they can be clustered based on their title heights to obtain various cluster sets, and each cluster set can contain at least one layout element "title" with related title heights (for example, the title heights are roughly the same). Then, the cluster sets are sorted according to the title heights to generate a preliminary title hierarchy. It should be noted that the layout element blocks (such as rectangular boxes, etc., in short, "title blocks") represented by the coordinate information obtained through the aforementioned layout analysis may be slightly larger than the actual size, which affects the accuracy of judging the title height. Please refer to Figure 2b , Figure 2b 1 is a schematic diagram of the process of resetting the height of the title of this application. Figure 2b As shown, as a possible implementation example, the document page can be converted into a grayscale image first, and the grayscale image can be binarized to enhance the contrast of the text area. After that, the layout element block of the layout element "title" can be horizontally scanned to obtain the actual title height of the layout element "title", and the title height reset can be completed. On this basis, before high-level clustering, in order to avoid the influence of title height errors caused by layout analysis errors on subsequent hierarchical distinctions as much as possible, abnormal titles with large height differences from regular titles can be processed first. Specifically, on the one hand, the mean and variance of the title height can be calculated, based on which the height range of regular titles can be calculated, such as [μ-3σ, μ+3σ], where μ represents the mean value and σ represents the variance. If the title height is not within this height range, it can be determined as an abnormal title, otherwise it can be determined as a regular title; on the other hand, the average number of words in the regular title can be combined to screen abnormal titles. If the number of words in the title deviates significantly from the average number of words, it can be determined as an abnormal title, otherwise it can be determined as a regular title. Finally, the title heights can be clustered based on a clustering algorithm such as K-Means (eg, the K value can be determined using the "elbow method") to obtain a preliminary title hierarchy. As a possible implementation example, the preliminary title hierarchy can be used as the first title sequence.

[0035] In addition, as mentioned above, in order to further refine the title level, regular expressions can also be used to optimize the preliminary title level. Figure 2c , Figure 2c Schematic diagram of the process of refining the title level in this application. Figure 2cAs shown, common title patterns can be matched based on regular expressions, such as "Chapter [One Two Three Four]", "[1-9]", "(One)", etc., to identify and extract the structural features of the title. Then, according to the matching results, the preliminary title hierarchy is corrected to ensure the rationality of the title hierarchy. In addition, the context semantics of the title can be combined for analysis to evaluate the relevance between the title and the text content. If it is found that the hierarchical division does not match the semantic logic (for example, the text content is long but misjudged as a low-level title), the relevant title hierarchy can be corrected, so as to ensure that the hierarchical division conforms to the document logic.

[0036] It should be noted that the above examples are only several possible examples of obtaining the first title sequence based on the recognition results of the layout element "title". In actual application, it can be selected according to the actual situation. In addition, it does not exclude the use of other methods to obtain the first title sequence, and other possible implementation methods are not listed one by one here.

[0037] Step S13: Correct the title hierarchy of the first title sequence to obtain the second title sequence.

[0038] Specifically, a sequence verification model can be pre-trained to correct the title hierarchy of the first title sequence based on the sequence verification model to obtain the second title sequence.

[0039] In an implementation scenario, to train the sequence verification model, a sample title sequence can be obtained first. The sample title sequence can at least include an incorrect title sequence, and the sample title sequence can also be provided with annotation information, and the annotation information can at least include the modification method of the incorrect title sequence. Specifically, the correct title sequence can be obtained by exporting the title based on the sample document first, and then the correct title sequence can be disrupted to obtain the sample title sequence and the modification method. Exemplarily, taking the correct title sequence {"Chapter One XXXX", "1.1 XXXX", "1.2 XXXX", "Chapter Two XXXX", "2.1 XXXX", "2.2 XXXX"} as an example, it can be disrupted to obtain the incorrect title sequence {"Chapter One XXXX", "1.1 XXXX", "Chapter Two XXXX", "1.2 XXXX", "2.1 XXXX", "2.2 XXXX"} and its modification method {swap the order of "Chapter Two XXXX" and "1.2 XXXX"}. Of course, the above example is only one possible example of obtaining the sample title sequence in the actual application process. Other acquisition methods are not limited here and are not listed one by one.

[0040] In one implementation scenario, after obtaining the sample title sequence, the annotation information marked by the sample title sequence can be used as the training objective, and a sequence verification model can be trained based on the sample title sequence. Specifically, the sample title sequence can be processed based on the sequence verification model to obtain prediction information, and the prediction information includes whether the sample title sequence is an incorrect title sequence and the modification method in the case where the sample title sequence is predicted to be an incorrect title sequence. Then, based on the difference between the prediction information of the sample title sequence and the annotation information, the network parameters of the sequence verification model can be adjusted. Exemplarily, the difference between the prediction information and the annotation information can be measured based on a loss function such as cross-entropy to obtain the training loss of the sequence verification model, and the network parameters of the sequence verification model can be adjusted based on the training loss. Through the above method, the sequence verification model can be forced to learn the semantic relationships between title levels and common logical rules, such as the order of increasing title levels and the context dependence between levels, so that the sequence verification model can detect errors in the title level sequence and provide optimization suggestions.

[0041] In one implementation scenario, after training the sequence verification model, the title level of the first title sequence can be corrected based on the sequence verification model to obtain a second title sequence. Specifically, the first title sequence can be processed with the sequence verification model to predict the modification method of the first title sequence, and then the title level of the first title sequence can be corrected based on the modification method of the first title sequence to obtain a second title sequence. For example, if the predicted modification method of the first title sequence is to exchange the order of several titles in the first title sequence, the order of these titles in the first title sequence can be exchanged according to this method to obtain a second title sequence. Of course, the above example is only a possible example of using the sequence verification model to correct the title level in the actual application process, and other possible situations will not be exemplified one by one here.

[0042] Step S14: Verify based on the second title sequence to obtain a verification result.

[0043] In the embodiments of the present disclosure, the verification result indicates whether the second title sequence is correct. Specifically, as described above, the title level of the first title sequence can be corrected using the sequence verification model, and then the second title sequence can also be verified based on the sequence verification model to obtain a prediction mark indicating whether the second title sequence is incorrect as the verification result. For example, the aforementioned sequence verification model can be trained using a sample title sequence with annotation information, so that the sequence verification model not only has the ability to correct the title level but also has the ability to identify whether the sequence is incorrect. Therefore, the second title sequence can be directly verified using the aforementioned sequence verification model to obtain the verification result of whether the second title sequence is incorrect.

[0044] Step S15: In response to the verification result indicating that the second title sequence is incorrect, select the second title sequence as the new first title sequence, and iterate the step of returning to correct the title hierarchy of the first title sequence to obtain the second title sequence until the end condition is met.

[0045] In an implementation scenario, if the verification result indicates that the second title sequence is incorrect, the second title sequence can be selected as the new first title sequence, and the step of returning to correct the title hierarchy of the first title sequence to obtain the second title sequence is iterated until the end condition is met. That is to say, if the verification result indicates that the second title sequence is incorrect, the title hierarchy of the second title sequence can be continuously corrected, and then verified after correction, and so on in a loop until the end condition is met. It should be noted that the end condition can be set according to actual applications. For example, it can be set to iterate a preset number of times (such as 5 times, 10 times, etc.); or it can be set that the verification result of the latest second title sequence indicates that the latest second title sequence is correct. Of course, the above examples are only several possible examples of the second title sequence, and other possible situations are not limited here, nor will they be exemplified one by one.

[0046] In another implementation scenario, different from the foregoing situation, if the verification result indicates that the second title sequence is correct, the latest second title sequence can be used as the final hierarchical title sequence of the document to be recognized.

[0047] In an implementation scenario, the recognition result may also include the element type and the circumscribed block of the layout element. Then, based on the element type of the layout element to which the circumscribed block crossing the layout center line belongs, the column layout of the document to be recognized can be determined, and the column layout can be any one of a single-column layout and a multi-column layout. Then, based on the division method matching the column layout, the reading order of each layout element in the document to be recognized is determined. Therefore, an appropriate division method can be adaptively selected according to the document to be recognized to determine the reading order of the layout elements, which helps to improve the adaptability to various documents.

[0048] In a specific implementation scenario, to determine the column layout, it is possible to first count based on the element types of the external blocks that cross the center line of the layout, obtain the statistical results, and the statistical results include the proportion of the quantities of various element types. For example, after statistics, among the various element types that cross the center line of the layout: the proportion of the quantity of the layout element "title", the proportion of the quantity of the layout element "paragraph", the proportion of the quantity of the layout element "table", the proportion of the quantity of the layout element "picture", etc., and no more examples will be given here. In response to the statistical results indicating that the element types with the proportion of the quantity exceeding the proportion threshold only involve the target type, determine the column layout as a multi-column layout. It should be noted that the proportion threshold can be set according to the actual application. For example, the proportion threshold can be set to 40%, 50%, etc., and the specific value of the proportion threshold is not limited here. In addition, the target type can include but is not limited to at least one of the title and the picture. In addition, in response to the statistical results indicating that the element types with the proportion of the quantity exceeding the proportion threshold are more than the target type, the column layout can be determined as a single-column layout. For example, after statistics, there are also a large number of layout elements "paragraphs" that cross the center line of the layout, and this can determine the column layout as a single-column layout.

[0049] In a specific implementation scenario, after determining the column layout and before determining the reading order, it is possible to first sort the external blocks of each layout element according to the writing norms and the position information of the external blocks of each layout element. For example, it is possible to sort the external blocks of each layout element according to the writing norms of "from top to bottom, from left to right", and according to the "upper edge coordinate (y value)" and "left edge coordinate (x value)" (i.e., the upper left corner coordinate) of the external block, so as to lay a foundation for subsequent more refined division.

[0050] In a specific implementation scenario, when the field layout is a single-column layout, it is possible to determine whether to split an external block into multiple external blocks based on the gap between adjacent projection segments after projecting the external block along the horizontal direction, and determine whether to split the external block into multiple external blocks based on the gap between adjacent projection segments after projecting the external block along the vertical direction. Moreover, the layout elements correspond one-to-one with the external blocks. After each split, each layout element is re-sorted based on the position information of the external block that corresponds one-to-one with it respectively, and then the next split is performed. As a possible implementation example, taking the projection of the external block along the horizontal direction as an example, if the gap between adjacent projection segments after projection is too large (e.g., greater than the set threshold), it can be determined that the external block should be regarded as different rows or different horizontal bands, so the external block can be split into multiple external blocks along this gap; similarly, taking the projection of the external block along the vertical direction as an example, if the gap between adjacent projection segments after projection is too large (e.g., greater than the set threshold), it can be determined that the external block should be regarded as different columns or different vertical bands, so the external block can be split into multiple external blocks along this gap. Splitting the external block in this way can gradually split it into rows, columns, and smaller granularities by means of projection segmentation even when there is partial overlap of the external block, thus effectively avoiding the problems of omission or misclassification, and further greatly improving the recognition accuracy of the document layout.

[0051] In a specific implementation scenario, when the layout of the fields is a multi-column layout, the initial sorting can be performed based on the position information of the external blocks according to the writing plan to obtain an initial block sequence. For specific details, please refer to the relevant descriptions above and will not be elaborated here. Based on this, each external block in the initial block sequence can be traversed in turn. Based on the relative positions between the position information of the two vertices on both sides of the external block and several reference lines, the external block can be incorporated into one of the column subsequences, and the several reference lines include at least one of the center line of the layout, the quarter line of the layout, and the three-quarter line of the layout. The column subsequences at least include a left column subsequence and a right column subsequence. For the convenience of description, the position information of the two vertices on both sides of the external block can be respectively denoted as bbox[0] (representing the position information of the left vertex) and bbox[2] (representing the position information of the right vertex). Then, if it can be determined according to the position information that the left vertex is on the left side of the quarter line of the layout (i.e., bbox[0] < w / 4) and the right vertex is on the left side of the three-quarter line of the layout (i.e., bbox[2] < 3*w / 4), the external block can be incorporated into the left column subsequence; similarly, if it can be determined according to the position information that the left vertex is on the right side of the quarter line of the layout (i.e., bbox[0] > w / 4) and the right vertex is on the right side of the center line of the layout (i.e., bbox[2] > w / 2), the external block can be incorporated into the right column subsequence. Here, w represents the width of the entire page. Of course, the above examples are only one possible example in the actual application process, and other possible incorporation methods will not be exemplified one by one here. On this basis, each column subsequence can be spliced according to the writing specifications to obtain the reading order of each layout element. For example, each column subsequence can be spliced in the order of "from top to bottom, from left to right" to obtain the reading order of each layout element. In addition, during the traversal process, if an external block straddles the center line of the layout or there is a large gap between its upper edge and the lower edge of the previous external block, it means that this external block is more likely to be a new line or an independent block. In response to this situation, after collecting the left column subsequence and the right column subsequence, this external block can be inserted into the final column sequence (i.e., the column sequence after splicing) to obtain the reading order of each layout element.

[0052] In an implementation scenario, the recognition result of a layout element may further include the element type and element content of the layout element. To improve the semantic integrity of the element content, especially the semantic integrity of cross-page layout elements, a layout element located at the bottom of the first page can be selected as the first element, and a layout element located at the top of the second page and having the same element type as the first element can be selected as the second element. The first page and the second page are adjacent pages in the document to be recognized. Then, based on the element content of the first element and the element content of the second element, an analysis result indicating whether the first element and the second element are semantically related is obtained. Furthermore, based on the analysis result, it can be determined whether to merge the first element and the second element into one layout element. In this way, layout elements that span multiple pages but actually belong to the same semantic segment can be merged into a complete layout element to improve the semantic integrity of the layout elements.

[0053] In a specific implementation scenario, the analysis result can be obtained by a semantic analysis model analyzing the element content of the first element and the element content of the second element. It should be noted that the semantic analysis model can include, but is not limited to, pre-trained language models such as BERT (Bidirectional Encoder Representations from Transformers). The network structure of the semantic analysis model is not limited here. To train the semantic analysis model, the sample paragraph text can be first split to obtain sample paragraph combinations, and the sample paragraph combinations can include a first sample paragraph and a second sample paragraph. At least one second sample paragraph in the sample paragraph combination is randomly replaced with a third sample paragraph that is semantically unrelated to the first sample paragraph. Then, based on the difference between the prediction result of whether the two sample paragraphs in the sample paragraph combination are semantically related by the semantic analysis model and the true result, the network parameters of the semantic analysis model are adjusted. As a possible implementation example, when the semantic analysis model processes the sample paragraph combination, a prediction score (such as 0.8, 0.9, etc.) indicating whether the two sample paragraphs in the sample paragraph combination are semantically related can be obtained as the prediction result. The true result of whether the two sample paragraphs in the sample paragraph combination are semantically related can include a sample score (for example, using "0" to represent semantic unrelatedness and "1" to represent semantic relatedness). Based on this, a loss function such as cross-entropy can be used to measure the difference between the prediction score and the sample score to obtain the training loss of the semantic analysis model. Then, based on the training loss, the network parameters of the semantic analysis model are adjusted. In this way, the semantic analysis model can be forced to learn the ability to analyze whether two paragraphs are semantically related through training.

[0054] In a specific implementation scenario, after training the semantic analysis model, the element content of the first element and the element content of the second element can be processed based on the semantic analysis model to obtain an output score representing whether there is semantic relevance between the first element and the second element. Thus, in response to the output score being higher than a set threshold (e.g., 0.9, 0.95, etc.), it can be determined that there is semantic relevance between the first element and the second element, and in response to the output score not being higher than the set threshold, it can be determined that there is no semantic relevance between the first element and the second element.

[0055] In a specific implementation scenario, when the analysis result indicates semantic relevance between the first element and the second element, the first element and the second element can be combined into one layout element. Conversely, when the analysis result indicates no semantic relevance between the first element and the second element, the first element and the second element can be retained without being combined to avoid incorrect combination.

[0056] In an implementation scenario, as described above, the layout element can also include a table. The recognition result of the table can include rows, columns, and text blocks. To extract table information, post-processing can be performed first based on the recognition result of the table to obtain cells, and then, based on the overlap between the text blocks and the cells, the cells to which the text blocks belong can be determined. And when a text block covers multiple cells, the text block can be split according to the boundaries of the cells, and the sub-text blocks after splitting can be respectively assigned to each cell covered by the text block. After that, the parsing information of the layout element "table" can be obtained: the boundaries of both rows and columns, the text content of each cell, and the information on whether it belongs to the table header or the data area. In addition, when there are super cells spanning multiple rows and columns, the parsing result can further include the identifier and coverage range of the super cells. The overall confidence score of table parsing can be based on the text matching degree of the cells, the integrity of the super cells, and the overall row and column alignment. It should be noted that a row is a horizontally distributed area containing multiple cells, a column is a vertically distributed area containing multiple cells, and a super cell spans multiple rows and columns.

[0057] In a specific implementation scenario, to improve the accuracy of post-processing and text block assignment, after obtaining the recognition result of the layout element "table", it can be cleaned first. It should be noted that the recognition result of the layout element "table" can further include the confidence levels of rows, columns, and text blocks. Then, screening can be performed first based on the confidence threshold, and the non-maximum suppression (NMS) algorithm can be used to process redundant detections to ensure the accuracy and rationality of each element category.

[0058] In a specific implementation scenario, when post-processing based on the recognition results of a table, to ensure table consistency as much as possible, vertical alignment of each row can be performed based on the consistency of the boundaries of each row in the vertical direction. For example, if the left boundary of a certain row is not on the same vertical line as the left boundaries of other rows in the vertical direction, the left boundary of this row can be vertically aligned with the left boundaries of other rows in the vertical direction; similarly, horizontal alignment of each column can be performed based on the consistency of the boundaries of each column in the horizontal direction. For example, if the upper boundary of a certain column is not on the same horizontal line as the upper boundaries of other columns in the horizontal direction, the upper boundary of this column can be horizontally aligned with the upper boundaries of other columns in the horizontal direction.

[0059] In a specific implementation scenario, when post-processing based on the recognition results of a table, to ensure table integrity as much as possible, it can be determined whether a row is missing based on whether the boundaries of each row are discontinuous in the vertical direction. In the case of a missing row, the missing row can be inserted based on the spacing pattern between adjacent rows. For example, if there is a situation where the left boundary is missing between a certain row and another row in the vertical direction, the corresponding number of rows can be inserted between these two rows based on the spacing pattern between adjacent rows (such as row spacing, row height, etc.); similarly, it can be determined whether a column is missing based on whether the boundaries of each column are discontinuous in the horizontal direction. In the case of a missing column, the missing column can be inserted based on the spacing pattern between adjacent columns. For example, if there is a situation where the upper boundary is missing between a certain column and another column in the horizontal direction, the corresponding number of columns can be inserted between these two columns based on the spacing pattern between adjacent columns (such as column spacing, column height, etc.).

[0060] In a specific implementation scenario, after the above post-processing, the bounding boxes of cells can be generated based on the row-column intersection area, and the super cells that span rows and columns can be marked to ensure the accuracy of cell attributes. For conflicts between super cells and ordinary cells, they can be resolved by gradually narrowing the coverage range and verifying their logical continuity.

[0061] In a specific implementation scenario, the overlapping area between each text block and a cell can be calculated to determine the cell to which it belongs. For a text block that covers multiple cells simultaneously (e.g., a long paragraph), the text block can be split according to the boundaries of the cells so that the sub-text blocks obtained by splitting the text block are assigned to the corresponding cells. In addition, for an empty cell that does not match a text block, the possible values of the empty cell (such as "total", "grand total", etc.) can be inferred based on the content of adjacent cells. In the text content calibration stage, the character recognition process will uniformly format the extracted text, such as removing extra spaces, correcting recognition errors, and assigning confidence scores to the text matching of each cell to evaluate its accuracy and mark cells with low confidence for further inspection. It should be noted that if it cannot be inferred, it is marked as an empty cell for subsequent manual verification.

[0062] In an implementation scenario, after obtaining the parsing information of various layout elements (such as the final hierarchical title sequence of the layout element "title", the reading order of each layout element, etc.), these parsing information can be output in a structured form such as JSON, and can include but are not limited to the following fields: label (element type), text (text information, html information of the table, etc.), coordinate (coordinate information), page (page number), etc. The fields in the structured file are not limited here.

[0063] It should be noted that as a possible implementation example, please refer to Figure 2d , Figure 2d is a schematic diagram of the process of an embodiment of the document recognition method of the present application. As Figure 2dAs shown, the document to be recognized can be split first to obtain several document pages, and then each document page is converted into a picture format for text recognition (such as using OCR recognition technology) and layout analysis (such as using the YOLO parsing model). Then, the recognition results of text recognition and the analysis results of layout analysis are merged to obtain the layout elements and the recognition results of layout elements in the document to be recognized, and the layout elements can at least include titles. Of course, the layout elements can also include paragraphs, pictures, tables, etc. For the layout element "title", the final hierarchical title sequence of the layout element "title" can be obtained through operations such as initially confirming the title level, refining the title level, and correcting the title sequence (such as the above-mentioned process steps of correcting the title level and verifying the title sequence). In addition, based on the division method matching the column layout of the document to be recognized, the reading order of each layout element in the document to be recognized can be determined. For the layout elements spanning pages (that is, the layout elements at the bottom of the previous document page and the layout elements at the top of the next document page among two adjacent document pages), it can be determined whether to merge them into one layout element according to whether the content of the two elements is semantically related. In addition, for the layout element "table", its structure can be analyzed. For specific details, please refer to the relevant descriptions above and will not be elaborated here. Finally, the structured information of the document to be recognized can be obtained according to the parsing information of various layout elements.

[0064] In the above solution, recognition is performed based on the document to be recognized to obtain the layout elements and the recognition results of the layout elements in the document to be recognized, and the layout elements at least include titles. Based on the recognition results of the titles, analysis is performed to obtain the first title sequence, and the title level of the first title sequence is corrected to obtain the second title sequence. Based on the second title sequence, verification is performed to obtain the verification result, and the verification result indicates whether the second title sequence is correct. Then, in response to the verification result indicating that the second title sequence is incorrect, the second title sequence is selected as the new first title sequence, and the step of returning the title level of correcting the first title sequence to obtain the second title sequence is iterated until the end condition is met. Therefore, by correcting the title level, verifying the title sequence, and iterating by returning the corrected title level in case of incorrect verification, even in the cross-page scenario, the coherence of title modeling can be ensured as much as possible. Therefore, the coherence of title modeling during document recognition can be improved to distinguish the hierarchical relationships of each title in the document, especially in the cross-page scenario.

[0065] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of an embodiment of the intelligent interaction method of the present application. Specifically, the following steps can be included:

[0066] Step S31: Obtain the first statement to be responded to and obtain the reference document required to respond to the first statement.

[0067] In an implementation scenario, the first statement can be obtained in forms such as text input or voice input. In addition, the specific content of the first statement can also vary according to different interaction scenarios. For example, when the interaction scenario is e-commerce, the specific content of the first statement can include, but is not limited to, "Does this XXXX model of headphones have a wireless function?"; or, when the interaction scenario is medical services, the specific content of the first statement can include, but is not limited to, "What is the daily dosage of XXX capsules?" Of course, the above examples are only several possible examples in the actual application process, and the specific content of the first statement will not be exemplified one by one here.

[0068] In an implementation scenario, similar to the aforementioned first statement, the reference document can also vary according to different interaction scenarios. For example, when the interaction scenario is e-commerce, taking the first statement "Does this XXXX model of headphones have a wireless function?" as an example, the reference document can be the instruction manual document of the XXXX model of headphones; or, when the interaction scenario is medical services, taking the first statement "What is the daily dosage of XXX capsules?" as an example, the reference document can be the instruction manual document of XXX capsules. Of course, the above examples are only several possible examples in the actual application process, and the specific content of the first statement will not be exemplified one by one here.

[0069] Step S32: Perform recognition based on the reference document to obtain the structured information of the reference document.

[0070] In the embodiments of the present disclosure, the structured information at least includes the hierarchical title sequence of the reference document, and the structured information is obtained through the process steps in the embodiments of the above document recognition method. For details, reference can be made to the foregoing embodiments of the document recognition method, which will not be elaborated here.

[0071] Step S33: Perform a search based on the structured information to obtain a reference fragment for responding to the first statement.

[0072] Specifically, the intent of the first statement can be recognized first to obtain the target intent of the first statement, and then the target intent can be used to search in the structured information to obtain a reference segment for responding to the first statement. For example, when the interaction scenario is e-commerce, still taking the first statement "Does this pair of headphones of model XXXX have a wireless function?" as an example, the target intent of the first statement is "Whether the headphones of model XXXX have a wireless function", and then search in the "Function Description" of the structured information with this target intent to obtain a reference segment; or, when the interaction scenario is medical service, still taking the first statement "What is the daily dosage of XXX capsules?" as an example, the target intent of the first statement is "The daily dosage of XXX capsules", and then search in the "Usage and Dosage" of the structured information with this target intent to obtain a reference segment. Of course, the above examples are only several possible examples in the actual application process, and no further examples will be given for the specific content of the first statement here.

[0073] Step S34: Generate a second statement for responding to the first statement based on the reference segment.

[0074] Specifically, as a possible implementation example, a generative large model can be used to generate a second statement for responding to the first statement according to the reference segment. For the specific principle, the technical details of the generative large model can be referred to and will not be elaborated here. In addition, the second statement can be output in forms such as text, voice, etc., which will not be limited here either.

[0075] In the above solution, the first statement to be responded to is obtained, and the reference document required to respond to the first statement is obtained. Recognition is performed based on the reference document to obtain the structured information of the reference document, and the structured information at least includes the hierarchical title sequence of the reference document. The structured information is obtained through the process steps in the embodiments of the above document recognition method. Then, search is performed based on the structured information to obtain a reference segment for responding to the first statement, so as to generate a second statement for responding to the first statement based on the reference segment. Since the structured information is obtained through the process steps in the embodiments of the above document recognition method and the reference segment is searched based on this, it is possible to ensure the coherence of the content of the reference segment as much as possible, and thus the accuracy of the response to the first statement can be improved when generating a response based on the reference segment.

[0076] Please refer to Figure 4 , Figure 4It is a schematic framework diagram of an embodiment of the document recognition device of the present application. The document recognition device 40 includes: an identification module 41, an analysis module 42, a correction module 43, a verification module 44, and a loop module 45. The identification module 41 is used to perform identification based on the document to be recognized, and obtain the layout elements and the recognition results of the layout elements in the document to be recognized; wherein, the layout elements at least include a title. The analysis module 42 is used to analyze based on the recognition result of the title to obtain a first title sequence. The correction module 43 is used to correct the title level of the first title sequence to obtain a second title sequence. The verification module 44 is used to perform verification based on the second title sequence to obtain a verification result; wherein, the verification result indicates whether the second title sequence is correct. The loop module 45 is used to, in response to the verification result indicating that the second title sequence is incorrect, select the second title sequence as the new first title sequence, and return to the step of correcting the title level of the first title sequence to obtain the second title sequence for iteration until the end condition is met.

[0077] In the above solution, the document recognition device 40 performs recognition based on the document to be recognized, obtains the layout elements and the recognition results of the layout elements in the document to be recognized, and the layout elements at least include a title. It analyzes based on the recognition result of the title to obtain a first title sequence, corrects the title level of the first title sequence to obtain a second title sequence, performs verification based on the second title sequence to obtain a verification result, and the verification result indicates whether the second title sequence is correct. Then, in response to the verification result indicating that the second title sequence is incorrect, it selects the second title sequence as the new first title sequence, and returns to the step of correcting the title level of the first title sequence to obtain the second title sequence for iteration until the end condition is met. Therefore, by correcting the title level, verifying the title sequence, and returning to correct the title level for iteration when the verification is incorrect, it is possible to ensure the coherence of title modeling as much as possible even in a cross-page scenario. Therefore, it is possible to improve the coherence of title modeling during document recognition to distinguish the hierarchical relationships of each title in the document, especially in a cross-page scenario.

[0078] In some publicly disclosed embodiments, the document recognition device 40 includes a training module, which is used to use the annotation information marked by the sample title sequence as the training target, and train a sequence verification model based on the sample title sequence; wherein, the sample title sequence at least includes an incorrect title sequence, and the annotation information at least includes the modification method of the incorrect title sequence. The correction module 43 is specifically used to correct the title level of the first title sequence based on the sequence verification model to obtain a second title sequence. The verification module 44 is specifically used to verify the second title sequence based on the sequence verification model, and obtain a prediction mark indicating whether the second title sequence is incorrect as the verification result.

[0079] In some disclosed embodiments, the document recognition device 40 includes an export module for exporting a title based on a sample document to obtain a correct title sequence; the document recognition device 40 includes a scrambling module for scrambling based on the correct title sequence to obtain a sample title sequence and a modification method.

[0080] In some disclosed embodiments, the correction module 43 includes a processing sub-module for processing the first title sequence based on a sequence verification model to predict a modification method for the first title sequence; the correction module 43 includes a modification sub-module for modifying the title level of the first title sequence based on the modification method of the first title sequence to obtain a second title sequence.

[0081] In some disclosed embodiments, the recognition result includes the element type and the bounding block of a layout element. The document recognition device 40 includes a layout module for determining the column layout of the document to be recognized based on the element type of the layout element to which the bounding block straddling the center line of the layout belongs; wherein the column layout is any one of a single-column layout and a multi-column layout; the document recognition device 40 includes a division module for determining the reading order of each layout element in the document to be recognized based on a division method matching the column layout.

[0082] In some disclosed embodiments, the layout module includes a proportion statistics sub-module for performing statistics based on the element type of the layout element to which the bounding block straddling the center line of the layout belongs to obtain a statistical result; wherein the statistical result includes the quantity proportion of various element types; the layout module includes a first response sub-module for determining that the column layout is a multi-column layout in response to the statistical result indicating that the element types with a quantity proportion exceeding the proportion threshold only involve a target type; the layout module includes a second response sub-module for determining that the column layout is a single-column layout in response to the statistical result indicating that the element types with a quantity proportion exceeding the proportion threshold are more than the target type; wherein the target type includes at least one of a title and a picture.

[0083] In some disclosed embodiments, the division module includes a splitting sub-module for, when the column layout is a single-column layout, determining whether to split the bounding block into multiple bounding blocks based on the gap between adjacent projection segments after projecting the bounding block along the horizontal direction, and determining whether to split the bounding block into multiple bounding blocks based on the gap between adjacent projection segments after projecting the bounding block along the vertical direction; wherein there is a one-to-one correspondence between the layout element and the bounding block, and after each split, each layout element is re-ordered based on the position information of the bounding block corresponding to each layout element respectively and then the next split is performed.

[0084] In some disclosed embodiments, the partitioning module includes a sorting sub-module, configured to perform an initial sorting according to the writing norms based on the position information of the externally-connected blocks when the column layout is a multi-column layout, so as to obtain an initial block sequence; the partitioning module includes a traversing sub-module, configured to sequentially traverse each externally-connected block in the initial block sequence, and incorporate the externally-connected block into one of the column sub-sequences based on the relative positions between the position information of the two vertices on both sides of the externally-connected block and a plurality of reference lines; wherein the plurality of reference lines includes at least one of the center line of the layout, the one-fourth line of the layout, and the three-fourths line of the layout, and the column sub-sequences at least include a left column sub-sequence and a right column sub-sequence; the partitioning module includes a splicing sub-module, configured to splice each column sub-sequence according to the writing norms to obtain the reading order of each layout element.

[0085] In some disclosed embodiments, the recognition result includes the element type and element content of the layout element. The document recognition device 40 includes a selection module, configured to select the layout element located at the bottom of the first page as the first element, and select the layout element located at the top of the second page and having the same element type as the first element as the second element; wherein the first page and the second page are adjacent pages in the document to be recognized; the document recognition device 40 includes a semantics module, configured to analyze based on the element content of the first element and the element content of the second element to obtain an analysis result representing whether there is semantic relevance between the first element and the second element; the document recognition device 40 includes a merging module, configured to determine whether to merge the first element and the second element into one layout element based on the analysis result.

[0086] In some disclosed embodiments, the analysis result is obtained by analyzing with a semantic analysis model. The document recognition device 40 includes a decomposition sub-module, configured to split based on the sample paragraph text to obtain a sample paragraph combination; wherein the sample paragraph combination includes a first sample paragraph and a second sample paragraph; the document recognition device 40 includes a replacement sub-module, configured to randomly replace the second sample paragraph in at least one sample paragraph combination with a third sample paragraph that has no semantic relevance to the first sample paragraph; the document recognition device 40 includes a parameter adjustment sub-module, configured to adjust the network parameters of the semantic analysis model based on the difference between the predicted result and the true result of whether the two sample paragraphs in the sample paragraph combination are semantically relevant.

[0087] In some disclosed embodiments, the layout element further includes a table. The recognition result of the table includes rows, columns, and text blocks. The document recognition device 40 includes a post-processing module, configured to perform post-processing based on the recognition result of the table to obtain table cells; the document recognition device 40 includes an assignment module, configured to determine the table cell to which the text block belongs based on the overlapping situation between the text block and the table cells; wherein when the text block covers multiple table cells, the text sub-blocks after the text block is split according to the boundaries of the table cells are respectively assigned to each table cell covered by the text block.

[0088] In some disclosed embodiments, the post-processing module is specifically configured to perform at least one of the following: perform vertical alignment on each row based on the consistency of the boundaries of each row in the vertical direction, and perform horizontal alignment on each column based on the consistency of the boundaries of each column in the horizontal direction; determine whether a row is missing based on whether the boundaries of each row are discontinuous in the vertical direction, and in the case of a missing row, insert the missing row based on the spacing pattern between adjacent rows, and determine whether a column is missing based on whether the boundaries of each column are discontinuous in the horizontal direction, and in the case of a missing column, insert the missing column based on the spacing pattern between adjacent columns.

[0089] Please refer to Figure 5 , Figure 5 , which is a schematic diagram of the framework of an embodiment of the intelligent interaction device of the present application. The intelligent interaction device 50 includes: an acquisition module 51, an identification module 52, a search module 53, and a generation module 54. The acquisition module 51 is configured to acquire a first statement to be responded to and acquire a reference document required to respond to the first statement; the identification module 52 is configured to perform identification based on the reference document to obtain the structured information of the reference document; wherein the structured information at least includes the hierarchical title sequence of the reference document, and the structured information is obtained through the above-mentioned document identification device; the search module 53 is configured to perform a search based on the structured information to obtain a reference segment for responding to the first statement; the generation module 54 is configured to generate a second statement for responding to the first statement based on the reference segment.

[0090] In the above solution, the intelligent interaction device 50 acquires the first statement to be responded to, acquires the reference document required to respond to the first statement, performs identification based on the reference document to obtain the structured information of the reference document, and the structured information at least includes the hierarchical title sequence of the reference document. The structured information is obtained through the above-mentioned document identification device, and then a search is performed based on the structured information to obtain a reference segment for responding to the first statement, so as to generate a second statement for responding to the first statement based on the reference segment. Since the structured information is obtained through the above-mentioned document identification device and the reference segment is searched based on this, it is possible to ensure as much as possible the coherence of the content of the reference segment, and thus the accuracy of the response to the first statement can be improved when generating a response based on the reference segment.

[0091] Please refer to Figure 6 , Figure 6It is a schematic diagram of the framework of an embodiment of the electronic device of the present application. The electronic device 60 at least includes a memory 61 and a processor 62 that are coupled to each other. At least program instructions are stored in the memory 61, and the processor 62 is configured to execute the program instructions to implement the steps in any of the above-described embodiments of the document recognition method or the steps in any of the above-described embodiments of the intelligent interaction method. For details, reference can be made to the foregoing disclosed embodiments, which will not be elaborated herein. As a possible example, the electronic device 60 may include, but is not limited to, a smart phone, a tablet computer, a learning machine, an office book, a translator, a smart watch, a server, etc. The specific type of the electronic device 60 is not limited herein.

[0092] Specifically, the processor 62 is configured to control itself and the memory 61 to implement the steps in any of the above-described embodiments of the document recognition method or the steps in any of the above-described embodiments of the intelligent interaction method. The processor 62 may also be referred to as a CPU (Central Processing Unit). The processor 62 may be an integrated circuit chip with signal processing capabilities. The processor 62 may also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field-Programmable Gate Array, FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 62 may be implemented jointly by integrated circuit chips.

[0093] In the above solution, the electronic device 60 performs recognition based on the document to be recognized, obtains the layout elements and the recognition results of the layout elements in the document to be recognized, and the layout elements at least include a title. Analyze based on the recognition result of the title to obtain a first title sequence, correct the title levels of the first title sequence to obtain a second title sequence, verify based on the second title sequence to obtain a verification result, and the verification result indicates whether the second title sequence is correct. Then, in response to the verification result indicating that the second title sequence is incorrect, select the second title sequence as the new first title sequence, and return to the step of correcting the title levels of the first title sequence to obtain the second title sequence for iteration until the end condition is met. Therefore, by correcting the title levels, verifying the title sequence, and returning to correct the title levels for iteration in the case of incorrect verification, even in a multi-page scenario, it is possible to ensure the coherence of title modeling as much as possible. Therefore, it is possible to improve the coherence of title modeling during document recognition to distinguish the hierarchical relationships of each title in the document, especially in a multi-page scenario. In addition, obtain the first statement to be responded to, obtain the reference document required to respond to the first statement, perform recognition based on the reference document to obtain the structured information of the reference document, and the structured information at least includes the hierarchical title sequence of the reference document. The structured information is obtained through the process steps of the above document recognition method, and then search based on the structured information to obtain a reference segment for responding to the first statement, so as to generate a second statement for responding to the first statement based on the reference segment. Since the structured information is obtained through the process steps of the above document recognition method and the reference segment is searched based on this, it is possible to ensure the coherence of the content of the reference segment as much as possible. Therefore, when generating a response based on the reference segment, the accuracy of the response to the first statement can be improved.

[0094] Please refer to Figure 7 , Figure 7 is a schematic framework diagram of an embodiment of the computer-readable storage medium 70 of the present application. The computer-readable storage medium 70 stores program instructions 71 that can be run by a processor. The program instructions 71 are used to implement the steps in any of the above document recognition method embodiments or the steps in any of the above intelligent interaction method embodiments.

[0095] In the above solution, the computer-readable storage medium 70 performs recognition based on the document to be recognized, obtains the layout elements and the recognition results of the layout elements in the document to be recognized, and the layout elements at least include a title. Based on the recognition result of the title, analysis is performed to obtain a first title sequence, and the title levels of the first title sequence are corrected to obtain a second title sequence. Verification is performed based on the second title sequence to obtain a verification result, and the verification result indicates whether the second title sequence is correct. Then, in response to the verification result indicating that the second title sequence is incorrect, the second title sequence is selected as the new first title sequence, and the step of returning the correction of the title levels of the first title sequence to obtain the second title sequence is iterated until the end condition is met. Therefore, by correcting the title levels, verifying the title sequence, and returning the correction of the title levels for iteration in the case of incorrect verification, even in a multi-page scenario, the coherence of title modeling can be ensured as much as possible. Thus, the coherence of title modeling during document recognition can be improved to distinguish the hierarchical relationships of each title in the document, especially in a multi-page scenario. In addition, a first statement to be responded to is obtained, and a reference document required to respond to the first statement is obtained. Recognition is performed based on the reference document to obtain the structured information of the reference document, and the structured information at least includes the hierarchical title sequence of the reference document. The structured information is obtained through the process steps of the above document recognition method. Then, a reference segment for responding to the first statement is obtained based on the structured information, so as to generate a second statement for responding to the first statement based on the reference segment. Since the structured information is obtained through the process steps of the above document recognition method and the reference segment is searched based on this, the coherence of the content of the reference segment can be ensured as much as possible. Therefore, the accuracy of the response to the first statement can be improved when generating a response based on the reference segment.

[0096] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0097] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0098] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0099] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0100] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0101] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical discs that can store program codes.

[0102] If the technical solution of this application involves personal information, before the product applying the technical solution of this application processes personal information, it has clearly informed the personal information processing rules and obtained the individual's independent consent. If the technical solution of this application involves sensitive personal information, before the product applying the technical solution of this application processes sensitive personal information, it has obtained the individual's separate consent and at the same time meets the requirements of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform that the personal information collection range has been entered and personal information will be collected. If an individual voluntarily enters the collection range, it is considered consent to the collection of their personal information; or on the device for personal information processing, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up messages or asking the individual to upload their personal information by themselves, etc.; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

Claims

1. A document recognition method, characterized in that, Including: Performing recognition based on a document to be recognized to obtain layout elements in the document to be recognized and recognition results of the layout elements; wherein, the layout elements at least include a title; Analyzing based on the recognition result of the title to obtain a first title sequence; Correcting the title levels of the first title sequence to obtain a second title sequence; Verifying based on the second title sequence to obtain a verification result; wherein, the verification result indicates whether the second title sequence is correct; In response to the verification result indicating that the second title sequence is incorrect, selecting the second title sequence as a new first title sequence, and returning to the step of correcting the title levels of the first title sequence to obtain a second title sequence for iteration until an end condition is met; Wherein, before correcting the title levels of the first title sequence to obtain a second title sequence, the method further includes: Training a sequence verification model based on a sample title sequence with the annotation information marked by the sample title sequence as a training target; wherein, the sample title sequence at least includes an incorrect title sequence, and the annotation information at least includes a modification method of the incorrect title sequence; The correcting the title levels of the first title sequence to obtain a second title sequence includes: Correcting the title levels of the first title sequence based on the sequence verification model to obtain the second title sequence; The verifying based on the second title sequence to obtain a verification result includes: Verifying the second title sequence based on the sequence verification model to obtain a prediction label indicating whether the second title sequence is incorrect as the verification result.

2. The method according to claim 1, wherein The obtaining step of the sample title sequence includes: Exporting titles based on a sample document to obtain a correct title sequence; Disrupting based on the correct title sequence to obtain the sample title sequence and the modification method.

3. The method according to claim 1, wherein The correcting the title levels of the first title sequence based on the sequence verification model to obtain the second title sequence includes: Processing the first title sequence based on the sequence verification model to predict a modification method of the first title sequence; Correcting the title levels of the first title sequence based on the modification method of the first title sequence to obtain the second title sequence.

4. The method according to claim 1, wherein The recognition result includes an element type and an external bounding block of the layout element, and the method further includes: Determining a column layout of the document to be recognized based on the element type of the layout element to which an external bounding block crossing a layout center line belongs; wherein, the column layout is any one of a single-column layout and a multi-column layout; Determining a reading order of each layout element in the document to be recognized based on a division method matching the column layout.

5. The method according to claim 4, wherein The determining a column layout of the document to be recognized based on the element type of the layout element to which an external bounding block crossing a layout center line belongs includes: Performing statistics based on the element type of the layout element to which an external bounding block crossing the layout center line belongs to obtain a statistical result; wherein, the statistical result includes a quantity proportion of each element type; In response to the statistical result indicating that the element types with the quantity ratio exceeding the ratio threshold only involve the target type, determine that the column layout is the multi-column layout; In response to the statistical result indicating that the element types with the quantity ratio exceeding the ratio threshold are more than the target type, determine that the column layout is the single-column layout; wherein, the target type includes at least one of a title and a picture.

6. The method according to claim 4, characterized in that When the column layout is the single-column layout, determining the reading order of each of the layout elements in the document to be recognized based on the division method matching the column layout includes: Based on the gap between adjacent projection segments after projecting the circumscribed block along the horizontal direction, determine whether to split the circumscribed block into multiple circumscribed blocks, and based on the gap between adjacent projection segments after projecting the circumscribed block along the vertical direction, determine whether to split the circumscribed block into multiple circumscribed blocks; Wherein, the layout elements correspond to the circumscribed blocks one by one, and after each split, each of the layout elements is re-sorted based on the position information of the circumscribed block corresponding to each of the layout elements one by one and then the next split is performed.

7. The method according to claim 4, wherein When the column layout is the multi-column layout, determining the reading order of each of the layout elements in the document to be recognized based on the division method matching the column layout includes: Based on the position information of the circumscribed block, perform an initial sorting according to the writing specification to obtain an initial block sequence; Traverse each of the circumscribed blocks in the initial block sequence in turn, and based on the relative positions between the position information of the two vertices on both sides of the circumscribed block and several reference lines, incorporate the circumscribed block into one of the column sub-sequences; wherein, the several reference lines include at least one of the center line of the layout, the one-fourth line of the layout, and the three-fourths line of the layout, and the column sub-sequences at least include a left column sequence and a right column sequence; Splice each of the column sub-sequences according to the writing specification to obtain the reading order of each of the layout elements.

8. The method according to claim 1, wherein The recognition result includes the element type and element content of the layout element, and the method further includes: Select the layout element located at the bottom of the first page as the first element, and select the layout element located at the top of the second page and having the same element type as the first element as the second element; wherein, the first page and the second page are adjacent pages in the document to be recognized; Based on the element content of the first element and the element content of the second element, analyze to obtain an analysis result indicating whether there is semantic relevance between the first element and the second element; Based on the analysis result, determine whether to merge the first element and the second element into one layout element.

9. The method according to claim 8, wherein The analysis result is obtained by analyzing with a semantic analysis model, and the training steps of the semantic analysis model include: Based on the sample paragraph text, perform splitting to obtain a sample paragraph combination; wherein, the sample paragraph combination includes a first sample paragraph and a second sample paragraph; Randomly replace the second sample paragraph in at least one of the sample paragraph combinations with a third sample paragraph that is semantically irrelevant to the first sample paragraph; Based on the difference between the prediction result and the true result of whether two sample paragraphs in the sample paragraph combination are semantically related by the semantic analysis model, adjust the network parameters of the semantic analysis model.

10. The method according to claim 1, wherein The layout element further includes a table, and the recognition result of the table includes rows, columns, and text blocks. The method further includes: Performing post-processing based on the recognition result of the table to obtain cells; Based on the overlapping situation between the text block and the cells, determine the cell to which the text block belongs; wherein, when the text block covers multiple cells, the text sub-blocks after the text block is split according to the boundaries of the cells are respectively assigned to each of the cells covered by the text block.

11. The method according to claim 10, wherein The implementation steps of the post-processing include at least one of the following: Vertically align each row based on the consistency of the boundaries of each row in the vertical direction, and horizontally align each column based on the consistency of the boundaries of each column in the horizontal direction; Based on whether the boundaries of each row are discontinuous in the vertical direction, determine whether a row is missing. In the case of a missing row, insert the missing row based on the spacing pattern between adjacent rows, and based on whether the boundaries of each column are discontinuous in the horizontal direction, determine whether a column is missing. In the case of a missing column, insert the missing column based on the spacing pattern between adjacent columns.

12. An intelligent interaction method, characterized in that, Include: Obtain a first statement to be responded to, and obtain a reference document required to respond to the first statement; Perform recognition based on the reference document to obtain the structured information of the reference document; wherein, the structured information at least includes the hierarchical title sequence of the reference document, and the structured information is obtained by the document recognition method according to any one of claims 1 to 11. Search the reference document based on the structured information to obtain a reference segment for responding to the first statement; Generate a second statement for responding to the first statement based on the reference segment.

13. A document recognition device, characterized in that, Include: A recognition module for performing recognition based on a document to be recognized to obtain the layout elements in the document to be recognized and the recognition results of the layout elements; wherein, the layout elements at least include a title; An analysis module for analyzing based on the recognition result of the title to obtain a first title sequence; A correction module for correcting the title levels of the first title sequence to obtain a second title sequence; A verification module for verifying based on the second title sequence to obtain a verification result; wherein, the verification result indicates whether the second title sequence is correct; A loop module for, in response to the verification result indicating that the second title sequence is incorrect, selecting the second title sequence as a new first title sequence, and returning to the step of correcting the title levels of the first title sequence to obtain a second title sequence for iteration until an end condition is met. The document recognition device further includes a training module, which is configured to use the annotation information labeled by the sample title sequence as the training target, and train a sequence verification model based on the sample title sequence; wherein the sample title sequence at least includes an incorrect title sequence, and the annotation information at least includes the modification method of the incorrect title sequence; the correction module is specifically configured to correct the title level of the first title sequence based on the sequence verification model to obtain the second title sequence; the verification module is specifically configured to verify the second title sequence based on the sequence verification model, and obtain a prediction marker indicating whether the second title sequence is incorrect as the verification result.

14. An intelligent interaction device, characterized in that, Comprising: An acquisition module, configured to acquire a first statement to be responded to, and acquire a reference document required to respond to the first statement; An identification module, configured to perform identification based on the reference document to obtain the structured information of the reference document; wherein the structured information at least includes the hierarchical title sequence of the reference document, and the structured information is obtained by the document recognition device described in claim 13; A search module, configured to search the reference document based on the structured information to obtain a reference segment for responding to the first statement; A generation module, configured to generate a second statement for responding to the first statement based on the reference segment.

15. An electronic device, characterized in that, At least including a memory and a processor coupled to each other, at least program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the document recognition method described in any one of claims 1 to 11, or implement the intelligent interaction method described in claim 12.

16. A computer-readable storage medium, characterized in that, Stored with program instructions that can be run by a processor, the program instructions are used to implement the document recognition method described in any one of claims 1 to 11, or implement the intelligent interaction method described in claim 12.

Citation Information

Patent Citations

  • Document processing method, computer terminal and computer readable storage medium

    CN117251538A

  • Power emergency unstructured document directory construction method and device, computer equipment, storage medium and computer program product

    CN119066141A

  • Document segmentation method and device, computer equipment and storage medium

    CN119474250A

  • PDF extraction method and system based on deep learning and layout analysis

    CN119598971A

  • Method and device for identifying layout structure, equipment and medium

    CN119720942A