Document identification method, intelligent interaction method, related device, equipment and medium

By analyzing, correcting and verifying title levels in document recognition, the problem of cross-page content title hierarchical modeling is solved, and the consistency and accuracy of title modeling in document recognition is achieved.

CN119990112AActive Publication Date: 2025-05-13ANHUI IFLYTEK INTELLIGENT SYST

Patent Information

Application Number
CN202510454667.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-13
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The prior art is difficult to effectively model the title hierarchy relationship of span page content in document recognition, which makes it difficult to reconstruct the document logical structure.

Method used

Through a document recognition method, the layout elements and their recognition results are obtained based on the document to be identified, the title recognition results are analyzed, the title level is corrected, the title sequence is verified, and the correction is iteratively corrected when there is error in verification until the end condition is met.

Benefits of technology

Improves the consistency of title modeling during document recognition, ensuring that the hierarchical relationships of each title in the document can be accurately distinguished in a cross-page scenario.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990112A_ABST
    Figure CN119990112A_ABST
Patent Text Reader

Abstract

The invention discloses a document recognition method, an intelligent interaction method, a related device, equipment and a medium, and the document recognition method comprises the steps: carrying out the recognition based on a to-be-recognized document, and obtaining a layout element in the to-be-recognized document and a recognition result of the layout element; analyzing based on the identification result of the title to obtain a first title sequence; correcting the title level of the first title sequence to obtain a second title sequence; performing verification based on the second title sequence to obtain a verification result; wherein the verification result represents whether the second title sequence is correct or not; and in response to the verification result representing that the second title sequence is wrong, selecting the second title sequence as a new first title sequence, and returning to the step of correcting the title level of the first title sequence and obtaining the second title sequence for iteration until an end condition is met. According to the scheme, the continuity of title modeling during document recognition can be improved, so that the hierarchical relation of all titles in the document can be distinguished, especially in a cross-page scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing technology, and in particular to a document recognition method and an intelligent interaction method and related devices, equipment, and media. Background Art

[0002] Document recognition is of great significance in many scenarios. For example, by identifying and extracting structured information in documents, it can ensure that retrieval enhancement generation can correctly identify the document structure when processing retrieval tasks in vertical fields, thereby accurately and completely extracting key knowledge from it.

[0003] Among them, the correct identification of document titles is an important part of document recognition. However, existing technologies usually analyze single pages as units, lacking coherence modeling of cross-page content, so that they can only identify the existence of titles, but it is difficult to distinguish their hierarchical relationships, which in turn makes it difficult to reconstruct the logical structure of the document. In view of this, how to improve the coherence of title modeling during document recognition to distinguish the hierarchical relationships of various titles in the document, especially in cross-page scenarios, has become an urgent problem to be solved. Summary of the invention

[0004] The main technical problem solved by the present application is to provide a document recognition method and an intelligent interaction method and related devices, equipment, and media, which can improve the consistency of title modeling during document recognition to distinguish the hierarchical relationship between each title in the document, especially in cross-page scenarios.

[0005] In order to solve the above technical problems, the first aspect of the present application provides a document recognition method, including: performing recognition based on the document to be recognized, obtaining layout elements and recognition results of the layout elements in the document to be recognized; wherein the layout elements at least include the title; performing analysis based on the recognition result of the title to obtain a first title sequence; correcting the title level of the first title sequence to obtain a second title sequence; performing verification based on the second title sequence to obtain a verification result; wherein the verification result indicates whether the second title sequence is correct; in response to the verification result indicating that the second title sequence is incorrect, selecting the second title sequence as the new first title sequence, and returning to correct the title level of the first title sequence, and iterating the steps of obtaining the second title sequence until the end condition is met.

[0006] In order to solve the above technical problems, the second aspect of the present application provides an intelligent interaction method, including: obtaining a first sentence to be responded to, and obtaining a reference document required to respond to the first sentence; identifying based on the reference document to obtain structured information of the reference document; wherein the structured information at least includes a hierarchical title sequence of the reference document, and the structured information is obtained through the document identification method in the above first aspect; searching based on the structured information to obtain a reference fragment for responding to the first sentence; based on the reference fragment, generating a second sentence for responding to the first sentence.

[0007] In order to solve the above technical problems, the third aspect of the present application provides a document recognition device, including: a recognition module, an analysis module, a correction module, a verification module, and a loop module, wherein the recognition module is used to perform recognition based on the document to be recognized, and obtain the layout elements and recognition results of the layout elements in the document to be recognized; wherein the layout elements at least include the title; the analysis module is used to perform analysis based on the recognition result of the title, and obtain a first title sequence; the correction module is used to correct the title level of the first title sequence, and obtain a second title sequence; the verification module is used to perform verification based on the second title sequence, and obtain a verification result; wherein the verification result indicates whether the second title sequence is correct; the loop module is used to select the second title sequence as the new first title sequence in response to the verification result indicating that the second title sequence is incorrect, and return to correct the title level of the first title sequence, and iterate the steps of obtaining the second title sequence until the end condition is met.

[0008] In order to solve the above-mentioned technical problems, the fourth aspect of the present application provides an intelligent interactive device, including: an acquisition module, an identification module, a search module, and a generation module, wherein the acquisition module is used to acquire a first sentence to be responded to and acquire a reference document required to respond to the first sentence; the identification module is used to perform identification based on the reference document to obtain structured information of the reference document; wherein the structured information at least includes a hierarchical title sequence of the reference document, and the structured information is obtained through the document identification device in the above-mentioned third aspect; the search module is used to perform search based on the structured information to obtain a reference fragment for responding to the first sentence; and the generation module is used to generate a second sentence for responding to the first sentence based on the reference fragment.

[0009] In order to solve the above-mentioned technical problems, the fifth aspect of the present application provides an electronic device, which at least includes a memory and a processor coupled to each other, and the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the document recognition method in the above-mentioned first aspect, or to implement the intelligent interaction method in the above-mentioned second aspect.

[0010] In order to solve the above technical problems, the sixth aspect of the present application provides a computer-readable storage medium, which stores program instructions that can be executed by a processor, and the program instructions are used to implement the document recognition method of the first aspect above, or to implement the intelligent interaction method in the second aspect above.

[0011] The above scheme is based on the document to be identified, and the layout elements and the identification results of the layout elements in the document to be identified are obtained, and the layout elements at least include the title, and the recognition result based on the title is analyzed to obtain the first title sequence, and the title level of the first title sequence is corrected to obtain the second title sequence, and the second title sequence is verified based on the second title sequence to obtain the verification result, and the verification result indicates whether the second title sequence is correct, and then in response to the verification result indicating that the second title sequence is incorrect, the second title sequence is selected as the new first title sequence, and the title level of the first title sequence is corrected, and the steps of obtaining the second title sequence are iterated until the end condition is met. Therefore, by correcting the title level, verifying the title sequence, and returning to the correct title level for iteration when the verification is incorrect, the consistency of the title modeling can be ensured as much as possible even in the cross-page scenario. Therefore, the consistency of the title modeling during document recognition can be improved to distinguish the hierarchical relationship of each title in the document, especially in the cross-page scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a flowchart of an embodiment of the document recognition method of the present application; Figure 2a It is a schematic diagram of the process of confirming the preliminary level one embodiment of the title of the present application; Figure 2b It is a process diagram of an embodiment of the title height reset of this application; Figure 2c This is a schematic diagram of the process of refining the title level in an embodiment of the present application; Figure 2d It is a process diagram of an embodiment of the document recognition method of the present application; Figure 3 It is a flowchart of an embodiment of the intelligent interaction method of the present application; Figure 4 It is a schematic diagram of the framework of an embodiment of a document recognition device of the present application; Figure 5 It is a schematic diagram of the framework of an embodiment of the intelligent interactive device of the present application; Figure 6 It is a schematic diagram of the framework of an embodiment of the electronic device of the present application; Figure 7 It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0013] The scheme of the embodiment of the present application is described in detail below in conjunction with the drawings of the specification.

[0014] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0015] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the fragment " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. In addition, "many" in this article means two or more than two.

[0016] See also Figure 1 , Figure 1 This is a flowchart of an embodiment of the document recognition method of the present application. Specifically, it may include the following steps: Step S11: performing recognition based on the document to be recognized, and obtaining layout elements in the document to be recognized and recognition results of the layout elements.

[0017] In the disclosed embodiment, the layout element may at least include a title. It should be noted that the layout element "title" may specifically include titles at various levels, such as a main title, a first-level title, a second-level title, a third-level title, etc., which will not be listed one by one here. In addition, the recognition result of the layout element "title" may specifically include a title area (such as the coordinates of a rectangular box surrounding the title text) and title text (such as "Chapter 1 XXXX", "1.1 XXX", "1.1.1 XXXX", etc.). Of course, in actual application, layout elements may also include other types besides titles, such as paragraphs, tables, pictures, etc., which will not be listed one by one here.

[0018] In one implementation scenario, the file format of the document to be identified may include but is not limited to: PDF, doc, epub, etc., and the file format of the document to be identified is not limited here.

[0019] In an implementation scenario, in order to obtain the layout elements and their recognition results in the document to be identified, the document to be identified can be first split to obtain the individual document pages in the document to be identified, and then layout analysis and text recognition are performed on each document page respectively, so as to merge the two types of information, namely the analysis results of the layout analysis and the recognition results of the text recognition, to obtain the layout elements and their recognition results in the document page. Finally, the layout elements and their recognition results in each document page are sequentially formed into a set, which can be used as the layout elements and their recognition results in the document to be identified.

[0020] In a specific implementation scenario, as a possible implementation example, a layout analysis model such as YOLO can be used to perform layout analysis on a document page to obtain each layout element (such as title, paragraph, table, picture, etc.) and its coordinate information in the document page, which is the analysis result of the layout analysis of the document page.

[0021] In a specific implementation scenario, as a possible implementation example, a text recognition tool such as TesseractOCR can be used to perform text recognition on a document page to obtain text strings in text blocks in the document page and coordinate information of the text blocks, which is the text recognition result of the document page.

[0022] In a specific implementation scenario, after obtaining the analysis results of layout analysis and the recognition results of text recognition, the above two types of information can be merged. Specifically, for each layout element in the document page, the coordinate information of the layout element can be matched with the coordinate information of each text block for overlap. If the overlap between the text block and the layout element is higher than the set threshold, the text block is attributed to the layout element. For example, if the overlap between the coordinate information of the text block "Chapter 1 XXXX" and the layout element "Title" is higher than the set threshold, the text block "Chapter 1 XXXX" can be attributed to the layout element "Title". Of course, the above example is only one possible example in actual application, and other possible situations will not be given one by one here. In addition, in the actual application process, there are also situations where the text block cannot be classified into any layout element based on the coordinate information of the layout element and the coordinate information of the text block. In this case, the text block can be temporarily classified into the layout element "text" (for example, the text block "1.1.2" cannot be classified into any layout element based on its coordinate information and the coordinate information of the layout element, so it can be temporarily classified into the layout element "text"), and it can be further classified later (for example, it can be classified into the layout element "title" according to the regular expression later). For details, please refer to the subsequent related descriptions, which will not be repeated here. Of course, in the actual application process, there may also be situations where any text block cannot be classified into a certain layout element based on the coordinate information of the layout element and the coordinate information of the text block. In this case, the layout element and its coordinate information can be temporarily retained for subsequent analysis of possible empty tables or empty titles. At this point, for the document to be identified, each layout element, the coordinate information of the layout element, and the text content of the layout element on each document page can be obtained.

[0023] Step S12: Analyze based on the recognition result of the title to obtain a first title sequence.

[0024] Specifically, as mentioned above, the level of each layout element "title" (i.e., the level of title) cannot be determined through the above identification, so we can first analyze based on the identification result of the layout element "title" to obtain the first title sequence to obtain the hierarchical title sequence. In other words, the first title sequence contains each layout element "title" organized in a hierarchical structure, such as the organization in order: "Chapter 1 XXXX", "1.1 XXXX", "1.2 XXXX", "Chapter 2 XXXX", "2.1XXXX", "2.2 XXXX", etc. The specific content of the first title sequence is not limited here. Please refer to Figure 2a , Figure 2a This is a schematic diagram of the process of confirming the preliminary level of the title in this application. Figure 2a As shown, for each layout element "title" in the document to be identified, the title height can be reset first, and then the abnormal titles can be identified based on this (the abnormal titles will be reset to the layout element "text"). Then, for the remaining layout elements "title" (that is, the layout elements "title" without abnormalities), they can be clustered based on their title heights to obtain various cluster sets, and each cluster set can contain at least one layout element "title" with related title heights (for example, the title heights are roughly the same). Then, the cluster sets are sorted according to the title heights to generate a preliminary title hierarchy. It should be noted that the layout element blocks (such as rectangular boxes, etc., in short, "title blocks") represented by the coordinate information obtained through the aforementioned layout analysis may be slightly larger than the actual size, which affects the accuracy of judging the title height. Please refer to Figure 2b , Figure 2b 1 is a schematic diagram of the process of resetting the height of the title of this application. Figure 2bAs shown, as a possible implementation example, the document page can be converted into a grayscale image first, and the grayscale image can be binarized to enhance the contrast of the text area. After that, the layout element block of the layout element "title" can be horizontally scanned to obtain the actual title height of the layout element "title", and the title height reset can be completed. On this basis, before high-level clustering, in order to avoid the influence of title height errors caused by layout analysis errors on subsequent hierarchical distinctions as much as possible, abnormal titles with large height differences from regular titles can be processed first. Specifically, on the one hand, the mean and variance of the title height can be calculated, based on which the height range of regular titles can be calculated, such as [μ-3σ, μ+3σ], where μ represents the mean value and σ represents the variance. If the title height is not within this height range, it can be determined as an abnormal title, otherwise it can be determined as a regular title; on the other hand, the average number of words in the regular title can be combined to screen abnormal titles. If the number of words in the title deviates significantly from the average number of words, it can be determined as an abnormal title, otherwise it can be determined as a regular title. Finally, the title heights can be clustered based on a clustering algorithm such as K-Means (eg, the K value can be determined using the "elbow method") to obtain a preliminary title hierarchy. As a possible implementation example, the preliminary title hierarchy can be used as the first title sequence.

[0025] In addition, as mentioned above, in order to further refine the title level, regular expressions can also be used to optimize the preliminary title level. Figure 2c , Figure 2c Schematic diagram of the process of refining the title level in this application. Figure 2c As shown, we can match common title patterns based on regular expressions, such as "Chapter

[1234] ", "[1-9]", "(I)", etc., identify and extract the structural features of the title, and then correct the preliminary title hierarchy based on the matching results to ensure the rationality of the title hierarchy. In addition, we can also analyze the contextual semantics of the title to evaluate the relevance of the title to the text content. If we find that the hierarchical division does not conform to the semantic logic (for example, the text content is long but is mistakenly judged as a low-level title), we can correct the relevant title hierarchy to ensure that the hierarchical division conforms to the document logic.

[0026] It should be noted that the above examples are only several possible examples of obtaining the first title sequence based on the recognition result analysis of the layout element "title", which can be selected and used according to the actual situation in the actual application process. In addition, it does not exclude the use of other methods for obtaining the first title sequence, and other possible implementation methods are not given examples one by one here.

[0027] Step S13: Modify the title level of the first title sequence to obtain a second title sequence.

[0028] Specifically, a sequence verification model may be pre-trained to correct the title hierarchy of the first title sequence based on the sequence verification model to obtain the second title sequence.

[0029] In an implementation scenario, in order to train the sequence verification model, a sample title sequence can be obtained first, and the sample title sequence can at least include an erroneous title sequence. The sample title sequence can also be provided with annotation information, and the annotation information can at least include a modification method of the erroneous title sequence. Specifically, the title can be derived based on the sample document first to obtain a correct title sequence, and then the correct title sequence can be disturbed to obtain a sample title sequence and a modification method. For example, taking the correct title sequence {"Chapter 1 XXXX", "1.1 XXXX", "1.2 XXXX", "Chapter 2 XXXX", "2.1 XXXX", "2.2 XXXX"} as an example, it can be disturbed to obtain an erroneous title sequence {"Chapter 1 XXXX", "1.1 XXXX", "Chapter 2 XXXX", "1.2 XXXX", "2.1 XXXX", "2.2 XXXX"} and a modification method {switch the order of "Chapter 2 XXXX" and "1.2 XXXX"}. Of course, the above example is only a possible example of obtaining a sample title sequence in an actual application process, and other acquisition methods are not limited here, and no examples are given one by one.

[0030] In one implementation scenario, after obtaining a sample title sequence, the annotation information annotated by the sample title sequence can be used as a training target, and a sequence verification model can be obtained based on the sample title sequence training. Specifically, the sample title sequence can be processed based on the sequence verification model to obtain prediction information, and the prediction information includes whether the sample title sequence is an erroneous title sequence and its modification method when the sample title sequence is predicted to be an erroneous title sequence, that is, based on the difference between the prediction information and the annotation information of the sample title sequence, the network parameters of the sequence verification model can be adjusted. Exemplarily, the difference between the prediction information and the annotation information can be measured based on a loss function such as cross entropy to obtain the training loss of the sequence verification model, so as to adjust the network parameters of the sequence verification model based on the training loss. In the above manner, the sequence verification model can be forced to learn the semantic relationship between title levels and common logical rules, such as the order of increasing title levels and the contextual dependency between levels, so that the sequence verification model can detect errors in the title level sequence and provide optimization suggestions.

[0031] In an implementation scenario, after the sequence verification model is trained, the title hierarchy of the first title sequence can be corrected based on the sequence verification model to obtain a second title sequence. Specifically, the first title sequence can be processed with the sequence verification model to predict the modification method of the first title sequence, and then the title hierarchy of the first title sequence can be corrected based on the modification method of the first title sequence to obtain the second title sequence. For example, it can be predicted that the modification method of the first title sequence is to exchange the order of certain titles in the first title sequence, and the order of these titles in the first title sequence can be modified and exchanged in this way to obtain the second title sequence. Of course, the above example is only a possible example of using the sequence verification model to correct the title hierarchy in actual application, and other possible situations will not be given examples one by one here.

[0032] Step S14: Perform verification based on the second title sequence to obtain a verification result.

[0033] In the disclosed embodiment, the verification result characterizes whether the second title sequence is correct. Specifically, as mentioned above, the sequence verification model can be used to correct the title level of the first title sequence, and the second title sequence can also be verified based on the sequence verification model to obtain a prediction mark characterizing whether the second title sequence is incorrect as the verification result. For example, the aforementioned sequence verification model can be trained using a sample title sequence with annotation information, so that the sequence verification model not only has the ability to correct the title level, but also has the ability to identify whether the sequence is incorrect. Therefore, the aforementioned sequence verification model can be directly used to verify the second title sequence to obtain a verification result of whether the second title sequence is incorrect.

[0034] Step S15: In response to the verification result indicating that the second title sequence is incorrect, the second title sequence is selected as the new first title sequence, and the title hierarchy of the first title sequence is corrected, and the steps of obtaining the second title sequence are iterated until the end condition is met.

[0035] In one implementation scenario, if the verification result indicates that the second title sequence is incorrect, the second title sequence can be selected as the new first title sequence, and the steps of correcting the title level of the first title sequence to obtain the second title sequence are iterated until the end condition is met. In other words, if the verification result indicates that the second title sequence is incorrect, the title level of the second title sequence can be corrected, and then verified again after correction, and the cycle is repeated until the end condition is met. It should be noted that the end condition can be set according to the actual application, such as being set to a preset number of loop iterations (such as 5 times, 10 times, etc.); or, it can be set to the verification result of the latest second title sequence indicating that the latest second title sequence is correct. Of course, the above examples are only a few possible examples of the second title sequence, and other possible situations are not limited here, and examples are not given one by one.

[0036] In another implementation scenario, different from the aforementioned situation, if the verification result indicates that the second title sequence is correct, the latest second title sequence can be used as the final hierarchical title sequence of the document to be identified.

[0037] In one implementation scenario, the recognition results may also include the element type and external blocks of the layout elements. The column layout of the document to be identified can be determined based on the element type of the layout element to which the external blocks crossing the center line of the layout belong, and the column layout can be either a single-column layout or a multi-column layout. The reading order of each layout element in the document to be identified is then determined based on the division method that matches the column layout. Therefore, the appropriate division method can be adaptively selected according to the document to be identified to determine the reading order of the layout elements, which helps to improve the adaptability to various documents.

[0038] In a specific implementation scenario, in order to determine the column layout, the element type of the layout element to which the external block crossing the center line of the layout belongs can be first counted to obtain the statistical results, and the statistical results include the quantity proportions of various element types. For example, according to statistics, among the various element types crossing the center line of the layout: the quantity proportion of the layout element "title", the quantity proportion of the layout element "paragraph", the quantity proportion of the layout element "table", the quantity proportion of the layout element "picture", etc., no longer exemplified one by one here. In response to the statistical results representing that the element type whose quantity proportion exceeds the ratio threshold only involves the target type, the column layout is determined to be a multi-column layout. It should be noted that the ratio threshold can be set according to the actual application. For example, the ratio threshold can be set to 40%, 50%, etc., and the specific value of the ratio threshold is not limited here. In addition, the target type may include but is not limited to at least one of the title and the picture. In addition, in response to the statistical results representing that the element type whose quantity proportion exceeds the ratio threshold is more than the target type, the column layout can be determined to be a single-column layout. For example, according to statistics, there are still a large number of layout elements "paragraphs" that cross the center line of the layout, which can determine that the column layout is a single-column layout.

[0039] In a specific implementation scenario, after determining the column layout and before determining the reading order, the external blocks of each layout element can be sorted according to the writing specifications and the position information of the external blocks of each layout element. For example, the external blocks of each layout element can be sorted according to the writing specifications of "from top to bottom, from left to right" according to the "upper edge coordinates (y value)" and "left edge coordinates (x value)" of the external blocks (i.e., the upper left corner is marked), so as to pave the way for subsequent more detailed divisions.

[0040] In a specific implementation scenario, when the column layout is a single-column layout, it can be determined whether to split the external block into multiple external blocks based on the gap between adjacent projection segments after projecting the external block in the horizontal direction, and based on the gap between adjacent projection segments after projecting the external block in the vertical direction, it can be determined whether to split the external block into multiple external blocks, and the layout elements correspond one-to-one to the external blocks. After each split, each layout element is reordered based on the position information of the external blocks that correspond one-to-one to each layout element before the next split. As a possible implementation example, taking the projection of the circumscribed block in the horizontal direction as an example, if the gap between adjacent projection segments after projection is too large (e.g., greater than a set threshold), it can be determined that the circumscribed block should be regarded as different rows or different horizontal bands, so the circumscribed block can be split into multiple circumscribed blocks along the gap; similarly, taking the projection of the circumscribed block in the vertical direction as an example, if the gap between adjacent projection segments after projection is too large (e.g., greater than a set threshold), it can be determined that the circumscribed block should be regarded as different columns or different vertical bands, so the circumscribed block can be split into multiple circumscribed blocks along the gap. Splitting the circumscribed blocks in this way can gradually split them into rows, columns, and smaller granularities by relying on projection segmentation even if there is partial overlap of the circumscribed blocks, thereby effectively avoiding omissions or misclassifications, and greatly improving the recognition accuracy of the document layout.

[0041] In a specific implementation scenario, when the column layout is a multi-column layout, the initial sorting can be first performed according to the writing plan based on the position information of the external blocks to obtain an initial block sequence. For specific details, please refer to the relevant descriptions above and will not be elaborated here. Based on this, each external block in the initial block sequence can be traversed in turn. Based on the relative positions between the position information of the two vertices on both sides of the external block and several reference lines, the external block is incorporated into one of the column sub-sequences, and the several reference lines include at least one of the center line of the layout, the one-fourth line of the layout, and the three-fourths line of the layout. The column sub-sequences at least include a left column sequence and a right column sequence. For the convenience of description, the position information of the two vertices on both sides of the external block can be respectively denoted as bbox[0] (representing the position information of the left vertex) and bbox[2] (representing the position information of the right vertex). Then, if it can be determined according to the position information that the left vertex is on the left side of the one-fourth line of the layout (i.e., bbox[0] < w / 4) and the right vertex is on the left side of the three-fourths line of the layout (i.e., bbox[2] < 3*w / 4), the external block can be incorporated into the left column sequence; similarly, if it can be determined according to the position information that the left vertex is on the right side of the one-fourth line of the layout (i.e., bbox[0] > w / 4) and the right vertex is on the right side of the center line of the layout (i.e., bbox[2] > w / 2), the external block can be incorporated into the right column sequence. Here, w represents the width of the entire page. Of course, the above example is only one possible example in the actual application process, and other possible incorporation methods will not be exemplified one by one here. On this basis, each column sub-sequence can be spliced according to the writing specification to obtain the reading order of each layout element. For example, each column sub-sequence can be spliced in the order of "from top to bottom, from left to right" to obtain the reading order of each layout element. In addition, during the traversal process, if an external block straddles the center line of the layout or there is a large gap between its upper edge and the lower edge of the previous external block, it means that this external block is more likely to be a new line or an independent block. In response to this situation, after collecting the left column sequence and the right column sequence, this external block can be inserted into the final column sequence (i.e., the column sequence after splicing) to obtain the reading order of each layout element.

[0042] In one implementation scenario, the recognition results of layout elements may also include the element type and element content of the layout elements. In order to improve the semantic integrity of the element content, especially the semantic integrity of the layout elements that span pages, the layout element located at the bottom of the first page may be selected as the first element, and the layout element located at the top of the second page and of the same element type as the first element may be selected as the second element. The first page and the second page are adjacent pages in the document to be identified. Based on the element content of the first element and the element content of the second element, an analysis result is obtained to obtain whether the first element and the second element are semantically related. Based on the analysis result, it can be determined whether to merge the first element and the second element into one layout element. In this way, the layout elements that span pages but actually belong to the same semantic segment can be merged into one complete layout element to improve the semantic integrity of the layout elements.

[0043] In a specific implementation scenario, the analysis result can be obtained by analyzing the element content of the first element and the element content of the second element by a semantic analysis model. It should be noted that the semantic analysis model may include but is not limited to pre-trained language models such as BERT (Bidirectional Encoder Representations from Transformers), and the network structure of the semantic analysis model is not limited here. In order to train a semantic analysis model, it can be first split based on the sample paragraph text to obtain a sample paragraph combination, and the sample paragraph combination can include a first sample paragraph and a second sample paragraph, and randomly replace the second sample paragraph in at least one sample paragraph combination with a third sample paragraph that is semantically irrelevant to the first sample paragraph, and then adjust the network parameters of the semantic analysis model based on the difference between the predicted result of whether the two sample paragraphs in the sample paragraph combination are semantically related and the actual result. As a possible implementation example, the semantic analysis model processes the sample paragraph combination to obtain a prediction score (e.g., 0.8, 0.9, etc.) of whether two sample paragraphs in the sample paragraph combination are semantically related, which can be used as a prediction result, and the actual result of whether the two sample paragraphs in the sample paragraph combination are semantically related can include the sample score (e.g., "0" represents semantic irrelevance, and "1" represents semantic relevance). Based on this, a loss function such as cross entropy can be used to measure the difference between the prediction score and the sample score to obtain the training loss of the semantic analysis model. Based on the training loss, the network parameters of the semantic analysis model are adjusted, thereby forcing the semantic analysis model to acquire the ability to analyze whether two paragraphs are semantically related through training and learning.

[0044] In a specific implementation scenario, after the semantic analysis model is trained, the element content of the first element and the element content of the second element can be processed based on the semantic analysis model to obtain an output score that characterizes whether the first element and the second element are semantically related, so that in response to the output score being higher than a set threshold (such as 0.9, 0.95, etc.), it can be determined that the first element and the second element are semantically related, and in response to the output score being not higher than the set threshold, it can be determined that the first element and the second element are semantically unrelated.

[0045] In a specific implementation scenario, when the analysis result indicates that the first element and the second element are semantically related, the first element and the second element can be merged into one layout element. Conversely, when the analysis result indicates that the first element and the second element are semantically unrelated, the first element and the second element can be retained without being merged to avoid erroneous merging.

[0046] In an implementation scenario, as mentioned above, the layout element may also include a table, and the recognition result of the table may include rows, columns, and text blocks. In order to extract the table information, the recognition result of the table may be post-processed to obtain cells, and then the cells to which the text blocks belong may be determined based on the overlap between the text blocks and the cells. In the case where the text blocks cover multiple cells, the sub-text blocks after the text blocks are segmented according to the boundaries of the cells are respectively allocated to the cells covered by the text blocks. After that, the parsing information of the layout element "table" can be obtained: the boundaries of both rows and columns, the text content of each cell, and whether it belongs to the header or data area. In addition, when there are supercells that span rows and columns, the parsing results may further include the identification and coverage of the supercells. The overall confidence score of the table parsing can be based on the text matching degree of the cells, the integrity of the supercells, and the overall row and column alignment. It should be noted that a row is a horizontal distribution area containing multiple cells, a column is a vertical distribution area containing multiple cells, and a supercell spans multiple rows and columns.

[0047] In a specific implementation scenario, in order to improve the accuracy of post-processing and text block allocation, after obtaining the recognition result of the layout element "table", it can be cleaned first. It should be noted that the recognition result of the layout element "table" can further include the confidence of rows, columns and text blocks, then it can be filtered based on the confidence threshold, and the non-maximum suppression (NMS) algorithm can be used to process redundant detection to ensure the accuracy and rationality of each element category.

[0048] In a specific implementation scenario, when post-processing is performed based on the recognition results of the table, in order to ensure the consistency of the table as much as possible, the rows can be vertically aligned based on the consistency of the boundaries of the rows in the vertical direction. For example, if the left boundary of a row is not on the same vertical line as the left boundaries of other rows in the vertical direction, the left boundary of the row can be vertically aligned with the left boundaries of other rows in the vertical direction; similarly, based on the consistency of the boundaries of the columns in the horizontal direction, the columns are horizontally aligned. For example, if the upper boundary of a column is not on the same horizontal line as the upper boundaries of other columns in the horizontal direction, the upper boundary of the column can be horizontally aligned with the upper boundaries of other columns in the horizontal direction.

[0049] In a specific implementation scenario, when post-processing is performed based on the recognition results of the table, in order to ensure the integrity of the table as much as possible, it is possible to determine whether rows are missing based on whether the boundaries of each row are discontinuous in the vertical direction, so that in the case of missing rows, the missing rows can be inserted based on the spacing pattern between adjacent rows. For example, if there is a missing left boundary between a row and another row in the vertical direction, a corresponding number of rows can be inserted between the two rows based on the spacing pattern between adjacent rows (such as row spacing, row height, etc.); similarly, it is possible to determine whether columns are missing based on whether the boundaries of each column are discontinuous in the horizontal direction, so that in the case of missing columns, the missing columns can be inserted based on the spacing pattern between adjacent columns. For example, if there is a missing upper boundary between a column and another column in the horizontal direction, a corresponding number of columns can be inserted between the two columns based on the spacing pattern between adjacent columns (such as column spacing, column height, etc.).

[0050] In a specific implementation scenario, after the above post-processing, the bounding box of the cell can be generated based on the row and column intersection area, and the super cells that cross rows and columns can be marked to ensure the accuracy of cell attributes. The conflict between super cells and ordinary cells can be resolved by gradually reducing the coverage area and verifying its logical continuity.

[0051] In a specific implementation scenario, the overlapping area between each text block and the cell can be calculated to determine the cell to which it belongs. For text blocks that cover multiple cells at the same time (such as long paragraph content), the text blocks can be segmented according to the boundaries of the cells to assign the sub-text blocks segmented from the text blocks to the corresponding cells. In addition, for empty cells that are not matched to text blocks, the possible values ​​of the empty cells (such as "total", "total", etc.) can be inferred based on the contents of adjacent cells. In the text content calibration stage, the text recognition link will uniformly format the extracted text, such as clearing excess spaces, correcting recognition errors, and assigning confidence scores to the text matches of each cell, evaluating its accuracy, and marking low-confidence cells for further inspection. It should be noted that if it cannot be inferred, it will be marked as an empty cell for subsequent manual verification.

[0052] In an implementation scenario, after obtaining the parsed information of various layout elements (such as the final hierarchical title sequence of the layout element "title", the reading order of each layout element, etc.), these parsed information can be output in a structured form such as JSON, which may include but is not limited to the following fields: label (element type), text (text information, table html information, etc.), coordinate (coordinate information), page (page number), etc., and there is no limitation on the fields in the structured file.

[0053] It should be noted that as a possible implementation example, please refer to Figure 2d , Figure 2d Schematic diagram of the process of an embodiment of the document recognition method of the present application. Figure 2dAs shown, the document to be identified can be split first to obtain several document pages, and then each document page can be converted into an image format to perform text recognition (such as using OCR recognition technology) and layout analysis (such as using the YOLO parsing model), and then the recognition results of the text recognition and the analysis results of the layout analysis are merged to obtain the layout elements in the document to be identified and the recognition results of the layout elements, and the layout elements can at least include titles. Of course, the layout elements can also include paragraphs, pictures, tables, etc. For the layout element "title", the final hierarchical title sequence of the layout element "title" can be obtained by successively confirming the preliminary hierarchy of the title, refining the title hierarchy, and correcting the title sequence (such as the aforementioned process steps of correcting the title hierarchy, verifying the title sequence, etc.). In addition, the reading order of each layout element in the document to be identified can also be determined based on the division method that matches the column layout of the document to be identified. For layout elements that span pages (i.e., the layout elements at the bottom of the previous document page and the layout elements at the top of the next document page in two adjacent document pages), it can be determined whether to merge the two elements into one layout element based on whether the contents of the two elements are semantically related. In addition, for the layout element "table", its structure can be analyzed. For details, please refer to the above-mentioned related description, which will not be repeated here. Finally, the structured information of the document to be identified can be obtained based on the parsed information of various layout elements.

[0054] The above scheme is based on the document to be identified, and the layout elements and the identification results of the layout elements in the document to be identified are obtained, and the layout elements at least include the title, and the recognition result based on the title is analyzed to obtain the first title sequence, and the title level of the first title sequence is corrected to obtain the second title sequence, and the second title sequence is verified based on the second title sequence to obtain the verification result, and the verification result indicates whether the second title sequence is correct, and then in response to the verification result indicating that the second title sequence is incorrect, the second title sequence is selected as the new first title sequence, and the title level of the first title sequence is corrected, and the steps of obtaining the second title sequence are iterated until the end condition is met. Therefore, by correcting the title level, verifying the title sequence, and returning to the correct title level for iteration when the verification is incorrect, the consistency of the title modeling can be ensured as much as possible even in the cross-page scenario. Therefore, the consistency of the title modeling during document recognition can be improved to distinguish the hierarchical relationship of each title in the document, especially in the cross-page scenario.

[0055] See also Figure 3 , Figure 3 This is a flow chart of an embodiment of the intelligent interaction method of the present application. Specifically, it may include the following steps: Step S31: Obtain a first statement to be responded to, and obtain a reference document required to respond to the first statement.

[0056] In an implementation scenario, the first sentence can be obtained through text input, voice input, etc. In addition, the specific content of the first sentence may also be different depending on the interaction scenario. For example, when the interaction scenario is e-commerce, the specific content of the first sentence may include but is not limited to "Does this XXXX model headset have wireless function?" etc.; or, when the interaction scenario is medical service, the specific content of the first sentence may include but is not limited to "What is the daily dosage of XXX capsule?" etc. Of course, the above examples are only a few possible examples in actual application, and the specific content of the first sentence will not be given one by one here.

[0057] In an implementation scenario, similar to the aforementioned first sentence, the reference document may also be different depending on the interaction scenario. For example, when the interaction scenario is e-commerce, taking the first sentence "Does this XXXX model headset have wireless function?" as an example, the reference document may be the instruction document of the XXXX model headset; or, when the interaction scenario is medical services, taking the first sentence "What is the daily dosage of XXX capsules?" as an example, the reference document may be the instruction document of XXX capsules. Of course, the above examples are only a few possible examples in actual application, and the specific content of the first sentence will not be given one by one here.

[0058] Step S32: performing identification based on the reference document to obtain structured information of the reference document.

[0059] In the disclosed embodiment, the structured information at least includes the hierarchical title sequence of the reference document, and the structured information is obtained through the process steps in the above-mentioned document recognition method embodiment. For details, please refer to the above-mentioned document recognition method embodiment, which will not be repeated here.

[0060] Step S33: Search based on the structured information to obtain a reference segment for responding to the first sentence.

[0061] Specifically, the intent recognition can be performed based on the first sentence to obtain the target intent of the first sentence, and then the target intent can be used to search in the structured information to obtain a reference fragment for responding to the first sentence. For example, when the interaction scenario is e-commerce, the first sentence "Does this XXXX model headset have a wireless function?" is still taken as an example. The target intent of the first sentence is "Does the XXXX model headset have a wireless function?" Then this target intent is used to search in the "Function Description" of the structured information to obtain a reference fragment; or, when the interaction scenario is a medical service, the first sentence "What is the daily dosage of XXX capsules?" is still taken as an example. The target intent of the first sentence is "The daily dosage of XXX capsules", and this target intent is used to search in the "Usage and Dosage" of the structured information to obtain a reference fragment. Of course, the above examples are only a few possible examples in actual application, and the specific content of the first sentence will not be given one by one here.

[0062] Step S34: Based on the reference segment, generate a second sentence in response to the first sentence.

[0063] Specifically, as a possible implementation example, a generative big model can be used to generate a second sentence for responding to the first sentence based on a reference fragment. The specific principle can refer to the technical details of the generative big model, which will not be repeated here. In addition, the second sentence can be output in the form of text, voice, etc., which is not limited here.

[0064] The above scheme obtains the first sentence to be responded to, and obtains the reference document required for responding to the first sentence, performs identification based on the reference document, obtains structured information of the reference document, and the structured information at least includes the hierarchical title sequence of the reference document. The structured information is obtained through the process steps in the above document recognition method embodiment, and then a search is performed based on the structured information to obtain a reference fragment for responding to the first sentence, and then a second sentence for responding to the first sentence is generated based on the reference fragment. Since the structured information is obtained through the process steps in the above document recognition method embodiment and the reference fragment is searched based on it, the coherence of the content of the reference fragment can be ensured as much as possible, and the accuracy of the response to the first sentence can be improved when a response is generated based on the reference fragment.

[0065] See also Figure 4 , Figure 4It is a schematic diagram of the framework of an embodiment of the document recognition device of the present application. The document recognition device 40 includes: a recognition module 41, an analysis module 42, a correction module 43, a verification module 44, and a loop module 45. The recognition module 41 is used to perform recognition based on the document to be recognized, and obtain the layout elements and the recognition results of the layout elements in the document to be recognized; wherein the layout elements at least include the title; the analysis module 42 is used to perform analysis based on the recognition result of the title to obtain the first title sequence; the correction module 43 is used to correct the title level of the first title sequence to obtain the second title sequence; the verification module 44 is used to perform verification based on the second title sequence to obtain the verification result; wherein the verification result indicates whether the second title sequence is correct; the loop module 45 is used to select the second title sequence as the new first title sequence in response to the verification result indicating that the second title sequence is incorrect, and return to the step of correcting the title level of the first title sequence to obtain the second title sequence for iteration until the end condition is met.

[0066] In the above scheme, the document recognition device 40 performs recognition based on the document to be recognized, obtains the layout elements and recognition results of the layout elements in the document to be recognized, and the layout elements at least include the title, analyzes based on the recognition result of the title, obtains the first title sequence, and corrects the title level of the first title sequence to obtain the second title sequence, verifies based on the second title sequence, obtains the verification result, and the verification result indicates whether the second title sequence is correct, and then in response to the verification result indicating that the second title sequence is incorrect, selects the second title sequence as the new first title sequence, and returns to correct the title level of the first title sequence, and iterates the steps of obtaining the second title sequence until the end condition is met. Therefore, by correcting the title level, verifying the title sequence, and returning to correct the title level for iteration when the verification is incorrect, even in the cross-page scenario, the consistency of the title modeling can be ensured as much as possible. Therefore, the consistency of the title modeling during document recognition can be improved to distinguish the hierarchical relationship of each title in the document, especially in the cross-page scenario.

[0067] In some disclosed embodiments, the document recognition device 40 includes a training module for obtaining a sequence verification model based on the sample title sequence training, using the annotation information annotated by the sample title sequence as a training target; wherein the sample title sequence includes at least an erroneous title sequence, and the annotation information includes at least a modification method for the erroneous title sequence; the correction module 43 is specifically used to correct the title hierarchy of the first title sequence based on the sequence verification model to obtain a second title sequence; the verification module 44 is specifically used to verify the second title sequence based on the sequence verification model to obtain a prediction mark representing whether the second title sequence is incorrect as a verification result.

[0068] In some disclosed embodiments, the document identification device 40 includes an export module for exporting titles based on sample documents to obtain a correct title sequence; the document identification device 40 includes a disruption module for disrupting the correct title sequence to obtain a sample title sequence and a modification method.

[0069] In some disclosed embodiments, the correction module 43 includes a processing submodule for processing the first title sequence based on the sequence verification model to predict the modification method of the first title sequence; the correction module 43 includes a modification submodule for modifying the title level of the first title sequence based on the modification method of the first title sequence to obtain a second title sequence.

[0070] In some disclosed embodiments, the recognition results include element types and external blocks of layout elements, and the document recognition device 40 includes a layout module for determining the column layout of the document to be recognized based on the element type of the layout element to which the external block crossing the centerline of the layout belongs; wherein the column layout is either a single-column layout or a multi-column layout; the document recognition device 40 includes a division module for determining the reading order of each layout element in the document to be recognized based on a division method matching the column layout.

[0071] In some disclosed embodiments, the layout module includes a proportion statistics submodule, which is used to perform statistics based on the element types of the layout elements belonging to the external blocks crossing the center line of the layout to obtain statistical results; wherein the statistical results include the quantity proportions of various element types; the layout module includes a first response submodule, which is used to determine that the column layout is a multi-column layout in response to the statistical results indicating that the element types whose quantity proportions exceed the proportion threshold only involve the target type; the layout module includes a second response submodule, which is used to determine that the column layout is a single-column layout in response to the statistical results indicating that the element types whose quantity proportions exceed the proportion threshold are more than the target type; wherein the target type includes at least one of a title or a picture.

[0072] In some disclosed embodiments, the division module includes a splitting submodule for determining whether to split an external block into multiple external blocks based on a gap between adjacent projection segments after projecting the external block in a horizontal direction, and determining whether to split an external block into multiple external blocks based on a gap between adjacent projection segments after projecting the external block in a vertical direction when the column layout is a single-column layout; wherein layout elements correspond one-to-one to external blocks, and after each split, each layout element is reordered based on position information of the external blocks corresponding one-to-one to each layout element before the next split.

[0073] In some disclosed embodiments, the division module includes a sorting submodule for performing initial sorting according to text specifications based on position information of external blocks when the column layout is a multi-column layout, so as to obtain an initial block sequence; the division module includes a traversal submodule for traversing each external block in the initial block sequence in turn, and incorporating the external block into one of the column subsequences based on the position information of vertices on both sides of the external block and the relative position between a number of reference lines; wherein the several reference lines include at least one of a layout center line, a layout quarter line, and a layout three-quarter line, and the column subsequence includes at least a left column sequence and a right column sequence; the division module includes a splicing submodule for splicing each column subsequence according to text specifications to obtain a reading order of each layout element.

[0074] In some disclosed embodiments, the recognition result includes the element type and element content of the layout element, and the document recognition device 40 includes a selection module for selecting the layout element located at the bottom of the first page as the first element, and selecting the layout element located at the top of the second page and having the same element type as the first element as the second element; wherein the first page and the second page are adjacent pages in the document to be recognized; the document recognition device 40 includes a semantic module for analyzing, based on the element content of the first element and the element content of the second element, to obtain an analysis result characterizing whether the first element and the second element are semantically related; the document recognition device 40 includes a merging module for determining, based on the analysis result, whether to merge the first element and the second element into one layout element.

[0075] In some disclosed embodiments, the analysis results are obtained by analyzing the semantic analysis model, and the document recognition device 40 includes a decomposition submodule for splitting based on the sample paragraph text to obtain a sample paragraph combination; wherein the sample paragraph combination includes a first sample paragraph and a second sample paragraph; the document recognition device 40 includes a replacement submodule for randomly replacing the second sample paragraph in at least one sample paragraph combination with a third sample paragraph that is semantically irrelevant to the first sample paragraph; the document recognition device 40 includes a parameter adjustment submodule for adjusting the network parameters of the semantic analysis model based on the difference between the predicted result and the actual result of whether two sample paragraphs in the sample paragraph combination are semantically related by the semantic analysis model.

[0076] In some disclosed embodiments, the layout elements also include a table, and the recognition results of the table include rows, columns, and text blocks. The document recognition device 40 includes a post-processing module for performing post-processing based on the recognition results of the table to obtain cells; the document recognition device 40 includes an allocation module for determining the cell to which the text block belongs based on the overlap between the text block and the cell; wherein, in the case where the text block covers multiple cells, the text sub-blocks obtained after the text block is segmented according to the boundaries of the cells are respectively allocated to each cell covered by the text block.

[0077] In some disclosed embodiments, the post-processing module is specifically used to perform at least one of the following: vertically aligning the rows based on the consistency of the boundaries of the rows in the vertical direction, and horizontally aligning the columns based on the consistency of the boundaries of the columns in the horizontal direction; determining whether a row is missing based on whether the boundaries of the rows are discontinuous in the vertical direction, and in the case of missing rows, inserting the missing row based on the spacing pattern between adjacent rows, and determining whether a column is missing based on whether the boundaries of the columns are discontinuous in the horizontal direction, and in the case of missing columns, inserting the missing column based on the spacing pattern between adjacent columns.

[0078] See also Figure 5 , Figure 5 It is a schematic diagram of the framework of an embodiment of the intelligent interactive device of the present application. The intelligent interactive device 50 includes: an acquisition module 51, an identification module 52, a search module 53, and a generation module 54. The acquisition module 51 is used to acquire the first sentence to be responded to and acquire the reference document required to respond to the first sentence; the identification module 52 is used to identify based on the reference document to obtain the structured information of the reference document; wherein the structured information at least includes the hierarchical title sequence of the reference document, and the structured information is obtained by the above-mentioned document identification device; the search module 53 is used to search based on the structured information to obtain the reference fragment for responding to the first sentence; the generation module 54 is used to generate the second sentence for responding to the first sentence based on the reference fragment.

[0079] In the above scheme, the intelligent interactive device 50 obtains the first sentence to be responded to, and obtains the reference document required for responding to the first sentence, and performs identification based on the reference document to obtain structured information of the reference document, and the structured information at least includes the hierarchical title sequence of the reference document. The structured information is obtained by the above-mentioned document recognition device, and then a search is performed based on the structured information to obtain a reference fragment for responding to the first sentence, and a second sentence for responding to the first sentence is generated based on the reference fragment. Since the structured information is obtained by the above-mentioned document recognition device and the reference fragment is searched based on it, the coherence of the content of the reference fragment can be ensured as much as possible, and the accuracy of the response to the first sentence can be improved when a response is generated based on the reference fragment.

[0080] See also Figure 6 , Figure 6It is a schematic diagram of the framework of an embodiment of an electronic device of the present application. The electronic device 60 includes at least a memory 61 and a processor 62 coupled to each other, the memory 61 at least stores program instructions, and the processor 62 is used to execute the program instructions to implement the steps in any of the above document recognition method embodiments, or to implement the steps in any of the above intelligent interaction method embodiments. For details, please refer to the aforementioned disclosed embodiments, which will not be repeated here. As a possible example, the electronic device 60 may include but is not limited to smart phones, tablet computers, learning machines, office books, translation machines, smart watches, servers, etc., and the specific type of the electronic device 60 is not limited here.

[0081] Specifically, the processor 62 is used to control itself and the memory 61 to implement the steps in any of the above document recognition method embodiments, or to implement the steps in any of the above intelligent interaction method embodiments. The processor 62 can also be called a CPU (Central Processing Unit). The processor 62 may be an integrated circuit chip with signal processing capabilities. The processor 62 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 62 can be implemented by an integrated circuit chip.

[0082] In the above scheme, the electronic device 60 performs identification based on the document to be identified, obtains the layout elements and the identification results of the layout elements in the document to be identified, and the layout elements at least include the title, analyzes based on the identification result of the title, obtains the first title sequence, and corrects the title level of the first title sequence to obtain the second title sequence, verifies based on the second title sequence, obtains the verification result, and the verification result indicates whether the second title sequence is correct, and then in response to the verification result indicating that the second title sequence is incorrect, selects the second title sequence as the new first title sequence, and returns to correct the title level of the first title sequence, and iterates the steps of obtaining the second title sequence until the end condition is met. Therefore, by correcting the title level, verifying the title sequence, and returning to correct the title level for iteration when the verification is incorrect, even in the cross-page scenario, the consistency of the title modeling can be ensured as much as possible. Therefore, the consistency of the title modeling during document recognition can be improved to distinguish the hierarchical relationship of each title in the document, especially in the cross-page scenario. In addition, a first sentence to be responded to is obtained, and a reference document required for responding to the first sentence is obtained, identification is performed based on the reference document, and structured information of the reference document is obtained, and the structured information at least includes a hierarchical title sequence of the reference document. The structured information is obtained through the process steps of the above-mentioned document recognition method, and then a search is performed based on the structured information to obtain a reference fragment for responding to the first sentence, and a second sentence for responding to the first sentence is generated based on the reference fragment. Since the structured information is obtained through the process steps of the above-mentioned document recognition method and the reference fragment is searched based on it, the coherence of the content of the reference fragment can be ensured as much as possible, and thus the accuracy of the response to the first sentence can be improved when a response is generated based on the reference fragment.

[0083] See also Figure 7 , Figure 7 The computer-readable storage medium 70 of the present application is a schematic diagram of a framework of an embodiment. The computer-readable storage medium 70 stores program instructions 71 that can be executed by a processor, and the program instructions 71 are used to implement the steps in any of the above document recognition method embodiments, or to implement the steps in any of the above intelligent interaction method embodiments.

[0084] In the above scheme, the computer-readable storage medium 70 performs identification based on the document to be identified, obtains the layout elements and the identification results of the layout elements in the document to be identified, and the layout elements at least include the title, analyzes based on the identification result of the title, obtains the first title sequence, and corrects the title hierarchy of the first title sequence to obtain the second title sequence, verifies based on the second title sequence, obtains the verification result, and the verification result indicates whether the second title sequence is correct, and then in response to the verification result indicating that the second title sequence is incorrect, selects the second title sequence as the new first title sequence, and returns to correct the title hierarchy of the first title sequence, and iterates the steps of obtaining the second title sequence until the end condition is met. Therefore, by correcting the title hierarchy, verifying the title sequence, and returning to correct the title hierarchy for iteration when the verification is incorrect, the consistency of the title modeling can be ensured as much as possible even in the cross-page scenario. Therefore, the consistency of the title modeling during document recognition can be improved to distinguish the hierarchical relationship of each title in the document, especially in the cross-page scenario. In addition, a first sentence to be responded to is obtained, and a reference document required for responding to the first sentence is obtained, identification is performed based on the reference document, and structured information of the reference document is obtained, and the structured information at least includes a hierarchical title sequence of the reference document. The structured information is obtained through the process steps of the above-mentioned document recognition method, and then a search is performed based on the structured information to obtain a reference fragment for responding to the first sentence, and a second sentence for responding to the first sentence is generated based on the reference fragment. Since the structured information is obtained through the process steps of the above-mentioned document recognition method and the reference fragment is searched based on it, the coherence of the content of the reference fragment can be ensured as much as possible, and thus the accuracy of the response to the first sentence can be improved when a response is generated based on the reference fragment.

[0085] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0086] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.

[0087] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0088] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0089] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0090] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.

[0091] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

Claims

1. A document recognition method, characterized in that: include: Performing recognition based on the document to be recognized, obtaining layout elements in the document to be recognized and recognition results of the layout elements; wherein the layout elements at least include a title; Analyze based on the recognition result of the title to obtain a first title sequence; Correcting the title level of the first title sequence to obtain a second title sequence; Performing verification based on the second title sequence to obtain a verification result; wherein the verification result indicates whether the second title sequence is correct; In response to the verification result indicating that the second title sequence is incorrect, the second title sequence is selected as the new first title sequence, and the title level of the first title sequence is corrected and the step of obtaining the second title sequence is iterated until an end condition is met.

2. The method according to claim 1, characterized in that Before modifying the title level of the first title sequence to obtain the second title sequence, the method further includes: Taking the annotation information of the sample title sequence as the training target, a sequence verification model is obtained based on the training of the sample title sequence; wherein the sample title sequence at least includes an erroneous title sequence, and the annotation information at least includes a modification method of the erroneous title sequence; The step of modifying the title level of the first title sequence to obtain a second title sequence includes: Correcting the title level of the first title sequence based on the sequence checking model to obtain the second title sequence; The verifying based on the second title sequence to obtain a verification result includes: The second title sequence is verified based on the sequence verification model to obtain a prediction mark indicating whether the second title sequence is incorrect as the verification result.

3. The method according to claim 2, characterized in that The step of obtaining the sample title sequence comprises: Export titles based on sample documents to obtain the correct title sequence; The correct title sequence is disturbed to obtain the sample title sequence and the modification method.

4. The method according to claim 2, characterized in that: The step of correcting the title level of the first title sequence based on the sequence checking model to obtain the second title sequence includes: Processing the first title sequence based on the sequence verification model to predict a modification method of the first title sequence; The title level of the first title sequence is modified based on the modification method of the first title sequence to obtain the second title sequence.

5. The method according to claim 1, characterized in that The recognition result includes the element type and the external block of the layout element, and the method further includes: Determine the column layout of the document to be identified based on the element type of the layout element to which the external block crossing the center line of the layout belongs; wherein the column layout is any one of a single-column layout and a multi-column layout; Based on a division method that matches the column layout, a reading order of each of the layout elements in the document to be identified is determined.

6. The method according to claim 5, characterized in that The determining of the column layout of the document to be identified based on the element type of the layout element to which the external block crossing the center line of the layout belongs includes: Based on the element types of the layout elements belonging to the external blocks crossing the center line of the layout, statistical results are obtained; wherein the statistical results include the proportion of the number of various element types; In response to the statistical result indicating that the element types whose quantity proportion exceeds the ratio threshold only involve the target type, determining that the column layout is the multi-column layout; In response to the statistical result indicating that the element types whose quantity proportion exceeds the ratio threshold are more than the target type, the column layout is determined to be the single-column layout; wherein the target type includes at least one of a title and a picture.

7. The method according to claim 5, characterized in that In the case where the column layout is the single-column layout, determining the reading order of each of the layout elements in the document to be identified based on the division method matching the column layout includes: Determine whether to split the circumscribed block into multiple circumscribed blocks based on a gap between adjacent projection segments after projecting the circumscribed block in a horizontal direction, and determine whether to split the circumscribed block into multiple circumscribed blocks based on a gap between adjacent projection segments after projecting the circumscribed block in a vertical direction; The layout elements correspond to the external blocks one by one, and after each split, the layout elements are reordered based on the position information of the external blocks corresponding to each layout element before the next split.

8. The method according to claim 5, characterized in that In the case where the column layout is the multi-column layout, determining the reading order of each of the layout elements in the document to be identified based on the division method matching the column layout includes: Based on the position information of the external blocks, an initial sorting is performed according to the text specification to obtain an initial block sequence; Sequentially traverse each of the external blocks in the initial block sequence, and include the external block into one of the column subsequences based on the position information of the vertices on both sides of the external block and the relative positions between a plurality of reference lines; wherein the plurality of reference lines include at least one of the layout center line, the layout quarter line, and the layout three quarter line, and the column subsequence includes at least a left column sequence and a right column sequence; The column subsequences are spliced ​​according to the writing specification to obtain the reading order of the layout elements.

9. The method according to claim 1, characterized in that: The recognition result includes the element type and element content of the layout element, and the method further includes: Selecting a layout element at the bottom of a first page as a first element, and selecting a layout element at the top of a second page and having the same element type as the first element as a second element; wherein the first page and the second page are adjacent pages in the document to be identified; Based on the element content of the first element and the element content of the second element, analyzing to obtain an analysis result indicating whether the first element and the second element are semantically related; Based on the analysis result, determine whether to merge the first element and the second element into one layout element.

10. The method according to claim 9, characterized in that The analysis result is obtained by analyzing a semantic analysis model, and the training steps of the semantic analysis model include: Splitting based on the sample paragraph text to obtain a sample paragraph combination; wherein the sample paragraph combination includes a first sample paragraph and a second sample paragraph; Randomly replacing the second sample paragraph in at least one of the sample paragraph combinations with a third sample paragraph that is semantically irrelevant to the first sample paragraph; Based on the difference between the prediction result of the semantic analysis model on whether two sample paragraphs in the sample paragraph combination are semantically related and the actual result, the network parameters of the semantic analysis model are adjusted.

11. The method according to claim 1, characterized in that: The layout elements further include a table, and the recognition result of the table includes rows, columns, and text blocks. The method further includes: Performing post-processing based on the recognition result of the table to obtain a cell; Based on the overlap between the text block and the cell, the cell to which the text block belongs is determined; wherein, in the case where the text block covers multiple cells, the text sub-blocks obtained after the text block is divided according to the boundaries of the cells are respectively allocated to each of the cells covered by the text block.

12. The method according to claim 11, characterized in that The post-processing implementation steps include at least one of the following: Align the rows vertically based on the consistency of the boundaries of the rows in the vertical direction, and align the columns horizontally based on the consistency of the boundaries of the columns in the horizontal direction; Based on whether the boundaries of each row are discontinuous in the vertical direction, determine whether a row is missing, so that in the case of a missing row, the missing row is inserted based on the spacing pattern between adjacent rows. Based on whether the boundaries of each column are discontinuous in the horizontal direction, determine whether a column is missing, so that in the case of a missing column, the missing column is inserted based on the spacing pattern between adjacent columns.

13. An intelligent interaction method, characterized in that: include: Obtaining a first statement to be responded to, and obtaining a reference document required to respond to the first statement; Based on the reference document, identification is performed to obtain structured information of the reference document; wherein the structured information at least includes a hierarchical title sequence of the reference document, and the structured information is obtained by the document identification method according to any one of claims 1 to 12; Searching the reference document based on the structured information to obtain a reference segment for responding to the first sentence; Based on the reference segment, a second sentence is generated to respond to the first sentence.

14. A document recognition device, characterized in that: include: A recognition module, used for performing recognition based on a document to be recognized, obtaining layout elements in the document to be recognized and recognition results of the layout elements; wherein the layout elements at least include a title; An analysis module, configured to perform analysis based on the recognition result of the title to obtain a first title sequence; A correction module, used for correcting the title level of the first title sequence to obtain a second title sequence; A verification module, configured to perform verification based on the second title sequence to obtain a verification result; wherein the verification result indicates whether the second title sequence is correct; A loop module is used for selecting the second title sequence as a new first title sequence in response to the verification result indicating that the second title sequence is incorrect, and returning to the title level of the first title sequence to correct the first title sequence, and iterating the step of obtaining the second title sequence until an end condition is met.

15. An intelligent interactive device, characterized in that: include: An acquisition module, used to acquire a first statement to be responded to, and acquire a reference document required to respond to the first statement; an identification module, configured to identify the reference document based on the reference document to obtain structured information of the reference document; wherein the structured information at least includes a hierarchical title sequence of the reference document, and the structured information is obtained by the document identification device of claim 14; A search module, configured to search the reference document based on the structured information to obtain a reference segment for responding to the first sentence; A generating module is used to generate a second statement for responding to the first statement based on the reference segment.

16. An electronic device, characterized in that: It at least comprises a memory and a processor coupled to each other, wherein the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the document recognition method described in any one of claims 1 to 12, or to implement the intelligent interaction method described in claim 13.

17. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the document recognition method described in any one of claims 1 to 12, or to implement the intelligent interaction method described in claim 13.

Citation Information

Patent Citations

  • Document hierarchy division method, document hierarchy division device and readable storage medium

    CN111079402A

  • Document header extraction method, system and device and storage medium

    CN113095061A

  • Document processing method, computer terminal and computer readable storage medium

    CN117251538A

  • Power emergency unstructured document directory construction method and device, computer equipment, storage medium and computer program product

    CN119066141A

  • Document segmentation method and device, computer equipment and storage medium

    CN119474250A

Cited By

  • Multi-dimensional data processing method and system

    CN121033876A

  • Hybrid PDF (Portable Document Format) document analysis and knowledge fragment construction method and device

    CN121527785A