Layout analysis method, device, electronic device and storage medium
By determining the candidate next sentence from the set of sentences of the image to be analyzed and sorting based on semantic information and image features, the problem of difficult to reproduce complex layout images in the prior art is solved, and an automated and more adaptable layout analysis is achieved.
Patent Information
- Application Number
- CN202210055957.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-18
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-01-18
AI Technical Summary
The prior art is difficult to realize the reproduction of layouts after text recognition, especially in images of complex layouts. Manually formulated rules require a lot of manpower and time, and it is difficult to adapt to structural changes and complex typesetting.
Layout sorting is achieved by determining the candidate next sentence for each sentence from the set of sentences to be analyzed and determining the next sentence based on semantic information and image features. This method does not require manual rules and can automate and adapt to complex layouts.
Automatic layout analysis and text sequence recovery of complex layout images are realized, saving manpower and time, and suitable for structural changes and complex typesetting images.
Smart Images

Figure CN114491129B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular, to a layout analysis method, apparatus, electronic device, and storage medium. Background Art
[0002] OCR (optical character recognition) refers to the process of examining printed characters on paper by an electronic device (such as a scanner or digital camera), and then translating the shape into computer text using character recognition methods.
[0003] With the development of image recognition technology based on deep learning, the performance of simple text recognition has approached over 95%. However, due to the ever-changing layout information of text, even if the text is recognized correctly, it may not be able to accurately reproduce the original text sequence structure, and the text sequence structure will greatly affect or even change the semantics of the original text, directly restricting the application of text recognition technology.
[0004] Currently, the acquisition of text layout information and the restoration of text sequences are mainly achieved based on manually formulated rules. However, such methods are mostly for images with fixed structures and simple layouts. For images with complex layouts such as newspapers, it is difficult to achieve layout reproduction using the above methods, and the manually formulated rules need to be customized by technical personnel, wasting manpower and time, and bringing great trouble to users. Summary of the Invention
[0005] The present invention provides a layout analysis method, apparatus, electronic device, and storage medium to solve the problem that text recognition in the prior art cannot achieve layout reproduction.
[0006] The present invention provides a layout analysis method, including:
[0007] Determining candidate next sentences for each sentence from the sentence set of the image to be analyzed;
[0008] Based on the semantic information of each sentence and its candidate next sentences, determining the next sentence for each sentence from the candidate next sentences of each sentence;
[0009] Based on the next sentences of each sentence, performing layout sorting on the sentence set.
[0010] According to a layout analysis method provided by the present invention, the determining the next sentence for each sentence from the candidate next sentences of each sentence based on the semantic information of each sentence and its candidate next sentences includes:
[0011] Based on the semantic information of each sentence and its candidate next sentence, as well as the image features of the image to be analyzed, determine the confidence of the candidate next sentence of each sentence, where the image features are used to characterize the layout distribution information of each sentence in the image to be analyzed;
[0012] Based on the confidence of the candidate next sentence of each sentence, determine the next sentence of each sentence from the candidate next sentences of each sentence.
[0013] According to a layout analysis method provided by the present invention, the determining the confidence of the candidate next sentence of each sentence based on the semantic information of each sentence and its candidate next sentence, as well as the image features of the image to be analyzed, includes:
[0014] Concatenate each sentence and its candidate next sentence to obtain the candidate text of each sentence;
[0015] Perform semantic extraction on the candidate text of each sentence to obtain the semantic features of the candidate text of each sentence, where the semantic features are used to characterize the semantic information of the corresponding sentence and its candidate next sentence;
[0016] Based on the semantic features of the candidate text of each sentence and the image features of the image to be analyzed, determine the confidence of the candidate next sentence of each sentence.
[0017] According to a layout analysis method provided by the present invention, the image features of the image to be analyzed are determined based on the following steps:
[0018] Based on an optical character recognition model, perform feature extraction on the image to be analyzed to obtain the image features, where the optical character recognition model is used to extract the image features and perform optical character recognition based on the image features.
[0019] According to a layout analysis method provided by the present invention, the determining the candidate next sentence of each sentence from the set of sentences of the image to be analyzed includes:
[0020] Based on the position of the last character of the current sentence in the image to be analyzed and the position of the first character of other sentences in the image to be analyzed, determine the distance between the current sentence and the other sentences, where the current sentence is a sentence in the set of sentences, and the other sentences are sentences in the set of sentences other than the current sentence;
[0021] Based on the distance between the current sentence and the other sentences, determine the candidate next sentence of the current sentence from the other sentences.
[0022] According to a layout analysis method provided by the present invention, the set of sentences is determined based on the following steps:
[0023] Determine the continuous state of the current text and the next text thereof based on the text positions of the current text and the next text in the image to be analyzed;
[0024] If the continuous state is continuous, place the next text into the sentence where the current text is located; otherwise, place the next text into a newly created blank sentence;
[0025] Use the next text as the new current text until sentence grouping is completed to obtain the sentence set.
[0026] According to a layout analysis method provided by the present invention, the determining the continuous state of the current text and the next text thereof based on the text positions of the current text and the next text in the image to be analyzed includes:
[0027] Determine the continuous state of the current text and the next text thereof based on the vertex positions of two vertices adjacent to the following text in the text position of the current text and the vertex positions of two vertices adjacent to the preceding text in the text position of the next text.
[0028] The present invention also provides a layout analysis device, including:
[0029] A candidate determination unit for determining candidate next sentences for each sentence from the sentence set of the image to be analyzed;
[0030] A next sentence determination unit for determining the next sentence for each sentence from the candidate next sentences of each sentence based on the semantic information of each sentence and its candidate next sentences;
[0031] A sorting unit for performing layout sorting on the sentence set based on the next sentence of each sentence.
[0032] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the layout analysis method as described in any one of the above are implemented.
[0033] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the layout analysis method as described in any one of the above are implemented.
[0034] The layout analysis method, device, electronic device, and storage medium provided by the present invention perform upper and lower sentence judgment based on the semantic information of each sentence and its candidate next sentence, so as to determine the next sentence of the sentence from the candidate next sentences, realize the layout sorting of the sentences in the image to be analyzed, and do not need to apply artificially formulated layout sorting rules throughout the process, avoiding the waste of manpower and time caused by artificially specified rules, and being equally applicable to images with structural changes or complex layouts, realizing automated and more adaptable layout analysis, which helps to broaden the application of layout analysis. Brief Description of the Drawings
[0035] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly describe the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 is one of the flowcharts of the layout analysis method provided by the present invention;
[0037] Figure 2 is one of the flowcharts of the next sentence selection method provided by the present invention;
[0038] Figure 3 is the second flowchart of the next sentence selection method provided by the present invention;
[0039] Figure 4 is the flowchart of the sentence set determination method provided by the present invention;
[0040] Figure 5 is the second flowchart of the layout analysis method provided by the present invention;
[0041] Figure 6 is the structural diagram of the layout analysis device provided by the present invention;
[0042] Figure 7 is the structural diagram of the electronic device provided by the present invention. Detailed Embodiments
[0043] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0044] In the related art, the acquisition of text layout information and the restoration of text sequences are mainly achieved based on manually formulated rules. However, such methods are mostly applicable to images with fixed structures and simple layouts. For images with complex layouts such as newspapers and magazines and without fixed formats, the layout may be adjusted in each issue. It is difficult to reproduce the layout using the above methods, and the manually formulated rules need to be customized by technical personnel, wasting manpower and time, which brings great trouble to users.
[0045] With the development of natural language processing technology, although text intelligent error correction technology has been widely applied, text intelligent error correction technology corrects text errors such as spelling, grammar, and wrong characters in the complete semantic sequence of long texts, and cannot solve the problem of sequence disorder of longer texts.
[0046] In view of this, an embodiment of the present invention provides a layout analysis method. Figure 1 It is one of the flow schematic diagrams of the layout analysis method provided by the present invention. As Figure 1 shown, the method includes:
[0047] Step 110, determine candidate next sentences for each sentence from the sentence set of the image to be analyzed.
[0048] Specifically, the image to be analyzed is the image that needs to perform layout analysis. The image to be analyzed may contain multiple sentences with context relationships. The image to be analyzed can be obtained by image acquisition devices such as scanners, mobile phones, and cameras, or can be obtained by web crawling or downloading. The embodiments of the present invention do not make specific limitations on this.
[0049] Text recognition can be performed on the image to be analyzed to obtain the text contained in the image to be analyzed. On this basis, according to the positions of the texts in the image to be analyzed, the sizes and fonts of the texts themselves, etc., the texts contained in the image to be analyzed are grouped into sentences, so as to obtain the sentences in the image to be analyzed to construct a sentence set; or a pre-trained text recognition model with a sentence grouping function can also be used to directly obtain each sentence in the image to be analyzed to construct a sentence set. The embodiments of the present invention do not make specific limitations on this.
[0050] Here, the sentences in the sentence set of the image to be analyzed are disordered. In order to restore the text semantics in the image to be recognized, the disordered sentences are sorted and connected. Each sentence in the image to be analyzed can be judged one by one whether there is a next sentence to be continued, so as to obtain the layout sorting of each sentence in the image to be analyzed.
[0051] In this process, one of the sentences can be used as the current sentence. Here, the selection of the current sentence can be based on the writing or typesetting habits during the initial analysis. The sentence distributed at the top or upper left corner of the image to be analyzed can be selected as the first sentence, that is, the initial current sentence; during the analysis process, the next sentence based on the previous current sentence can be used as the current sentence, so as to achieve orderly analysis of the sentences. It is also possible to analyze the next sentence of each sentence in parallel, and after determining the next sentence of each sentence. The overall ranking of the sentence set can be integrated.
[0052] The selection of candidate next sentences of each sentence can be explained by taking any sentence in the sentence set as the current sentence. For the current sentence, the sentence closest to the current sentence or the sentence closest to the current sentence in the printing format such as font and font size can be selected from the sentence set of the image to be analyzed as the candidate next sentence of the current sentence. The candidate next sentence here is the sentence that may be the following sentence of the current sentence, and the candidate next sentence can be one or more.
[0053] Step 120: Determine the next sentence of each sentence from the candidate next sentences of each sentence based on the semantic information of each sentence and its candidate next sentences.
[0054] Specifically, after obtaining the candidate next sentences of each sentence, the next sentence of each sentence can be selected from the candidate next sentences. Taking the current sentence as an example, the next sentence of the current sentence is selected based on the semantic information of the current sentence and the semantic information of the candidate next sentences of the current sentence. Here, the semantic information of the sentence can be extracted through a pre-trained language model, or can be extracted through a semantic understanding model, a question-answering model, or other models that require semantic feature extraction, and the embodiment of the present invention does not specifically limit this.
[0055] When selecting the next sentence of the current sentence, the correlation between the semantic information of the current sentence and the semantic information of the candidate next sentence can be analyzed. The more correlated the two are, the more likely the candidate next sentence is to be the next sentence of the current sentence. It is also possible to analyze whether the semantic information of the candidate next sentence inherits the semantic information of the current sentence, and judge whether the candidate next sentence may be the next sentence of the current sentence by analyzing whether the connection between the semantic information of the current sentence and the semantic information of the candidate next sentence is smooth. It is also possible to splice the semantic information of the current sentence and the candidate next sentence, and analyze whether the spliced semantic information is reasonable, so as to judge whether the candidate next sentence may be the next sentence of the current sentence. The embodiments of the present invention do not make specific limitations on this.
[0056] During the selection process, the confidence levels of each candidate next sentence can be compared horizontally as the current sentence, so as to select the candidate next sentence with the highest confidence level as the next sentence of the current sentence, or select the candidate next sentence with a confidence level greater than a preset threshold as the next sentence of the current sentence; alternatively, the smoothness or rationality of the semantic information connection of each candidate next sentence after the semantic information of the current sentence can also be compared horizontally, so as to select the smoothest or most reasonable candidate next sentence as the next sentence of the current sentence. The embodiments of the present invention do not make specific limitations on this.
[0057] Specifically, after determining the next sentence of the current sentence, the layout sorting of the current sentence is completed. By arranging the next sentence after the current sentence, the next sentence can be used as the new current sentence, and step 110 is returned to execute to obtain the candidate next sentence of the new current sentence, and then step 120 is executed to select the next sentence of the new current sentence from the new candidate next sentences, and so on in a loop until the layout sorting of each sentence in the image to be analyzed is completed, so as to obtain the text sequence of the image to be analyzed.
[0058] In addition, the determination of the next sentence for each sentence can also be executed disorderly or in parallel, as long as steps 110 and 120 are executed for each sentence in the sentence set of the image to be analyzed.
[0059] Step 130, perform layout sorting on the sentence set based on the next sentences of the respective sentences.
[0060] Specifically, after determining the next sentence for each sentence in the sentence set, there is a context relationship between the sentences in the sentence set. Thus, the sentences in the sentence set can be sorted to obtain the text sequence of the image to be analyzed, so as to restore the text semantics in the image to be recognized.
[0061] The method provided by the embodiments of the present invention determines the next sentence of a sentence from candidate next sentences based on the semantic information of each sentence and its candidate next sentences, realizes the layout sorting of sentences in the image to be analyzed, and does not need to apply artificially formulated layout sorting rules throughout the process, avoiding the waste of manpower and time caused by artificially specifying rules, and is also applicable to images with structural changes or complex layouts, realizing automated and more adaptable layout analysis, which helps to broaden the application of layout analysis.
[0062] Based on the above embodiment, in step 120, based on the semantic information of each sentence and its candidate next sentence, the next sentence of each sentence is determined from the candidate next sentence of each sentence, which can be implemented by a pre-trained model. The model here can be adjusted based on a pre-trained language model for general sentence verification or sentence pair verification. The sentence pair referred to here can be a sentence pair consisting of each sentence and its candidate next sentence. For example, the model can be used to determine whether the two sentences in the sentence pair are the upper and lower sentences, and determine whether the candidate next sentence is the next sentence of the corresponding sentence.
[0063] Here, the pre-trained language model can be obtained by multi-task model training based on massive data with the help of unsupervised data sampling technology and data back-annotation technology. The pre-trained language model can be used for the expression of text sequences; in the process of such model training, the prediction of the MASK words in the original text and the judgment of whether the sentence pairs of the original text are upper and lower sentences are usually adopted. Since the pre-trained language model has good text feature expression ability, simple fine-tuning based on specific data in different tasks can achieve better results. For example, in an embodiment of the present invention, the pre-trained language model can be adjusted based on the sample sentence and its candidate next sentence. Here, the candidate next sentence of the sample sentence is marked with a label indicating whether each candidate next sentence is the next sentence of the sample sentence.
[0064] Based on any of the above embodiments, in step 120, based on the semantic information of each sentence and its candidate next sentence, determining the next sentence of each sentence from the candidate next sentences of each sentence includes:
[0065] Determining the confidence of the candidate next sentences of each sentence based on the semantic information of each sentence and its candidate next sentences;
[0066] Based on the confidence of the candidate next sentences of the respective sentences, the next sentence of the respective sentences is determined from the candidate next sentences of the respective sentences.
[0067] Specifically, taking the current sentence as an example, when selecting the next sentence of the current sentence, the correlation between the semantic information of the current sentence and the semantic information of the candidate next sentence can be analyzed. The more correlated the two are, the higher the confidence of the candidate next sentence as the next sentence of the current sentence; it is also possible to analyze whether the semantic information of the candidate next sentence inherits the semantic information of the current sentence, and determine the confidence of the candidate next sentence as the next sentence of the current sentence by analyzing whether the connection between the semantic information of the current sentence and the semantic information of the candidate next sentence is smooth; it is also possible to splice the semantic information of the current sentence and the candidate next sentence, and analyze whether the spliced semantic information is reasonable, so as to determine the confidence of the candidate next sentence as the next sentence of the current sentence. The embodiments of the present invention do not make specific limitations on this.
[0068] During the selection process, the confidence levels of each candidate next sentence for the current sentence can be compared horizontally, so as to select the candidate next sentence with the highest confidence level as the next sentence of the current sentence, or select the candidate next sentence with a confidence level greater than a preset threshold as the next sentence of the current sentence; alternatively, the smoothness or rationality of the semantic information connection of each candidate next sentence after the semantic information of the current sentence can also be compared horizontally, so as to select the smoothest or most reasonable candidate next sentence as the next sentence of the current sentence. The embodiments of the present invention do not make specific limitations on this.
[0069] Based on any of the above embodiments, determining the confidence level of the candidate next sentence of each sentence based on the semantic information of each sentence and its candidate next sentence includes:
[0070] Concatenate each sentence and its candidate next sentence to obtain the candidate text of each sentence;
[0071] Extract the semantic features of the candidate text of each sentence to obtain the semantic features of the candidate text of each sentence, and the semantic features are used to represent the semantic information of the corresponding sentence and its candidate next sentence;
[0072] Based on the semantic features of the candidate text of each sentence, determine the confidence level of the candidate next sentence of each sentence.
[0073] Specifically, take one of the sentences as the current sentence. During the process of determining the confidence level of the candidate next sentence of the current sentence, for any candidate next sentence, the current sentence and this candidate next sentence can be used as a sentence pair, and the two can be concatenated to obtain the candidate text for this sentence pair.
[0074] By extracting the semantic features of the candidate text, the semantic features of the candidate text can be obtained. Since the candidate text is obtained by concatenating the current sentence and this candidate next sentence, the semantic features of the candidate text thus obtained reflect the semantic information after the combination of the current sentence and the candidate next sentence. Compared with the semantic information of the two being independent of each other, the semantic information reflected in the semantic features of the candidate text not only includes the semantic information of the two being independent of each other, but also includes the semantic information when the two are used as the front and back sentences, that is, the semantic information of the two being integrated into a sentence pair.
[0075] After obtaining the semantic features of the candidate text, the confidence level of this candidate next sentence can be determined based on the semantic features of the candidate text.
[0076] The method provided by the embodiments of the present invention enhances the semantic information of each sentence and its candidate next sentence as the upper and lower sentences by extracting the semantic features of the candidate text after concatenating each sentence and its candidate next sentence, which helps to further improve the reliability of the confidence level of the candidate next sentence.
[0077] Based on any of the above embodiments,Figure 2 It is one of the flow diagrams of the next sentence selection method provided by the present invention. As Figure 2 shown, in step 120, based on the semantic information of each sentence and its candidate next sentence, the next sentence of each sentence is determined from the candidate next sentences of each sentence, including:
[0078] Step 121, based on the semantic information of each sentence and its candidate next sentence, and the image features of the image to be analyzed, the confidence of the candidate next sentence of each sentence is determined, and the image features are used to characterize the layout distribution information of each sentence in the image to be analyzed.
[0079] Specifically, the image to be analyzed itself contains a large amount of text semantics, and this part of information can enhance the semantic information of the current sentence and the candidate next sentence, thereby improving the reliability of selecting the next sentence of the current sentence. In the embodiments of the present invention, the image features of the image to be analyzed can be combined with the semantic information of each sentence and its candidate next sentence, so as to determine the confidence of the candidate next sentence of each sentence as the next sentence of each sentence.
[0080] Here, the image features of the image to be analyzed can be obtained by performing image feature extraction on the image to be analyzed, or can be the image features extracted during the previous process of performing character recognition on the image to be analyzed. The embodiments of the present invention do not make specific limitations on this. The image features here reflect the layout distribution information of each sentence in the image to be analyzed, that is, the image features cover the position distribution of each sentence shown in the image to be analyzed on the layout, and the position distribution of each sentence on the layout can reflect the word order between each sentence. Therefore, the image features can reflect the semantic information of the image to be analyzed itself, that is, the overall semantic information of the text content contained in the image to be analyzed. Here, the overall text content in the image to be analyzed is composed of the content of each sentence shown in the image to be analyzed itself and the layout distribution.
[0081] Compared with the semantic information reflected by the current sentence and the candidate next sentence from the sentence text itself, the image features of the image to be analyzed reflect the semantic information contained in the image itself. The combination of the semantic information of the current sentence and the candidate next sentence and the image features can integrate the semantic information in the two modalities of text and image, thereby improving the reliability and accuracy of the semantic information used to judge the confidence of the candidate next sentence, and further improving the reliability of the confidence of the candidate next sentence.
[0082] The process of determining the confidence for each sentence is the same. Here, one of the sentences in each sentence is used as the current sentence for illustration. During the process of determining the confidence, the semantic features of the current sentence, the semantic information of the candidate next sentence, and the image features of the image to be analyzed can be fused, and the fused features can be applied to the confidence determination. Or, the confidence in the text modality and the confidence in the image modality can be determined respectively based on the semantic information of the current sentence, the semantic information of the candidate next sentence, and the image features of the image to be analyzed, and the confidences in the two modalities can be fused to obtain the final confidence. The embodiments of the present invention do not make specific limitations on this.
[0083] Step 122: Determine the next sentence of each sentence from the candidate next sentences of each sentence based on the confidence of the candidate next sentences of each sentence.
[0084] Specifically, taking one of the sentences in each sentence as the current sentence as an example, after obtaining the confidence of each candidate next sentence of the current sentence, the candidate next sentence with the highest confidence can be selected as the next sentence of the current sentence, or the candidate next sentence with a confidence greater than the preset threshold can be selected as the next sentence of the current sentence, thereby determining the next sentence of the current sentence. Based on the same method, the next sentence of each sentence can be determined.
[0085] The method provided by the embodiments of the present invention determines the confidence of the candidate next sentence by combining the semantic information of each sentence and its candidate next sentence, and the image features of the image to be analyzed, and performs feature expression from two modalities of text and image, so that the semantic information of the sentence can be combined with the semantic information of the original text in the image modality, thereby effectively improving the feature expression effect and ensuring the reliability and accuracy of the next sentence selection.
[0086] Based on any of the above embodiments, step 121 includes:
[0087] Concatenate each sentence and its candidate next sentence to obtain the candidate text of each sentence;
[0088] Extract the semantics of the candidate text of each sentence to obtain the semantic features of the candidate text of each sentence, and the semantic features are used to represent the semantic information of the corresponding sentence and its candidate next sentence;
[0089] Based on the semantic features of the candidate text of each sentence and the image features of the image to be analyzed, determine the confidence of the candidate next sentence of each sentence.
[0090] Specifically, one sentence in each sentence is used as the current sentence. In the process of determining the confidence of the candidate next sentence of the current sentence, for any candidate next sentence, the current sentence and the candidate next sentence can be used as a sentence pair, and the two can be concatenated to obtain the candidate text for this sentence pair.
[0091] By performing semantic extraction on the candidate text, the semantic features of the candidate text can be obtained. Since the candidate text is obtained by concatenating the current sentence and the candidate next sentence, the semantic features of the candidate text thus obtained reflect the semantic information after the combination of the current sentence and the candidate next sentence. Compared with the semantic information of the two being independent of each other, the semantic information reflected in the semantic features of the candidate text not only includes the semantic information of the two being independent of each other, but also includes the semantic information when the two are used as the previous and next sentences, that is, the semantic information of the two integrated into a sentence pair.
[0092] After obtaining the semantic features of the candidate text, the confidence of the candidate next sentence can be determined based on the semantic features of the candidate text and the image features of the image to be analyzed. For example, the semantic features and the image features can be fused, and the fused features can be used for confidence analysis.
[0093] The method provided by the embodiments of the present invention enhances the semantic information of each sentence and its candidate next sentence when they are used as the upper and lower sentences by extracting the semantic features of the candidate text obtained by concatenating each sentence and its candidate next sentence, which helps to further improve the reliability of the confidence of the candidate next sentence.
[0094] Based on any of the above embodiments, the image features of the image to be analyzed are determined based on the following steps:
[0095] Based on the text recognition model, feature extraction is performed on the image to be analyzed to obtain the image features. The text recognition model is used to extract the image features and perform text recognition based on the image features.
[0096] Specifically, the text recognition model is the OCR model. The process of applying the text recognition model to perform text recognition on the image to be analyzed is to first extract the image features of the image to be analyzed, and then perform text recognition based on the image features. Here, the image features of the image to be analyzed are the intermediate features in the process of performing text recognition on the image to be analyzed.
[0097] Considering that the text recognition model aims at text recognition, it is inevitable to tend to extract the semantic information contained in the image to be analyzed itself during feature extraction. Therefore, the image features extracted based on the text recognition model can also reflect the semantic information of all the text contents contained in the image to be analyzed as a whole.
[0098] Moreover, in the process of obtaining each sentence in the image to be analyzed, it is necessary to first perform optical character recognition on the image to be analyzed to obtain each character in the image to be analyzed, and then group the characters to obtain each sentence. The image features of the image to be analyzed, as an intermediate process of optical character recognition of the image to be analyzed, can be obtained without relying on additional steps and without occupying additional computing resources. Thus, by combining the image features with the semantic information of the current sentence and the candidate next sentence, a better layout analysis result can be obtained, while optimizing the layout analysis effect and avoiding additional computing expenses.
[0099] Based on any of the above embodiments, Figure 3 is the second flowchart of the method for selecting the next sentence provided by the present invention. As Figure 3 shown, the current sentence and the candidate next sentence can be concatenated to obtain a candidate text, and the candidate text is input into a pre-trained language model (Pre-Trained Language Model) to extract semantic features. In addition, the image to be analyzed is input into an optical character recognition model (OCR) to obtain the image features extracted during the optical character recognition process. Subsequently, the semantic features and the image features can be used to determine the confidence, and the candidate next sentences are sorted according to the determined confidence of each candidate next sentence, so as to obtain the next sentence of the current sentence.
[0100] In this process, the input of the pre-trained language model, that is, the candidate text in the form of the concatenation of the current sentence and the candidate next sentence, can be represented as the following sequence Seq:
[0101] Seq = [Sen, Sen 1 , [SEP], Sen 2 , [SEP], …, Sen N
[0102] where Sen is the current sentence, and Sen 1 , Sen 2 , …, Sen N are N candidate next sentences of the current sentence Sen, and N is an integer greater than 1, such as 5, 10, 12, 20, etc. Each candidate next sentence is concatenated by the [SEP] sentence flag.
[0103] The pre-trained language model extracts features from the input sequence Seq, so as to obtain the vector corresponding to the entire sequence, that is, the semantic feature H = PLM(Seq).
[0104] After obtaining the semantic feature H, the confidence can be determined by combining the semantic feature H with the image feature O obtained by the optical character recognition model. Specifically, the semantic feature H and the image feature O can be concatenated first to obtain the concatenated combined vector HO = contact(H, O).
[0105] Subsequently, the combined vector HO is passed through a fully connected layer to obtain the vector corresponding to each candidate next sentence as follows:
[0106] HI = Linear(HO)
[0107] Here, HI is a vector with the same length as the number of candidate next sentences. Then, the confidence of each candidate next sentence is obtained through the softmax function softmax(HI). Based on this, the candidate next sentence with the highest confidence in the candidate next sentences can be sorted and determined. The next sentence of the current sentence can be expressed as: re = max(softmax(HI)).
[0108] Based on any of the above embodiments, step 110 includes:
[0109] Based on the position of the last character of the current sentence in the image to be analyzed, and the position of the first character of other sentences in the image to be analyzed, determine the distance between the current sentence and the other sentences. The current sentence is one sentence in the sentence set, and the other sentences are the sentences other than the current sentence among all the sentences;
[0110] Based on the distance between the current sentence and the other sentences, determine the candidate next sentence of the current sentence from the other sentences.
[0111] Specifically, for the current sentence, the selection of the candidate next sentence can be determined according to the distribution positions of the current sentence and other sentences in the image to be analyzed. And considering the conventional text sequence arrangement rules, the next sentence of the current sentence is usually arranged after the end of the current sentence, that is, the end of the current sentence is usually close to the position of the beginning of the next sentence. Therefore, in the embodiments of the present invention, the distance between the current sentence and other sentences is determined according to the position of the last character of the current sentence in the image to be analyzed and the position of the first character of other sentences in the image to be analyzed.
[0112] Among them, the position of the last character is the position of the last character of the sentence, and the position of the first character is the position of the first character of the sentence. Based on these two to measure the distance between two sentences, the probability that the two sentences are context sentences can be judged. Specifically, when selecting the candidate next sentence, the first N other sentences sorted in ascending order of distance can be selected. For example, the 10 other sentences with the smallest distance can be selected as the candidate sentences of the current sentence.
[0113] The method provided by the embodiments of the present invention selects the candidate next sentence of the current sentence through the distance between the position of the last character of the current sentence and the position of the first character of other sentences, which narrows the selection range for the selection of the next sentence of the current sentence and helps to improve the efficiency of layout analysis.
[0114] Based on any of the above embodiments, the distance between the current sentence and other sentences, i.e., the distance between the position of the last character of the current sentence and the position of the first character of other sentences, can be further denoted as the difference between the maximum abscissa in the position of the last character of the current sentence and the minimum abscissa in the position of the first character of other sentences, and the difference between the maximum and / or minimum ordinate in the position of the last character of the current sentence and the maximum and / or minimum ordinate in the position of the first character of other sentences.
[0115] Based on any of the above embodiments, Figure 4 is a schematic flowchart of the method for determining a sentence set provided by the present invention. As Figure 4 shown, the sentence set is determined based on the following steps:
[0116] Step 410: Determine the continuous state of the current text and the next text thereof based on the text positions of the current text and the next text in the image to be analyzed.
[0117] Step 420: If the continuous state is continuous, place the next text into the sentence where the current text is located; otherwise, place the next text into a newly created blank sentence.
[0118] Step 430: Use the next text as the new current text until sentence grouping is completed to obtain the sentence set.
[0119] Specifically, after the image to be analyzed is subjected to text recognition, the text positions of each character contained therein can be obtained. Here, for any character, the text position may refer to the coordinate position of the corresponding area of the character in the image to be analyzed, and the corresponding area of the character is usually the minimum bounding box of the character, and its coordinate position can be expressed as the coordinates of the four vertices of the minimum bounding box. For example, the text position of text t can be expressed as:
[0120] token = {t, (x 1 , y 1 , (x 2 , y 2 , (x 3 , y 3 , (x 4 , y 4 )}
[0121] where t corresponds to the character in the dictionary, x and y respectively represent the abscissa and ordinate of text t in the image to be analyzed, and the area corresponding to text t in the image to be analyzed is composed of the positions of four coordinate points (x 1 , y 1 ), (x 2 , y 2 , (x 3 , y 3 , (x 4,y 4 ) for representation.
[0122] After determining the text positions of each text in the image to be analyzed, the texts in the image to be analyzed can be grouped into sentences. In this process, one text among the texts can be used as the current text, that is, the text for which sentence grouping judgment needs to be performed currently, and another arbitrary text can be used as the next text, that is, the text for which it is necessary to determine whether it can be grouped with the current text. Based on the text positions of these two, it can be determined whether the current text and the next text are two consecutive texts from the position level, thereby obtaining the consecutive state of the current text and the next text.
[0123] The consecutive state here can be consecutive or non - consecutive. Among them, when the consecutive state is consecutive, it means that the current text and the next text are two consecutive texts in the same sentence. At this time, the next text can be used as the text following the current text and placed into the sentence where the current text is located, so as to be the consecutive text of the current text. And when the consecutive state is non - consecutive, it means that the current text and the next text are not two consecutive texts, and a new sentence needs to be created for the next text to place the next text.
[0124] After completing the above operations, the next text can be used as the new current text, and a new next text can be selected to perform a new round of sentence grouping until all texts in the image to be analyzed are placed into the corresponding sentences and the sentence grouping is completed.
[0125] The method provided by the embodiments of the present invention combines the texts in the image obtained by text recognition through text positions, thereby realizing fast and reliable sentence grouping and providing a pre - processing operation for layout analysis.
[0126] Based on any of the above - mentioned embodiments, in step 410, determining the consecutive state of the current text and its next text based on the text positions of the current text and the next text in the image to be analyzed includes:
[0127] Determining the consecutive state of the current text and its next text based on the vertex positions of two vertices adjacent to the following text in the text position of the current text and the vertex positions of two vertices adjacent to the preceding text in the text position of the next text.
[0128] Specifically, when determining whether two texts are consecutive, for the text expected to be in the front, that is, the current text, the vertex coordinates of two vertices adjacent to the following text can be selected from the text position of the current text. For the text expected to be in the back, that is, the next text, the vertex positions of two vertices adjacent to the preceding text can be selected from the text position of the next text.
[0129] For example, when writing from left to right, the two vertices adjacent to the following text are the upper right and lower right vertices of the text bounding box, and the two vertices adjacent to the previous text are the upper left and lower left vertices of the text bounding box; for another example, when writing from top to bottom, the two vertices adjacent to the following text are the lower left and lower right vertices of the text bounding box, and the two vertices adjacent to the previous text are the upper left and upper right vertices of the text bounding box.
[0130] By comparing the vertex positions of the two vertices adjacent to the following text of the current text and the vertex positions of the two vertices adjacent to the previous text of the next text, the continuous state between the two can be determined. For example, the distance between the vertex positions of the two vertices adjacent to the following text of the current text and the vertex positions of the two vertices adjacent to the previous text of the next text can be compared with a preset threshold. If the distance is greater than the threshold, the connection state is not connected; otherwise, the connection state is connected.
[0131] Based on any of the above embodiments, the determination of the sentence set can be achieved based on the following steps:
[0132] Assume any text token i As the current text, its text position, that is, the coordinates of the four vertices are expressed as: (x i,1 , y i,1 ), (x i,2 , y i,2 ), (x i,3 , y i,3 ), (x i,4 , y i,4 ), which are the upper left, upper right, lower left, and lower right in sequence.
[0133] In addition, any text token j As the next text, its text position, that is, the coordinates of the four vertices are expressed as: (x j,1 , y j,1 ), (x j,2 , y j,2 ), (x j,3 , y j,3 ), (x j,4 , y j,4 ).
[0134] The continuous state of token i and token j can be judged by the following formula. If they are continuous, then continue to move backward until there are no two consecutive characters, and form sentences with all consecutive two characters:
[0135]
[0136] In the formula, Sen m is the current text token iThe sentence where, |x j,l -x i,l-1 | <thx and|y j,l -y i,l-2 | <thy is a token i and the token j The continuous state judgment condition of, thx and thy are preset thresholds, and l takes values of 3 and 4. If the above conditions are met, then the token j is placed into the sentence where the token i is located, that is, Sen m = Sen m + token j If not satisfied, another blank sentence is created to place the token j , that is, Sen m+1 = token j .
[0137] Through the above method for the image to be analyzed, sentence grouping can be completed, thereby obtaining each sentence in the image to be analyzed. Among them, each sentence is represented by the sentence content and four coordinate regions, that is, Sen i = (x i,1 , y i,1 ), (x i,2 , y i,2 ), (x j,3 , y j,3 ), (x j,4 , y j,4 ), where, (x i,1 , y i,1 ), (x i,2 , y i,2 ) are the upper left and lower left coordinates of the first character in the sentence, and (x j,3 , y j,3 ), (x j,4 , y j,4 ) are the upper right and lower right coordinates of the last character in the sentence.
[0138] Based on any of the above embodiments, Figure 5 This is the second flow diagram of the layout analysis method provided by the present invention. As Figure 5 shown, the method includes:
[0139] First, apply the OCR model to perform character recognition on the image to be analyzed, obtain the image features extracted during the character recognition process, and the position information of each character in the image to be analyzed obtained by the character recognition.
[0140] Based on the position information of each character in the image to be analyzed, perform sentence grouping, thereby obtaining each sentence in the image to be analyzed.
[0141] On this basis, the positions of the last characters of each sentence and the positions of the first characters of other sentences corresponding to each sentence are used to determine the distances between each sentence, so as to screen candidate next sentences for each sentence.
[0142] After that, based on the semantic information of each sentence and its candidate next sentence, as well as the image features of the image to be analyzed, the confidence levels of the candidate next sentences of each sentence are determined, and the confidence levels of the candidate next sentences of each sentence are sorted, so as to determine the next sentence of each sentence.
[0143] On this basis, in combination with the up-and-down sentence relationship between each sentence and the next sentence, each sentence is sorted in layout, so as to restore the original text sequence.
[0144] Based on any of the above embodiments, Figure 6 is a structural schematic diagram of the layout analysis device provided by the present invention, as Figure 6 shown, the device includes:
[0145] A candidate determination unit 610, configured to determine candidate next sentences for each sentence from the set of sentences of the image to be analyzed;
[0146] A next sentence determination unit 620, configured to determine the next sentence of each sentence from the candidate next sentences of each sentence based on the semantic information of each sentence and its candidate next sentence;
[0147] A sorting unit 630, configured to perform layout sorting on the set of sentences based on the next sentence of each sentence.
[0148] The device provided by the embodiment of the present invention judges the up-and-down sentences based on the semantic information of each sentence and its candidate next sentence, so as to determine the next sentence of the sentence from the candidate next sentences, realize the layout sorting of the sentences in the image to be analyzed, and do not need to apply artificially formulated layout sorting rules throughout the process, avoiding the waste of manpower and time caused by artificially specified rules, and being equally applicable to images with structural changes or complex layouts, realizing automated and more adaptable layout analysis, and helping to broaden the application of layout analysis.
[0149] Based on any of the above embodiments, the next sentence determination unit includes:
[0150] A confidence level determination subunit, configured to determine the confidence levels of the candidate next sentences of each sentence based on the semantic information of each sentence and its candidate next sentence, and the image features of the image to be analyzed, where the image features are used to characterize the layout distribution information of each sentence in the image to be analyzed;
[0151] A selection subunit, configured to determine the next sentence of each sentence from the candidate next sentences of each sentence based on the confidence levels of the candidate next sentences of each sentence.
[0152] Based on any of the above embodiments, the confidence determination subunit is configured to:
[0153] Concatenate each of the sentences and its candidate next sentences to obtain candidate texts for each of the sentences;
[0154] Extract semantics from the candidate texts of each of the sentences to obtain semantic features of the candidate texts of each of the sentences, where the semantic features are used to characterize the semantic information of the corresponding sentence and its candidate next sentence;
[0155] Based on the semantic features of the candidate texts of each of the sentences and the image features of the image to be analyzed, determine the confidence of the candidate next sentences of each of the sentences.
[0156] Based on any of the above embodiments, the apparatus further includes a character recognition unit, configured to:
[0157] Extract features from the image to be analyzed based on a character recognition model to obtain the image features, where the character recognition model is used to extract the image features and perform character recognition based on the image features.
[0158] Based on any of the above embodiments, the candidate determination unit is configured to:
[0159] Based on the position of the last character of the current sentence in the image to be analyzed and the position of the first character of other sentences in the image to be analyzed, determine the distance between the current sentence and the other sentences, where the current sentence is one sentence in the sentence set, and the other sentences are sentences in the sentence set other than the current sentence;
[0160] Based on the distance between the current sentence and the other sentences, determine the candidate next sentence of the current sentence from the other sentences.
[0161] Based on any of the above embodiments, the apparatus further includes a sentence composition unit, configured to:
[0162] Based on the character positions of the current character and the next character of the current character in the image to be analyzed, determine the continuous state of the current character and the next character;
[0163] If the continuous state is continuous, place the next character into the sentence where the current character is located; otherwise, place the next character into a newly created blank sentence;
[0164] Use the next character as the new current character until sentence composition is completed to obtain the sentence set.
[0165] Based on any of the above embodiments, the sentence composition unit is configured to:
[0166] Determine the continuous state of the current text and its next text based on the vertex positions of two vertices adjacent to the following text in the text position of the current text and the vertex positions of two vertices adjacent to the preceding text in the text position of the next text.
[0167] Figure 7 An example of a schematic diagram of the physical structure of an electronic device is shown as Figure 7 shown. The electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communications interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call the logical instructions in the memory 730 to execute a layout analysis method, and the method includes:
[0168] Determine the candidate next sentences of each sentence from the sentence set of the image to be analyzed;
[0169] Based on the semantic information of each sentence and its candidate next sentences, determine the next sentence of each sentence from the candidate next sentences of each sentence;
[0170] Based on the next sentence of each sentence, perform layout sorting on the sentence set.
[0171] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0172] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the layout analysis method provided by the above-mentioned various methods. The method includes:
[0173] Determine the candidate next sentences for each sentence from the set of sentences in the image to be analyzed;
[0174] Based on the semantic information of each of the sentences and its candidate next sentences, determine the next sentence for each of the sentences from the candidate next sentences of each of the sentences;
[0175] Based on the next sentence of each of the sentences, perform a layout sorting on the set of sentences.
[0176] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it is configured to execute the layout analysis method provided above, and the method includes:
[0177] Determine the candidate next sentences for each sentence from the set of sentences in the image to be analyzed;
[0178] Based on the semantic information of each of the sentences and its candidate next sentences, determine the next sentence for each of the sentences from the candidate next sentences of each of the sentences;
[0179] Based on the next sentence of each of the sentences, perform a layout sorting on the set of sentences.
[0180] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0181] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A layout analysis method, characterized in that, comprising: determining candidate next sentences for each sentence from a set of sentences in an image to be analyzed; determining the confidence levels of the candidate next sentences for each sentence based on the semantic information of each sentence and its candidate next sentences, and the image features of the image to be analyzed, and determining the next sentence for each sentence from the candidate next sentences for each sentence based on the confidence levels of the candidate next sentences for each sentence, where the image features are used to characterize the layout distribution information of each sentence in the image to be analyzed; performing layout sorting on the set of sentences based on the next sentence for each sentence.
2. The layout analysis method according to claim 1, characterized in that, the determining the confidence levels of the candidate next sentences for each sentence based on the semantic information of each sentence and its candidate next sentences, and the image features of the image to be analyzed, comprises: concatenating each sentence and its candidate next sentence to obtain candidate texts for each sentence; performing semantic extraction on the candidate texts for each sentence to obtain semantic features of the candidate texts for each sentence, where the semantic features are used to characterize the semantic information of the corresponding sentence and its candidate next sentence; determining the confidence levels of the candidate next sentences for each sentence based on the semantic features of the candidate texts for each sentence and the image features of the image to be analyzed.
3. The layout analysis method according to claim 1 or 2, characterized in that, the image features of the image to be analyzed are determined based on the following steps: performing feature extraction on the image to be analyzed based on a text recognition model to obtain the image features, where the text recognition model is used to extract the image features and perform text recognition based on the image features.
4. The layout analysis method according to claim 1 or 2, characterized in that, the determining candidate next sentences for each sentence from a set of sentences in an image to be analyzed, comprises: determining the distance between the current sentence and other sentences based on the position of the last character of the current sentence in the image to be analyzed and the position of the first character of the other sentences in the image to be analyzed, where the current sentence is a sentence in the set of sentences, and the other sentences are sentences in the set of sentences other than the current sentence; determining candidate next sentences for the current sentence from the other sentences based on the distance between the current sentence and the other sentences.
5. The layout analysis method according to claim 1 or 2, characterized in that, the set of sentences is determined based on the following steps: determining the continuous state of the current character and the next character of the current character in the image to be analyzed based on the text positions of the current character and the next character of the current character in the image to be analyzed; if the continuous state is continuous, placing the next character into the sentence where the current character is located, otherwise placing the next character into a newly created blank sentence; using the next character as the new current character until sentence formation is completed to obtain the set of sentences.
6. The layout analysis method according to claim 5, characterized in that, Determining the consecutive state of the current text and its next text based on the text positions of the current text and the next text in the image to be analyzed includes: Determining the consecutive state of the current text and its next text based on the vertex positions of two vertices adjacent to the following text in the text position of the current text and the vertex positions of two vertices adjacent to the preceding text in the text position of the next text.
7. A layout analysis device Characterized in that Comprising: A candidate determination unit for determining candidate next sentences for each sentence from the sentence set of the image to be analyzed; A next sentence determination unit for determining the confidence level of the candidate next sentences for each sentence based on the semantic information of each sentence and its candidate next sentences and the image features of the image to be analyzed, and determining the next sentence for each sentence from the candidate next sentences for each sentence based on the confidence level of the candidate next sentences for each sentence, where the image features are used to characterize the layout distribution information of each sentence in the image to be analyzed; A sorting unit for performing layout sorting on the sentence set based on the next sentence of each sentence.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor Characterized in that When the processor executes the program, the steps of the layout analysis method according to any one of claims 1 to 6 are implemented.
9. A non-transitory computer-readable storage medium, on which a computer program is stored Characterized in that When the computer program is executed by a processor, the steps of the layout analysis method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Document sequence identification method and device, storage medium and electronic equipment
CN110399601A