Intelligent positioning method and device for a page
Through the text image detection model, the problem of low recognition accuracy in the prior art is solved, and the fast and accurate positioning effect is achieved.
Patent Information
- Application Number
- CN202111269270.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-10-29
AI Technical Summary
When existing optical character recognition technology recognizes page text information with watermarks, shadows and tilt angles, the recognition accuracy is low, which affects positioning accuracy.
The text image detection model directly identifies the text information of the target page, and then locates the target page of the key page after integrating the text information, improving the accuracy of recognition and positioning.
There is no need to manually identify the text information on the target page, improve the recognition rate and accuracy of the text information, and quickly and accurately locate the text information of the target page corresponding to the key page.
Smart Images

Figure CN114140800B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to an intelligent page positioning method and device. Background Art
[0002] With the transformation of enterprise digitization, artificial intelligence technology has gradually become an important driving force for enterprise industrial transformation and digital development. In the review work of project contract scanned copies within an enterprise, artificial intelligence technology is often combined, and through artificial intelligence technology, problems such as high labor costs and long review cycles caused by previous pure manual reviews can be well solved, thus ensuring the improvement of enterprise economic benefits.
[0003] Currently, for the review of key information in enterprise contract scanned copies, generally, optical character recognition technology is first used to directly recognize the text information in the scanned copies, and then the keyword regular matching method is used to locate the key pages in the scanned copies, which helps to quickly find the key pages of the scanned copies that need to be reviewed. However, it is found in practice that the current optical character recognition technology has a low recognition accuracy when recognizing the text information of pages with watermarks, shadows, and tilt angles, which affects the positioning accuracy of the key pages of the subsequent scanned copies. Therefore, it is particularly important to improve the recognition accuracy of the key pages of the scanned copies. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method and device for intelligent page positioning, which can directly recognize the text information of the text images in the target pages through a text image detection model, and after determining the integrated text information of the target pages, further locate the integrated text information of the target pages corresponding to the key pages. In this way, it is not necessary to manually recognize the text information of the target pages, thereby improving the recognition rate and recognition accuracy of the text information of the target pages, and thus quickly and accurately locating the text information of the target pages corresponding to the key pages.
[0005] To solve the above technical problems, a first aspect of the present invention discloses an intelligent page positioning method, and the method includes:
[0006] Detecting a plurality of target pages according to the determined text image detection model and detection element information, and obtaining at least one text image corresponding to each target page; the detection element information includes the detection identifier of each target page, and there is a corresponding original page for each target page;
[0007] Performing a text recognition operation on each text image corresponding to each target page according to the determined text recognition model, and obtaining the text information of each text image corresponding to each target page;
[0008] Determine the integrated text information of each target page according to the text information of all the text images corresponding to each target page.
[0009] As an optional implementation manner, in the first aspect of the present invention, after determining the integrated text information of each target page according to the text information of all the text images corresponding to each target page, the method further includes:
[0010] Calculate the text edit distance between the integrated text information of all the target pages and the text information of at least one key page determined; all the key pages are derived from the original pages corresponding to all the target pages;
[0011] According to the text edit distance between the integrated text information of all the target pages and the text information of any one of the key pages, screen out the integrated text information of all the target pages whose text edit distance is less than or equal to a preset text edit distance threshold from the integrated text information of all the target pages, as the text information of the target output page corresponding to this key page; all the target pages include the target output page.
[0012] As an optional implementation manner, in the first aspect of the present invention, the detecting of multiple target pages according to the determined text image detection model and detection element information to obtain at least one text image corresponding to each target page includes:
[0013] Input the multiple target pages and the detection element information into the determined text image detection model for analysis to obtain the position information of at least one target detection region corresponding to each target page;
[0014] According to the position information of each target detection region corresponding to each target page, extract the text image of this target detection region from this target page and use it as at least one text image corresponding to this target page.
[0015] As an optional implementation manner, in the first aspect of the present invention, the performing of a text recognition operation on each text image corresponding to each target page according to the determined text recognition model to obtain the text information of each text image corresponding to each target page includes:
[0016] Input each text image corresponding to each target page into the determined text recognition model to extract the text features of each text image corresponding to each target page;
[0017] Fuse the text features of each text image corresponding to each target page to obtain the fused text features of each text image corresponding to each target page;
[0018] Perform sequence analysis on the fused text features of each of the text images corresponding to each of the target pages to obtain a fused text feature sequence for each of the text images corresponding to each of the target pages;
[0019] Perform text recognition operations on the fused text feature sequences of each of the text images corresponding to each of the target pages to obtain the text information of each of the text images corresponding to each of the target pages.
[0020] As an alternative implementation manner, in the first aspect of the present invention, the determining the integrated text information of each of the target pages according to the text information of all of the text images corresponding to each of the target pages includes:
[0021] Sort the text information of all of the text images corresponding to each of the target pages according to the position information of all of the target detection regions corresponding to each of the target pages and the determined sorting factor information to obtain a sorting result of the text information corresponding to each of the target pages;
[0022] Integrate the sorting result of the text information corresponding to each of the target pages according to a pre-determined text information integration method to obtain the integrated text information of each of the target pages.
[0023] As an alternative implementation manner, in the first aspect of the present invention, the text information of all of the key pages is determined by the following method:
[0024] Calculate the matching degree between the text information of at least one of the original pages and a preset feature condition;
[0025] According to the matching degree between the text information of all of the original pages and the preset feature condition, screen out the text information of all of the original pages whose matching degree is greater than or equal to a preset matching degree threshold from the text information of all of the original pages as the text information of all of the key pages;
[0026] And, the method further includes:
[0027] Determine the number of pages of the target output page corresponding to any one of the key pages according to the text information of the target output page corresponding to the key page, and determine whether the number of pages of the target output page is greater than a preset number of pages threshold;
[0028] When it is determined that the number of pages of the target output page is greater than the preset number of pages threshold, perform a change operation on the detection factor information to obtain the changed detection factor information;
[0029] Update the detection element information to the changed detection element information to trigger the operation of re - executing the detection of all the target pages according to the text image detection model and the changed detection element information.
[0030] As an optional implementation manner, in the first aspect of the present invention, the method further includes:
[0031] For any target detection area corresponding to any target page, determine whether the rotation angle of the text image is greater than a preset rotation angle threshold according to the position information of the text image in the target detection area;
[0032] When it is determined that the tilt angle of the text image is greater than the preset rotation angle threshold, correct the rotation angle of the text image to obtain the corrected position information of the text image.
[0033] The second aspect of the present invention discloses an intelligent positioning device for a page. The device includes:
[0034] A detection module, configured to detect multiple target pages according to the determined text image detection model and detection element information, and obtain at least one text image corresponding to each target page; the detection element information includes the detection identifier of each target page, and there is a corresponding original page for each target page;
[0035] An identification module, configured to perform a text recognition operation on each text image corresponding to each target page according to the determined text recognition model, and obtain the text information of each text image corresponding to each target page;
[0036] A determination module, configured to determine the integrated text information of each target page according to the text information of all the text images corresponding to each target page.
[0037] As an optional implementation manner, in the second aspect of the present invention, the device further includes:
[0038] A calculation module, configured to calculate the text edit distance between the integrated text information of all the target pages and the text information of at least one determined key page after the determination module determines the integrated text information of each target page according to the text information of all the text images corresponding to each target page; all the key pages are from the original pages corresponding to all the target pages;
[0039] A screening module, configured to screen out, from the integrated text information of all the target pages, the integrated text information of all the target pages whose text edit distance from the text information of any one of the key pages is less than or equal to a preset text edit distance threshold, as the text information of the target output page corresponding to this key page; all the target pages include the target output page.
[0040] As an alternative implementation manner, in the second aspect of the present invention, the manner in which the detection module detects multiple target pages according to the determined text image detection model and detection element information to obtain at least one text image corresponding to each target page is specifically as follows:
[0041] Input the multiple target pages and the detection element information into the determined text image detection model for analysis to obtain the position information of at least one target detection region corresponding to each target page;
[0042] According to the position information of each target detection region corresponding to each target page, extract the text image of this target detection region from this target page and use it as at least one text image corresponding to this target page.
[0043] As an alternative implementation manner, in the second aspect of the present invention, the manner in which the recognition module performs a text recognition operation on each text image corresponding to each target page according to the determined text recognition model to obtain the text information of each text image corresponding to each target page is specifically as follows:
[0044] Input each text image corresponding to each target page into the determined text recognition model to extract the text features of each text image corresponding to each target page;
[0045] Fuse the text features of each text image corresponding to each target page to obtain the fused text features of each text image corresponding to each target page;
[0046] Perform a sequence analysis on the fused text features of each text image corresponding to each target page to obtain the fused text feature sequence of each text image corresponding to each target page;
[0047] Perform a text recognition operation on the fused text feature sequence of each text image corresponding to each target page to obtain the text information of each text image corresponding to each target page.
[0048] As an alternative implementation, in the second aspect of the present invention, the specific manner in which the determination module determines the integrated text information of each target page according to the text information of all the text images corresponding to each target page is as follows:
[0049] Sort the text information of all the text images corresponding to each target page according to the position information of all the target detection regions corresponding to each target page and the determined sorting factor information, to obtain the sorting result of the text information corresponding to each target page;
[0050] Integrate the sorting result of the text information corresponding to each target page according to the pre-determined text information integration method, to obtain the integrated text information of each target page.
[0051] As an alternative implementation, in the second aspect of the present invention, the text information of all the key pages is determined by the following method:
[0052] Calculate the matching degree between the text information of at least one of the original pages and a preset feature condition;
[0053] According to the matching degree between the text information of all the original pages and the preset feature condition, screen out the text information of all the original pages whose matching degree is greater than or equal to a preset matching degree threshold from the text information of all the original pages, as the text information of all the key pages;
[0054] Moreover, the device further includes:
[0055] The determination module is further configured to determine the number of pages of the target output page corresponding to any one of the key pages according to the text information of the target output page corresponding to the key page;
[0056] The judgment module is configured to judge whether the number of pages of the target output page is greater than a preset number-of-pages threshold;
[0057] The change module is configured to, when the judgment module judges that the number of pages of the target output page is greater than the preset number-of-pages threshold, perform a change operation on the detection factor information to obtain the changed detection factor information;
[0058] The update module is configured to update the detection factor information to the changed detection factor information, so as to trigger the detection module to re-execute the operation of detecting all the target pages according to the text image detection model and the changed detection factor information.
[0059] As an alternative implementation, in the second aspect of the present invention, the device further includes:
[0060] The determination module is further configured to, for any one of the target detection regions corresponding to any one of the target pages, determine whether the rotation angle of the text image is greater than a preset rotation angle threshold according to the position information of the text image in the target detection region;
[0061] The correction module is configured to correct the rotation angle of the text image to obtain the corrected position information of the text image when the determination module determines that the tilt angle of the text image is greater than the preset rotation angle threshold.
[0062] A third aspect of the present invention discloses another intelligent positioning device for a page, and the device includes:
[0063] A memory storing executable program code;
[0064] A processor coupled to the memory;
[0065] The processor calls the executable program code stored in the memory and executes the intelligent positioning method for a page disclosed in the first aspect of the present invention.
[0066] A fourth aspect of the present invention discloses a computer storage medium, and the computer storage medium stores computer instructions, which are used to execute the intelligent positioning method for a page disclosed in the first aspect of the present invention when called.
[0067] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0068] In the embodiments of the present invention, the text information of the text image in the target page is directly recognized through the text image detection model, and after determining the integrated text information of the target page, the integrated text information of the target page corresponding to the key page is further located, which can improve the recognition rate and recognition accuracy of the text information of the target page, reduce the workload of finding the text information of the target page corresponding to the key page, and thus quickly and accurately locate the text information of the target page corresponding to the key page. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0070] Figure 1 It is a schematic flowchart of an intelligent positioning method for a page disclosed in an embodiment of the present invention;
[0071] Figure 2It is a schematic flowchart of another intelligent page positioning method disclosed in an embodiment of the present invention;
[0072] Figure 3 It is a schematic structural diagram of an intelligent page positioning device disclosed in an embodiment of the present invention;
[0073] Figure 4 It is a schematic structural diagram of another intelligent page positioning device disclosed in an embodiment of the present invention;
[0074] Figure 5 It is a schematic structural diagram of yet another intelligent page positioning device disclosed in an embodiment of the present invention. Detailed implementation manners
[0075] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0076] The terms "first", "second", etc. in the description and claims of the present invention and the above drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or terminal that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes unlisted steps or units, or optionally further includes other steps or units inherent to these processes, methods, products or terminals.
[0077] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0078] The present invention discloses an intelligent positioning method and device for pages, which can identify text information of multiple target pages through a text image detection model and determine the integrated text information of the target pages, and then locate the integrated text information of the corresponding target pages for key pages. In this way, it is not necessary to manually identify the text information of the target pages, which is beneficial to shortening the time for identifying text information and improving the identification accuracy, so as to quickly and accurately locate the text information of the target pages corresponding to the key pages according to the identified text information. The following will be described in detail respectively.
[0079] Embodiment 1
[0080] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an intelligent positioning method for pages disclosed in an embodiment of the present invention. Optionally, Figure 1 the described intelligent positioning method for pages can be applied to the positioning work of the review pages required in the project contract scan, can also be applied to the positioning work of the key pages of the specification scan, and can also be applied to the positioning work of the key information of the web page image. Further optionally, this method can be applied to the positioning of key pages of files such as Portable Document Format (PDF), Tagged Image File Format (TIFF), Graphics Interchange Format (GIF), etc., and the embodiments of the present invention do not make limitations. Still further, this method can be implemented by an image processing system, and the image processing system can be integrated in an image processing device, which can be a local server or a cloud server for managing the image processing process, etc., and the embodiments of the present invention do not make limitations. As Figure 1 shown, the intelligent positioning method for pages can include the following operations:
[0081] 101. Detect multiple target pages according to the determined text image detection model and detection element information, and obtain at least one text image corresponding to each target page.
[0082] In an embodiment of the present invention, after receiving the text image detection model and the detection element information input by relevant business personnel, multiple target pages can be detected. Optionally, the text image detection model can be a CTPN network model, that is, a text image detection model that combines a CNN network model and an RNN network model, which can detect text images containing text information in the target pages and mark the text images in the form of text boxes as detection regions. Further optionally, the detection element information can include the detection identifiers of each target page. Specifically, it can include at least one detection identifier of the specific text content (such as the transaction amount of the project contract, the signing date of the project contract, the title of the web page image) of each target page to be detected, the location information of the target page (such as the home page, the end page), and the image identifier information of the target page (such as the contract seal pattern, the web page QR code image, the instruction flow chart).
[0083] Furthermore, each target page can correspond to one or more text images. Specifically, for any detection region with a detection identifier on any target page, the text image in this detection region of this target page can be detected, that is, the detection region corresponds one-to-one with the text image. Still further, there is a corresponding original page for each target page. For example, after the original page of the project contract is scanned, the scanned page of the project contract can be obtained, which can be used as the target page of the project contract; and after the original web page is screenshot, the web page screenshot page can be obtained, which can be used as the target page of the web page.
[0084] 102. Perform a text recognition operation on each text image corresponding to each target page according to the determined text recognition model to obtain the text information of each text image corresponding to each target page.
[0085] In an embodiment of the present invention, after receiving the text recognition model input by relevant business personnel, the text information in each text image corresponding to each target page can be recognized. Optionally, the text recognition model can be a CRNN+CTC model, that is, a text recognition model that combines a CNN network model, an RNN network model, and a CTC network model, which has a high recognition accuracy for text images with less complex backgrounds. Specifically, it can extract text features of multiple text images of each target page simultaneously or in the determined text recognition order, and then fuse the text features of each extracted text image, and then obtain the character sequence of each text image, and finally perform transcription processing on the character sequence of each text image to obtain the text information. It should be noted that the text information corresponding to the text image in any detection region of any target page can be recognized, that is, this detection region of this target page also corresponds one-to-one with the recognized text information.
[0086] 103. Determine the integrated text information of each target page according to the text information of all text images corresponding to each target page.
[0087] In the embodiments of the present invention, after sorting the text information of all text images corresponding to each target page according to the position information of the detection area and the determined sorting method, integration can be performed to obtain the integrated text information of each target page. Optionally, the text information of all text images corresponding to each target page can be sorted according to the position information of the detection area and the sorting method of the text information of the original page, or the text information of all text images corresponding to each target page can be sorted according to specific requirements.
[0088] Further optionally, after sorting the text information of all text images corresponding to each target page, the text information of all text images corresponding to each target page can be integrated according to the determined text information integration method. Optionally, the text information of all text images corresponding to the target page can be integrated according to the rule of splicing the head and tail of each line of text from top to bottom, or according to the rule of splicing the head and tail of each column of text from left to right. For example, the text information of the original page is usually composed according to the rule of splicing the head and tail of each line of text from top to bottom, and the text information of all text images corresponding to each target page can also be integrated according to this text information integration method. In this way, the accuracy of the subsequent comparison of the two text information can be ensured by the assistance of the text information of the original page.
[0089] It can be seen that implementing the present invention can detect text images of multiple target pages through a text image detection model and a detection identifier, and after identifying the text information of each text image corresponding to each target page, sort and integrate the text information in all text images corresponding to each target page, so as to determine the integrated text information corresponding to each target page. In this way, it is possible to avoid manually identifying the text information of the target page, improve the speed of identifying the text information in the text image, and improve the recognition accuracy of the text information of the target page by comparing with the text information of the original page, so as to correctly locate the text information of the target page corresponding to the key page in the original page.
[0090] In an optional implementation, after performing the text recognition operation on each text image corresponding to each target page according to the determined text recognition model in step 102 above to obtain the text information of each text image corresponding to each target page, the method may further include:
[0091] For the text information of any text image corresponding to any target page, determine the weight of the text information of the text image, and make a matching display mark for the text information of the text image according to the weight.
[0092] In this alternative embodiment, the key text information in each target page can be determined according to the weights of the text information of each text image corresponding to each target page, and various display marks can be made. Optionally, the display mark can be a font color different from that of the non-key text information, or a fluorescent display mark at the bottom of the key text information, or a text box mark for the key text information.
[0093] It can be seen that this alternative embodiment can directly mark the key text information of any text image corresponding to any target page, so as to speed up the subsequent review work of relevant staff for this key text information.
[0094] Embodiment Two
[0095] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another intelligent page positioning method disclosed in the embodiments of the present invention. Optionally, Figure 2 the described intelligent page positioning method can be applied to the positioning work of the review pages required in the project contract scan, or to the positioning work of the key pages of the specification scan, or to the positioning work of the key information of the web page image. Further optionally, this method can be applied to the positioning of the key pages of files such as portable document format, tag image format, and graphics interchange format, which is not limited in the embodiments of the present invention. Further still, this method can be implemented by an image processing system, which can be integrated in an image processing device, and can be a local server or a cloud server for managing the image processing process, etc., which is not limited in the embodiments of the present invention. As Figure 2 shown, the intelligent page positioning method can include the following operations:
[0096] 201. Detect multiple target pages according to the determined text image detection model and detection element information, and obtain at least one text image corresponding to each target page.
[0097] 202. Perform text recognition operations on each text image corresponding to each target page according to the determined text recognition model, and obtain the text information of each text image corresponding to each target page.
[0098] 203. Determine the integrated text information of each target page according to the text information of all text images corresponding to each target page.
[0099] In the embodiments of the present invention, for other descriptions of steps 201 - 203, please refer to the detailed descriptions of steps 101 - 103 in Embodiment One, and the embodiments of the present invention will not be repeated here.
[0100] 204. Calculate the text edit distance between the integrated text information of all target pages and the text information of at least one determined key page.
[0101] In an embodiment of the present invention, the text edit distance matrix E of the integrated text information of all target pages and the text information of each key page can be obtained by calculating the text edit distance between the integrated text information of all target pages and the text information of one or more determined key pages. Among them, all key pages can be obtained by screening the original pages. Optionally, all the determined key pages can be pages with detection identifiers, such as the project contract transaction amount, the project contract signing date, the web page image title, the contract seal pattern, the web page QR code image, the instruction flow chart, etc. Further optionally, the text edit distance matrix E can be constructed with the number of pages of the target page as the number of rows and the key pages as the number of columns. Further, in the order of the page numbers of the target page and the key page, the element E at each position in the text edit distance matrix E ij can be obtained according to the calculated text edit distance between the integrated text information of the corresponding target page and the text information of the key page. The calculation formula is:
[0102] E ij = edit_dist(Ω i , Γ j )
[0103] Among them, E ij represents the element value of the i-th row and j-th column in the matrix E, edit_dist(*) represents the edit distance, Ω i represents the integrated text information of the i-th page of the target page, and Γ j represents the text information of the j-th key page. In this way, through the text edit distance matrix formed by the integrated text information of all target pages and the text information of each key page, it is beneficial to compare the global text information of the pages, can reduce the risk of positioning errors caused by misidentifying the text information of the target page, and improve the robustness of positioning.
[0104] 205. According to the text edit distance between the integrated text information of all target pages and the text information of any key page, screen out the integrated text information of all target pages with a text edit distance less than or equal to a preset text edit distance threshold from the integrated text information of all target pages, as the text information of the target output page corresponding to this key page.
[0105] In the embodiments of the present invention, the text information of one or more target pages corresponding to the key page can be located by screening the integrated text information of all target pages, and used as the text information of the target output page corresponding to the key page. Optionally, the text information of the target page most similar to the key page can be screened out according to the principle of the minimum text edit distance, or a numerical range of the text edit distance can be set to screen out the text information of multiple target pages similar to the key page. For example, when screening out the text information of the target page most similar to the key page according to the principle of the minimum text edit distance, for each column in the text edit distance matrix E, the row number where the minimum value in the column is located can be located, so as to locate the specific page in the key page. After all columns are processed, the specific pages of all target pages corresponding to the key page can be obtained, and then the text information of the target page corresponding to the key page can be output.
[0106] It can be seen that implementing the present invention can form a comparison of the global text information of the two pages by means of the text edit distance matrix formed by the integrated text information of all target pages and the text information of each key page, and match the text information of one or more target pages similar to the key page. It can not only reduce the positioning error rate caused by the misrecognition of the text information of the target page and improve the robustness of the positioning, but also perform flexible positioning according to specific screening requirements.
[0107] In an alternative embodiment, detecting multiple target pages according to the determined text image detection model and detection element information in step 201 above to obtain at least one text image corresponding to each target page may include:
[0108] Inputting the multiple target pages and the detection element information into the determined text image detection model for analysis to obtain the position information of at least one target detection area corresponding to each target page;
[0109] According to the position information of each target detection area corresponding to each target page, extracting the text image of the target detection area from the target page and using it as at least one text image corresponding to the target page.
[0110] In this optional embodiment, multiple target pages and detection element information can be input into a text image detection model (such as a CTPN network model) for text image detection, so as to obtain all the detection regions containing text images in each target page, and obtain the position information of each detection region in the target page. Specifically, the depth features of the input target page can be extracted first, and then regions containing text images in each target page are detected using anchors with a fixed width. The features corresponding to the anchors in the same row are concatenated into a sequence, and then a fully connected layer is used to classify or regress the sequence to obtain each pre-detection region of each target page. Finally, each pre-detection region of each target page is merged to obtain the position information of each target detection region of each target page. For example, the position information of the target detection region may include the center point coordinates (x c , y c ) of the target detection region, height h, width w, and rotation angle θ, which can be specifically expressed as: (x c , y c , h, w, θ). Further optionally, for each target detection region, one or more text images of the target detection region can be extracted according to the position information of the target detection region.
[0111] It can be seen that this optional embodiment can, according to specific requirements, target the target detection regions containing text images in each target page to be detected, which is beneficial to improving the detection efficiency and flexibility of text images in the target page, and further improving the reliability of the extracted text images.
[0112] In another optional embodiment, the step of performing text recognition operations on each text image corresponding to each target page according to the determined text recognition model in step 202 to obtain the text information of each text image corresponding to each target page may include:
[0113] Input each text image corresponding to each target page into the determined text recognition model to extract the text features of each text image corresponding to each target page;
[0114] Fuse the text features of each text image corresponding to each target page to obtain the fused text features of each text image corresponding to each target page;
[0115] Perform sequence analysis on the fused text features of each text image corresponding to each target page to obtain the fused text feature sequence of each text image corresponding to each target page;
[0116] Perform text recognition operations on the fused text feature sequence of each text image corresponding to each target page to obtain the text information of each text image corresponding to each target page.
[0117] In this alternative embodiment, a text recognition model of CRNN+CTC can be used to perform text recognition operations on each text image corresponding to each target page. Specifically, the text features of each text image corresponding to each target page can be extracted first through the CNN network model in the text recognition model, and then the text features of each text image corresponding to each target page can be fused into feature vectors through the RNN network model in the text recognition model. Furthermore, the character sequence of each text image corresponding to each target page can be extracted based on the fused text features of each text image corresponding to each target page. Finally, the character sequence of each text image corresponding to each target page can be transcribed according to the context features of the character sequence of each text image corresponding to each target page and the CTC network model in the text recognition model to obtain the text information of each text image corresponding to each target page.
[0118] It can be seen that this alternative embodiment can identify the text information of each text image corresponding to each target page according to the text feature distribution of each text image corresponding to each target page, which is beneficial to improving the recognition accuracy and reliability of the text information of the text image, so that the positioning operation can be performed according to the text information of the correct target page.
[0119] In another alternative embodiment, determining the integrated text information of each target page according to the text information of all text images corresponding to each target page in step 203 above may include:
[0120] Sorting the text information of all text images corresponding to each target page according to the position information of all target detection regions corresponding to each target page and the determined sorting element information to obtain the sorting result of the text information corresponding to each target page;
[0121] Integrating the sorting result of the text information corresponding to each target page according to the pre-determined text information integration method to obtain the integrated text information of each target page.
[0122] In this alternative embodiment, the position information of all target detection regions corresponding to each detected target page and the determined sorting element information (such as the sorting method of the text information of the original page) can be used to sort the text information of all text images corresponding to each target page to obtain the sorting result of the text information corresponding to each target page. Optionally, for the sorting result of the text information corresponding to each target page, the sorting result can be integrated according to the row text merging algorithm or the column text merging algorithm.
[0123] For example, when using the line text merging algorithm for integration, the sorting result of the text information corresponding to each target page can be input into the line text merging algorithm for analysis to obtain the position information (x c , y c , h, w, θ) of each target detection area corresponding to each target page and the set W of the text information ω of each target detection area corresponding to each target page, that is, W = {(x ic , y ic , h i , w i , θ i , ω i ) | 1 ≤ i ≤ M}.
[0124] Among them, the specific processing flow of using the line text merging algorithm for integration can be as follows:
[0125] (1) Initialize the stack S as an empty stack, and the counter t = 1;
[0126] (2) Create an empty stack L t and push it onto the stack S. Initialize the upper and lower bounds y top = 0 and y bottom = 0, and proceed to step (3);
[0127] (3) Traverse the elements in the set W. For the current element W i = (x ic , y ic , h i , w i , θ i , ω i ), if L t is an empty stack, then push the element W i onto the stack L t , and set y top = y ic + h / 2, y bottom = y ic - h / 2; otherwise, proceed to step (4). When the traversal is completed, set t = t + 1 and return to step (2); where, i in W i refers to the number of elements.
[0128] (4) If the current element W i = (x ic , y ic , h i , w i , θ i , ω i ) satisfies: y ic + h / 2 < y top or y ic - h / 2 > y bottom , then push the element Wi Push L onto the stack t , set y top = max(y top , y ic + h / 2), y bottom = min(y bottom , y ic - h / 2), and remove the element W i from the set W. If the set then go to step (5); otherwise return to step (3) and continue traversing.
[0129] (5) For each stack L in the stack S t , for each element W in it i sort them in ascending order according to the x - axis coordinate x of the center point ic , and concatenate the text information ω of the sorted elements i in sequence to form the text information Ω t of the stack L t , calculate the average y - axis coordinate of the center points of the stack L t where N represents the number of elements in the stack L t . t
[0130] (6) Sort all the stacks in the stack S in ascending order according to the average y - axis coordinate y t of the center points, and concatenate the text information Ω t of the sorted stacks in sequence to form the integrated text information of the entire target page.
[0131] It can be seen that this alternative embodiment can integrate the text information in all target detection regions corresponding to each target page to obtain the integrated text information corresponding to each target page. In this way, all local text information in each target page can be flexibly processed, thereby improving the reliability and accuracy of the integrated text information corresponding to each target page.
[0132] In yet another alternative embodiment, the method may further include:
[0133] Determine the number of pages of the target output page corresponding to any key page according to the text information of the target output page corresponding to the key page, and determine whether the number of pages of the target output page is greater than a preset page number threshold;
[0134] When it is determined that the number of pages of the target output page is greater than the preset page number threshold, perform a change operation on the detection element information to obtain the changed detection element information;
[0135] Update the detection element information to the changed detection element information to trigger the operation of re-executing the step of detecting all target pages according to the text image detection model and the changed detection element information in the above step 201;
[0136] Among them, the text information of all key pages can be determined by the following methods:
[0137] Calculate the matching degree between the text information of at least one original page and the preset feature conditions;
[0138] According to the matching degree between the text information of all original pages and the preset feature conditions, screen out the text information of all original pages with a matching degree greater than or equal to the preset matching degree threshold from the text information of all original pages as the text information of all key pages.
[0139] In this optional embodiment, by judging the number of pages of the target output page corresponding to any key page and the change of the detection element, the text information of the target output page with a higher similarity to the key page can be further screened out. Optionally, the preset feature conditions can be input according to actual needs, and then the matching degree between the text information of all original pages and the preset feature conditions can be calculated, so as to screen and determine the text information of all key pages from the text information of all original pages. Specifically, since the page distribution (such as the sealed signature page, contract amount page, page flow chart, etc.) and the text information included (such as contract amount, page title, etc.) of each key page are different, the matching feature conditions can be set first to screen out the key pages. For example, a regular expression for describing the features of the required key pages can be formulated, and the text information of specific key pages can be matched from the text information of the original pages through this regular expression.
[0140] It can be seen that this optional embodiment can further narrow the number of pages of the target output page corresponding to the key page by adjusting the number of pages of the target output page, so as to further screen out the target output page with a higher similarity to the key page. In this way, it is beneficial to further improve the recognition accuracy of the text information of the target output page, thereby improving the positioning accuracy of the target output page corresponding to the key page.
[0141] In another optional embodiment, the method may further include:
[0142] For any target detection area corresponding to any target page, judge whether the rotation angle of the text image is greater than the preset rotation angle threshold according to the position information of the text image in the target detection area;
[0143] When it is judged that the tilt angle of the text image is greater than the preset rotation angle threshold, correct the rotation angle of the text image to obtain the corrected position information of the text image.
[0144] In this optional embodiment, the rotation angle of the text image can be corrected according to the position information of the text image in each target detection area. For example, the distribution of text images on each target page is different, and for a text image whose rotation angle exceeds the preset rotation angle range, there may be a situation where the text information is misrecognized. By correcting the rotation angle, the recognition rate of the text information of the text image can be improved.
[0145] It can be seen that this optional embodiment can improve the recognition accuracy of the text information of the text image through the correction of the position information of the text image, thereby improving the reliability of the text information of the text image.
[0146] Embodiment III
[0147] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of an intelligent page positioning device disclosed in an embodiment of the present invention. As Figure 3 shown, the intelligent page positioning device may include:
[0148] A detection module 301, configured to detect a plurality of target pages according to the determined text image detection model and detection element information, and obtain at least one text image corresponding to each target page;
[0149] An identification module 302, configured to perform a text recognition operation on each text image corresponding to each target page according to the determined text recognition model, and obtain the text information of each text image corresponding to each target page;
[0150] A determination module 303, configured to determine the integrated text information of each target page according to the text information of all text images corresponding to each target page.
[0151] In an embodiment of the present invention, optionally, the detection element information may include a detection identifier of each target page; further optionally, there is a corresponding original page for each target page.
[0152] It can be seen that the implementation Figure 3The intelligent positioning device for the described page can detect text images of multiple target pages through a text image detection model and detection identifiers. After identifying the text information of each text image corresponding to each target page, it sorts and integrates the text information in all text images corresponding to each target page, so as to determine the integrated text information corresponding to each target page. In this way, it is possible to identify the text information of the target page without manual intervention, improve the rate of identifying text information in text images, and improve the recognition accuracy of the text information of the target page by comparing it with the text information of the original page, so as to correctly locate the text information of the target page corresponding to the key page in the original page.
[0153] In an optional embodiment, the device may further include:
[0154] A calculation module 304, configured to calculate the text edit distance between the integrated text information of all target pages and the text information of at least one determined key page after the determination module 303 executes the above-mentioned step of determining the integrated text information of each target page according to the text information of all text images corresponding to each target page;
[0155] A screening module 305, configured to screen out the integrated text information of all target pages whose text edit distance is less than or equal to a preset text edit distance threshold from the integrated text information of all target pages according to the text edit distance between the integrated text information of all target pages and the text information of any key page, as the text information of the target output page corresponding to this key page.
[0156] In this optional embodiment, optionally, all key pages are derived from the original pages corresponding to all target pages; further optionally, all target pages include target output pages.
[0157] It can be seen that implementing Figure 4 The intelligent positioning device for the described page can form a comparison of the global text information of the two pages through the text edit distance matrix formed by the integrated text information of all target pages and the text information of each key page, and match the text information of one or more target pages similar to the key page. It can not only reduce the positioning error rate caused by incorrect recognition of the text information of the target page and improve the robustness of positioning, but also perform flexible positioning according to specific screening requirements.
[0158] In another optional embodiment, the specific manner in which the detection module 301 detects multiple target pages according to the determined text image detection model and detection element information to obtain at least one text image corresponding to each target page is as follows:
[0159] Input multiple target pages and detection element information into the determined text image detection model for analysis to obtain the position information of at least one target detection area corresponding to each target page;
[0160] According to the position information of each target detection area corresponding to each target page, extract the text image of the target detection area from the target page and use it as at least one text image corresponding to the target page.
[0161] It can be seen that implementing Figure 4 The intelligent positioning device for the described page can, according to specific requirements, targetedly locate the target detection areas containing text images in each target page to be detected, which is beneficial to improving the detection efficiency and flexibility of the text images on the target page, and further improving the reliability of the extracted text images.
[0162] In another optional embodiment, the manner in which the above recognition module 302 performs text recognition operations on each text image corresponding to each target page according to the determined text recognition model to obtain the text information of each text image corresponding to each target page is specifically as follows:
[0163] Input each text image corresponding to each target page into the determined text recognition model to extract the text features of each text image corresponding to each target page;
[0164] Fuse the text features of each text image corresponding to each target page to obtain the fused text features of each text image corresponding to each target page;
[0165] Perform sequence analysis on the fused text features of each text image corresponding to each target page to obtain the fused text feature sequence of each text image corresponding to each target page;
[0166] Perform text recognition operations on the fused text feature sequences of each text image corresponding to each target page to obtain the text information of each text image corresponding to each target page.
[0167] It can be seen that implementing Figure 4 The intelligent positioning device for the described page can identify the text information of each text image corresponding to each target page according to the text feature distribution of each text image corresponding to each target page, which is beneficial to improving the recognition accuracy and reliability of the text information of the text images, so that the positioning operation can be performed according to the correct text information of the target page.
[0168] In another optional embodiment, the manner in which the above determination module 303 determines the integrated text information of each target page according to the text information of all text images corresponding to each target page is specifically as follows:
[0169] Sort the text information of all text images corresponding to each target page according to the position information of all target detection areas corresponding to each target page and the determined sorting factor information, and obtain the sorting result of the text information corresponding to each target page;
[0170] Integrate the sorting results of the text information corresponding to each target page according to the pre-determined text information integration method to obtain the integrated text information of each target page.
[0171] It can be seen that implementing Figure 4 The described intelligent positioning device for pages can integrate the text information in all target detection areas corresponding to each target page to obtain the integrated text information corresponding to each target page. In this way, all local text information in each target page can be flexibly processed, thereby improving the reliability and accuracy of the integrated text information corresponding to each target page.
[0172] In another optional embodiment, the device may further include:
[0173] The determination module 303 is further configured to determine the number of pages of the target output page corresponding to any key page according to the text information of the target output page corresponding to the key page;
[0174] The judgment module 306 is configured to judge whether the number of pages of the target output page is greater than a preset page number threshold;
[0175] The change module 307 is configured to perform a change operation on the detection factor information to obtain the changed detection factor information when the judgment module 306 judges that the number of pages of the target output page is greater than the preset page number threshold;
[0176] The update module 308 is configured to update the detection factor information to the changed detection factor information to trigger the detection module 301 to re-perform the above operation of detecting all target pages according to the text image detection model and the changed detection factor information.
[0177] In this optional embodiment, optionally, the text information of all key pages can be determined by the following method:
[0178] Calculate the matching degree between the text information of at least one original page and a preset feature condition;
[0179] According to the matching degree between the text information of all original pages and the preset feature condition, screen out the text information of all original pages whose matching degree is greater than or equal to the preset matching degree threshold from the text information of all original pages as the text information of all key pages.
[0180] It can be seen that implementing Figure 4The intelligent positioning device for the described page can further narrow down the page number range of the target output page corresponding to the key page by adjusting the page number of the target output page, so as to further screen out the target output page with a higher similarity to the key page. In this way, it is beneficial to further improve the recognition accuracy of the text information of the target output page, thereby improving the positioning accuracy of the target output page corresponding to the key page.
[0181] In yet another alternative embodiment, the device may further include:
[0182] The determination module 306 is further configured to determine, for any target detection area corresponding to any target page, whether the rotation angle of the text image is greater than a preset rotation angle threshold according to the position information of the text image in the target detection area;
[0183] The correction module 309 is configured to correct the rotation angle of the text image to obtain the corrected position information of the text image when the determination module 306 determines that the inclination angle of the text image is greater than the preset rotation angle threshold.
[0184] It can be seen that implementing Figure 4 the intelligent positioning device for the described page can improve the recognition accuracy of the text information of the text image by correcting the position information of the text image, thereby improving the reliability of the text information of the text image.
[0185] Embodiment Four
[0186] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of another intelligent positioning device for a page disclosed in an embodiment of the present invention. As Figure 5 shown, the intelligent positioning device for the page may include:
[0187] A memory 401 storing executable program code;
[0188] A processor 402 coupled to the memory 401;
[0189] The processor 402 calls the executable program code stored in the memory 401 to execute the steps in the intelligent positioning method for the page described in Embodiment One or Embodiment Two of the present invention.
[0190] Embodiment Five
[0191] An embodiment of the present invention discloses a computer-storable medium. The computer storage medium stores computer instructions, which are used to execute the steps in the intelligent positioning method for the page described in Embodiment One or Embodiment Two of the present invention when called.
[0192] Embodiment Six
[0193] An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps in the intelligent positioning method of the page described in Embodiment 1 or Embodiment 2.
[0194] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0195] Through the above specific description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium. The storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data.
[0196] Finally, it should be noted that: What is disclosed by an intelligent positioning method and device for a page disclosed in the embodiments of the present invention is only the preferred embodiments of the present invention, and is only used to illustrate the technical solutions of the present invention, rather than limiting it; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: They can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An intelligent positioning method for a page, characterized in that, the method includes: detecting a plurality of target pages according to a determined text image detection model and detection element information, and obtaining at least one text image corresponding to each target page; the detection element information includes a detection identifier of each target page, and there is a corresponding original page for each target page; performing a text recognition operation on each text image corresponding to each target page according to a determined text recognition model, and obtaining text information of each text image corresponding to each target page; determining integrated text information of each target page according to the text information of all the text images corresponding to each target page; and, after determining the integrated text information of each target page according to the text information of all the text images corresponding to each target page, the method further includes: calculating a text edit distance between the integrated text information of all the target pages and the text information of at least one determined key page; all the key pages are derived from the original pages corresponding to all the target pages; screening, from the integrated text information of all the target pages, the integrated text information of all the target pages whose text edit distance is less than or equal to a preset text edit distance threshold according to the text edit distance between the integrated text information of all the target pages and the text information of any one of the key pages, as the text information of the target output page corresponding to this key page; all the target pages include the target output page; wherein, the text information of all the key pages is determined by the following method: calculating the matching degree between the text information of at least one of the original pages and a preset feature condition; screening, from the text information of all the original pages, the text information of all the original pages whose matching degree is greater than or equal to a preset matching degree threshold according to the matching degree between the text information of all the original pages and the preset feature condition, as the text information of all the key pages; and, the method further includes: determining the number of pages of the target output page corresponding to any one of the key pages according to the text information of the target output page corresponding to this key page, and judging whether the number of pages of the target output page is greater than a preset number of pages threshold; when it is judged that the number of pages of the target output page is greater than the preset number of pages threshold, performing a change operation on the detection element information to obtain changed detection element information; updating the detection element information to the changed detection element information to trigger re-executing the operation of detecting all the target pages according to the text image detection model and the changed detection element information.
2. The intelligent positioning method for a page according to claim 1, characterized in that, the detecting a plurality of target pages according to a determined text image detection model and detection element information, and obtaining at least one text image corresponding to each target page includes: Input multiple target pages and detection element information into the determined text image detection model for analysis to obtain the position information of at least one target detection area corresponding to each of the target pages; Extract the text image of the target detection area from the target page according to the position information of each target detection area corresponding to each target page, and use it as at least one text image corresponding to the target page.
3. The intelligent positioning method of a page according to claim 1, wherein, The performing text recognition operations on each text image corresponding to each target page according to the determined text recognition model to obtain the text information of each text image corresponding to each target page includes: Input each text image corresponding to each target page into the determined text recognition model to extract the text features of each text image corresponding to each target page; Fuse the text features of each text image corresponding to each target page to obtain the fused text features of each text image corresponding to each target page; Perform sequence analysis on the fused text features of each text image corresponding to each target page to obtain the fused text feature sequence of each text image corresponding to each target page; Perform text recognition operations on the fused text feature sequence of each text image corresponding to each target page to obtain the text information of each text image corresponding to each target page.
4. The intelligent positioning method of a page according to claim 2, wherein, The determining the integrated text information of each target page according to the text information of all the text images corresponding to each target page includes: Sort the text information of all the text images corresponding to each target page according to the position information of all the target detection areas corresponding to each target page and the determined sorting element information to obtain the sorting result of the text information corresponding to each target page; Integrate the sorting result of the text information corresponding to each target page according to the pre-determined text information integration method to obtain the integrated text information of each target page.
5. The intelligent positioning method of a page according to claim 2, wherein, The method further includes: For any target detection area corresponding to any target page, determine whether the rotation angle of the text image is greater than a preset rotation angle threshold according to the position information of the text image of the target detection area; When it is determined that the tilt angle of the text image is greater than the preset rotation angle threshold, correct the rotation angle of the text image to obtain the corrected position information of the text image.
6. An intelligent positioning device for a page, wherein, The device includes: A detection module, configured to detect a plurality of target pages according to a determined text image detection model and detection element information, and obtain at least one text image corresponding to each of the target pages; the detection element information includes a detection identifier of each of the target pages, and there is a corresponding original page for each of the target pages; An identification module, configured to perform text recognition operations on each of the text images corresponding to each of the target pages according to a determined text recognition model, and obtain text information of each of the text images corresponding to each of the target pages; A determination module, configured to determine integrated text information of each of the target pages according to the text information of all the text images corresponding to each of the target pages; Moreover, the apparatus further includes: A calculation module, configured to calculate a text edit distance between the integrated text information of all the target pages and the text information of at least one determined key page after the determination module determines the integrated text information of each of the target pages according to the text information of all the text images corresponding to each of the target pages; all the key pages are from the original pages corresponding to all the target pages; A screening module, configured to screen, from the integrated text information of all the target pages, the integrated text information of all the target pages whose text edit distance is less than or equal to a preset text edit distance threshold as the text information of the target output page corresponding to this key page according to the text edit distance between the integrated text information of all the target pages and the text information of any one of the key pages; all the target pages include the target output page; Wherein, the text information of all the key pages is determined by the following method: Calculating a matching degree between the text information of at least one of the original pages and a preset feature condition; According to the matching degree between the text information of all the original pages and the preset feature condition, screening out the text information of all the original pages whose matching degree is greater than or equal to a preset matching degree threshold from the text information of all the original pages as the text information of all the key pages; Moreover, the apparatus further includes: The determination module is further configured to determine the number of pages of the target output page corresponding to any one of the key pages according to the text information of the target output page corresponding to this key page; A judgment module, configured to judge whether the number of pages of the target output page is greater than a preset number of pages threshold; A modification module, configured to perform a modification operation on the detection element information to obtain the modified detection element information when the judgment module judges that the number of pages of the target output page is greater than the preset number of pages threshold; An update module, configured to update the detection element information to the modified detection element information to trigger the detection module to re-perform the operation of detecting all the target pages according to the text image detection model and the modified detection element information.
7. An intelligent positioning device for pages, Characterized in that The apparatus includes: A memory storing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory and executes the intelligent positioning method for the page according to any one of claims 1-5.
8. A computer storage medium, characterized in that the computer storage medium stores computer instructions which, when called, are used to execute the intelligent positioning method for the page according to any one of claims 1-5.
Citation Information
Patent Citations
Advertorial presenting frequency statistical method and device
CN106815196A
Method and device for positioning chart in PDF document and computer equipment
CN110348294A
Text processing method and device
CN113362026A