Text splicing method and device, computer equipment and computer readable storage medium

By detecting and sorting text lines on text images, determining candidate text boxes and calculating position offsets, the precise filtering and splicing of text lines is solved, and the problem of insufficient accuracy and stability of text lines merging in the existing technology is overcome, and the limitations of document normativeness is achieved, and efficient processing of non-standardized document pictures is achieved.

CN119990074APending Publication Date: 2025-05-13ZHUHAI MOJIE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311494710.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-09
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, when processing non-standardized document pictures, the accuracy and stability of text line merging are low, and due to the limitation of document normativeness, it is difficult to effectively process non-standardized document pictures or indoor and outdoor images with text.

Method used

By obtaining text images, input them into the preset text detection model for text line detection, obtain a collection of text boxes, sort the text boxes, determine the candidate text boxes, calculate the position offset, and perform text lines splicing if the preset splicing conditions are met.

Benefits of technology

Improve the accuracy and stability of text line merging, overcome the limitations of document normativeness, and can effectively handle non-standardized document pictures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990074A_ABST
    Figure CN119990074A_ABST
Patent Text Reader

Abstract

The invention relates to the field of text processing, and provides a text splicing method and device, computer equipment and a computer readable storage medium, and the text splicing method comprises the steps: obtaining a text image, inputting the text image into a preset text detection model for text line detection, and obtaining a first textbox set, each textbox in the first textbox set corresponds to one text line; sorting the textboxes in the first textbox set to obtain a second textbox set; sequentially determining each textbox in the second textbox set as a current textbox, and determining candidate textboxes corresponding to the current textbox; determining a position offset between the current textbox and the candidate textbox; and if the position offset meets a preset splicing condition, splicing the text line in the current textbox and the text line in the candidate textbox. By setting the preset splicing condition, the accuracy and stability of text line merging are improved, and the limitation of document normalization is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of text processing, and in particular to a text splicing method, a device, a computer equipment and a computer-readable storage medium. Background Art

[0002] At present, relevant text paragraph recognition technologies include rule-based text line merging technology and deep learning-based technology. Although rule-based text line merging technology is effective in some scenarios, its generalization is poor and it is prone to errors. Especially when processing non-standardized document images, its accuracy and stability need to be improved. Deep learning-based technology is limited by the standardization of documents and cannot be used for non-standardized document images or indoor and outdoor images with text. Therefore, when processing non-standardized document images, how to improve the accuracy and stability of text line merging and overcome the limitations of document standardization has become an urgent problem to be solved. Summary of the invention

[0003] The present application provides a text splicing method, apparatus, computer equipment and storage medium to improve the accuracy and stability of text line merging when processing text and overcome the limitations of document standardization.

[0004] In a first aspect, the present application provides a text splicing method, the method comprising:

[0005] Acquire a text image, input the text image into a preset text detection model to perform text line detection, and obtain a first text box set, each text box in the first text box set corresponds to a text line; sort the text boxes in the first text box set to obtain a second text box set; determine each text box in the second text box set as a current text box in turn, and determine a candidate text box corresponding to the current text box; determine a position offset between the current text box and the candidate text box; if the position offset meets a preset splicing condition, splice the text line in the current text box with the text line in the candidate text box.

[0006] In a second aspect, the present application further provides a text splicing device, the device comprising:

[0007] A first text box set acquisition module acquires a first text box set corresponding to a text image, each text box in the first text box set corresponds to a text line; a second text box set determination module sorts the text boxes in the first text box set to obtain a second text box set; a candidate text box determination module sequentially determines each text box in the second text box set as a current text box, and determines a candidate text box corresponding to the current text box; a position offset determination module determines the position offset between the current text box and the candidate text box; a text line splicing module splices the text line in the current text box with the text line in the candidate text box if the position offset meets a preset splicing condition.

[0008] In a third aspect, the present application also provides a computer device, which includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the text splicing method as described above when executing the computer program.

[0009] In a fourth aspect, the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the text splicing method as described above.

[0010] The present application discloses a text splicing method, device, computer equipment and storage medium. The text splicing method includes: obtaining a text image, inputting the text image into a preset text detection model to perform text line detection, obtaining a first text box set, each text box in the first text box set corresponds to a text line; further sorting the text boxes in the first text box set to obtain a second text box set; sequentially determining each text box in the second text box set as a current text box, and determining the candidate text box corresponding to the current text box; determining the position offset between the current text box and the candidate text box; finally, according to a preset splicing condition, if the position offset meets the preset splicing condition, the text line in the current text box and the text line in the candidate text box are spliced. The method sorts the text boxes, determines the candidate text box corresponding to each text box, further determines the position offset between each text box and the candidate text box, and performs splicing when the position offset meets the splicing condition, thereby realizing accurate screening of text boxes according to the position offset, improving the accuracy and stability of text line merging, and overcoming the limitation of document standardization. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0012] Figure 1 is a schematic flow chart of a text splicing method provided in an embodiment of the present application;

[0013] Figure 2 yes Figure 1 A schematic flow chart of the sub-steps of the text splicing method in ;

[0014] Figure 3 is a schematic flowchart of sub-steps for determining a first offset provided by an embodiment of the present application;

[0015] Figure 4 It is a schematic diagram of determining the position of a text box provided in an embodiment of the present application;

[0016] Figure 5 is a schematic diagram of a text box position offset provided in an embodiment of the present application;

[0017] Figure 6 is a schematic diagram of another text box position provided in an embodiment of the present application;

[0018] Figure 7 is a schematic diagram of a text splicing device provided in an embodiment of the present application;

[0019] Figure 8 It is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0020] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0021] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.

[0022] It should be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0023] It should be further understood that the term “and / or” used in the specification and appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0024] The embodiments of the present application provide a text splicing method, apparatus, computer equipment and storage medium. The text splicing method can be applied to a computer equipment to realize accurate merging of text lines by performing text recognition on text images.

[0025] Exemplarily, the computer device may be a server or a terminal. The server may be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The terminal may be a smart phone, smart glasses, tablet computer, laptop computer, desktop computer, and other devices.

[0026] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0027] See also Figure 1 , Figure 1 This is a schematic flow chart of a text splicing method provided by an embodiment of the present application. The text splicing method sorts text boxes, determines the candidate text box corresponding to each text box, further determines the position offset between each text box and the candidate text box, and performs splicing when the position offset meets the splicing condition, thereby achieving accurate screening of text boxes according to the position offset, improving the accuracy and stability of text line merging, and overcoming the limitations of document standardization.

[0028] like Figure 1 As shown, the text splicing method specifically includes steps S101 to S105.

[0029] S101, obtaining a text image, inputting the text image into a preset text detection model to perform text line detection, and obtaining a first text box set, wherein each text box in the first text box set corresponds to a text line.

[0030] When it is necessary to perform text recognition on non-standardized document images, for example, when taking pictures of learning content on PPT in class, due to time constraints, the pictures taken are often not standard pictures, or when reading in the library, it is a waste of time to directly copy the text content you like, and when taking pictures with a mobile phone, the middle part of the picture will be more prominent and the two sides will be darker. For such non-standardized document images, the text splicing method in this application is needed for text line detection and text splicing. The text image is input into the text detection model for text line detection. For example, the text detection model can use any one of PaddleOCR, Tesseract and mmocr to perform text detection on the text image to obtain the position of each text line in the image, where the position of the text line is represented by a text box.

[0031] Specifically, the text box is a rectangle consisting of an upper left point, an upper right point, a lower right point and a lower left point. The output result of the text detection model corresponds to n text boxes, corresponding to the n text lines in the image, and the n text boxes constitute a first text box set.

[0032] The above embodiment inputs the text image into the text detection model to perform text line detection, thereby obtaining a first text box set consisting of text boxes for subsequent text line merging, thereby overcoming the limitation of document standardization.

[0033] S102: Sort the text boxes in the first text box set to obtain a second text box set.

[0034] Specifically, all text boxes in the first text box set are sorted in ascending order according to the size of the ordinate of the upper left point. If the ordinates are the same, they are sorted in ascending order according to the abscissa of the upper left point to obtain a sorted second text box set.

[0035] For example, suppose there are three text boxes in the first text box set, namely text box A, text box B, and text box C. The coordinate value of the upper left point of text box A is (1, 3), the coordinate value of the upper left point of text box B is (2, 1), and the coordinate value of the upper left point of text box C is (3, 2). All text boxes in the first text box set are sorted from small to large according to the vertical coordinate of the upper left point. The sorted second text box set is: text box B < text box C < text box A.

[0036] For example, assume that there are 4 text boxes in the first text box set, namely text box a, text box b, text box c and text box d. The coordinate value of the upper left point of text box a is (2, 3), the coordinate value of the upper left point of text box b is (1, 3), the coordinate value of the upper left point of text box c is (3, 2) and the coordinate value of the upper left point of text box d is (4, 7). All text boxes in the first text box set are sorted in ascending order according to the size of the ordinate of the upper left point. The ordinate of the upper left point of text box c is the smallest, and the ordinate of the upper left point of text box d is the largest. The ordinates of the upper left points of text box a and text box b are the same and are located between text box c and text box d. It is necessary to make further judgments on text box a and text box b. Text box a and text box b are sorted in ascending order according to the abscissa of the upper left points of text box a and text box b. Because the abscissa of the upper left point of text box a is greater than the abscissa of the upper left point of text box b, the sorting result of text box a and text box b is: text box b < text box a. Therefore, the sorted second text box set is text box c<text box b<text box a<text box d.

[0037] In the above embodiment, by sorting the text boxes in the first text box set from small to large to obtain the sorted second text box set, orderly screening of the text boxes can be achieved, thereby improving the efficiency of text box screening.

[0038] S103: sequentially determine each text box in the second text box set as a current text box, and determine a candidate text box corresponding to the current text box.

[0039] Exemplarily, each text box in the second text box set can be determined as the current text box in turn, and the candidate text box corresponding to the current text box can be determined. First, the first text box corresponding to the current text box needs to be determined, where the first text box here is a text box whose distance value between the center point of the remaining text boxes and the midline of the current text box is less than a preset distance threshold, where the remaining text boxes are any other text boxes in the second text box set excluding the current text box, and the center point of the remaining text box is determined by the average value of the coordinates of the four corner points of the upper left point, the lower left point, the upper right point, and the lower right point of the remaining text box; the second text box corresponding to the current text box is determined from the first text box, where the second text box is located below the current text box; then a third text box that satisfies a preset overlap needs to be determined from the second text box; finally, a candidate text box that satisfies a preset height condition is determined from the third text box.

[0040] In the above embodiment, the current text box in the second text box set is determined, and candidate text boxes are obtained by screening according to screening conditions such as distance threshold conditions, preset overlap conditions, and preset height conditions. The candidate text boxes can be used for subsequent judgment of whether the text boxes satisfy the splicing relationship, thereby improving the accuracy of text line merging.

[0041] S104. Determine the position offset between the current text box and the candidate text box.

[0042] In some embodiments, determining the position offset between the current text box and the candidate text box may include: determining a first offset of the left side of the candidate text box relative to the left side of the current text box; determining a second offset of the right side of the candidate text box relative to the right side of the current text box.

[0043] S105. If the position offset meets a preset splicing condition, splice the text lines in the current text box and the text lines in the candidate text box.

[0044] Exemplarily, the text line in the current text box is: "Input the image into the text detection model to obtain the", and the text line in the candidate text box is: "position, and the position of the text line is represented by a text box, and the text box is a rectangle". If the position offset between the current text box and the candidate text box meets the preset splicing condition, splice the text line in the current text box and the text line in the candidate text box, and the spliced text line is: "Input the image into the text detection model to obtain the position of each text line in the image, and the position of the text line is represented by a text box, and the text box is a rectangle".

[0045] Exemplarily, the preset splicing condition is -1 < left_offset_standard < 4 and right_offset_standard < 2. Of course, the magnitudes of these boundary values -1, 4, and 2 can be set to any other suitable values as needed. It should be noted that left_offset_standard represents the first offset and right_offset_standard represents the second offset.

[0046] Only when the text lines in the current text box and the candidate text box meet the above preset splicing condition can the text lines in the current text box and the candidate text box be spliced. Refer to Figure 6 , Figure 6 which is another schematic diagram of the text box position provided by the embodiments of the present application. Only the part of the dotted text box meets the text line splicing condition, and then the text lines can be spliced.

[0047] The text splicing method provided in the above embodiment overcomes the limitation of document standardization by inputting the text image into the text detection model for text line detection to obtain a first text box set composed of text boxes. The text boxes in the first text box set are sorted from small to large to improve the efficiency of text box screening. According to screening conditions such as distance threshold conditions, preset overlap conditions, and preset height conditions, candidate text boxes are obtained, thereby improving the accuracy of text line merging.

[0048] See also Figure 2 , Figure 2 yes Figure 1 The sub-step schematic flowchart of the text splicing method in the text splicing method may specifically include the following steps S201 to S204.

[0049] S201: Determine a first text box corresponding to the current text box from the second text box set, where the first text box is a text box whose distance value between the center point of the remaining text boxes and the midline of the current text box is less than a preset distance threshold, wherein the remaining text boxes are any other text boxes in the second text box set excluding the current text box.

[0050] Specifically, the preset distance threshold is a preset multiple of the first height. When d<2h, the distance value between the center point of the remaining text box and the midline of the current text box is less than the preset distance threshold, and the remaining text box is determined to be the first text box. Wherein, d represents the distance value between the center point of the first text box and the midline of the current text box, h represents the first height, and the coefficient 2 represents a suitable value for judging whether the distance between the current text box and the remaining text box is close. The coefficient 2 can also be replaced with a number such as 1.9 or 2.1 according to actual needs. It should be noted that the first height is the height of the current text box.

[0051] It should be noted that the center point is determined by the average value of the coordinates of the four corner points of the remaining text boxes, and the midline represents the line connecting the midpoint of the upper left point and lower left point of the current text box and the midpoint of the upper right point and lower right point.

[0052] S202: Determine a second text box corresponding to the current text box from the first text box, where the second text box is a text box below the current text box.

[0053] Specifically, determine the centerline vector a of the current text box, a=(ax,ay,0), where the centerline represents the line connecting the midpoint of the upper left point and lower left point of the text box and the midpoint of the upper right point and lower right point. According to the centerline of the current text box, determine the centerline vector of the current text box. Determine the midpoint vector b of the first text box, b=(bx,by,0), where the midpoint vector b represents the line connecting the midpoint of the upper left point and lower left point of the current text box and the midpoint of the upper right point and lower right point of the candidate text box. Please refer to Figure 4 , Figure 4It is a text box position judgment diagram provided by an embodiment of the present application. Among them, vector a represents the centerline vector of the current text box, a=(ax,ay,0); vector b represents the midpoint vector of the first text box, b=(bx,by,0). Calculate the vector value between the centerline vector a and the midpoint vector b; if the vector value meets the preset vector condition, the first text box is determined as the second text box. Among them, the preset vector condition is that the vector value between the centerline vector a of the current text box and the midpoint vector b of the first text box is greater than 0, that is, ax*by-ay*bx>0.

[0054] Exemplarily, the vector cross product formula can be used to determine whether the first text box is below the current text box. The midline vector of the current text box a=(ax,ay,0), and the midpoint vector of the first text box b=(bx,by,0). According to the vector cross product formula: a×b=(ax*by-ay*bx)k, when ax*by-ay*bx>0, the preset vector condition is met, and the first text box is below the current text box, and the first text box is determined to be the second text box.

[0055] S203: Determine a third text box from the second text box, where the third text box is a text box whose overlap with the current text box is greater than a preset overlap.

[0056] Specifically, based on the preset overlap calculation formula, the overlap between the current text box and the second text box is calculated; if the overlap between the current text box and the second text box meets the preset overlap condition, the second text box is determined as the third text box. The preset overlap calculation formula is: d1 = min(x2, x4) - max(x1, x3), x1 and x2 represent the x-coordinates of the upper left point and the upper right point of the current text box, respectively, x3 and x4 represent the x-coordinates of the upper left point and the upper right point of the second text box, min(x2, x4) represents the minimum value of x2 and x4, and max(x1, x3) represents the maximum value of x1 and x3. Assuming that the preset overlap is 0, if the overlap between the current text box and the second text box is greater than 0, the second text box is determined as the third text box. It should be noted that the preset overlap here can be set to other suitable values ​​such as 0.1, 0.2, etc. as needed.

[0057] For example, the x-coordinates of the upper left and upper right points of the current text box are: x1 = 2, x2 = 5, and the x-coordinates of the upper left and upper right points of the second text box are: x3 = 1, x4 = 3. According to the overlap calculation formula d1 = min(x2, x4) - max(x1, x3), d1 = min(5, 3) - max(2, 1) = 1, d1>0, the current text box and the second text box overlap in the x-axis direction, and the current text box and the second text box meet the preset overlap condition in the x-axis direction, then the second text box is determined as the third text box.

[0058] S204. If the height of the third text box and the first height of the current text box meet the preset height condition, then determine the third text box as the candidate text box.

[0059] Specifically, the preset height condition is 2 / 3 < h2 / h < 3 / 2 or 1 / 2 < h2 / h < 2, or any other appropriate range, which can be set according to needs. The purpose is to screen candidate text boxes with the same font size and close to the height of the current text box. Among them, it is assumed that h2 represents the height of the third text box.

[0060] Exemplarily, when the current text box is the last text box, when judging whether the text can be spliced, in step S201, determine the first text box corresponding to the current text box from the second text box set. According to the judgment condition of the first text box, if the distance threshold condition is met, step S202 can be continued, and the second text box corresponding to the current text box is determined from the first text box. At this time, the midline vector a of the current text box = (ax, ay, 0). Since the current text box is the last text box, there is no first text box at this time, and the midpoint vector b of the first text box = (0, 0, 0). According to the vector cross product formula a×b = (ax*by - ay*bx)k, ax*by - ay*bx = 0, which does not meet the preset vector condition ax*by - ay*bx > 0. Then the first text box below the current text box cannot be determined, and the execution ends at step S202, and the remaining steps S203 and S204 are no longer continued. The last current text box will be used as a single text box and will not participate in text merging.

[0061] It should be noted that if any one of the above steps S201 to S204 does not meet the judgment condition of the corresponding step, then the candidate text box cannot be obtained.

[0062] The candidate text box screening method provided by the above embodiment realizes the precise screening of text boxes by judging whether the text boxes in the second text box set meet the preset distance threshold condition, preset vector condition, preset overlap condition, and preset height condition, and improves the accuracy and stability of text line merging.

[0063] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of sub-steps for determining the first offset provided by an embodiment of the present application, and specifically may include the following steps S301 to S304.

[0064] S301. Determine the first distance between the left midpoint of the candidate text box and the left midpoint of the current text box.

[0065] Exemplarily, the left midpoint of the current text box refers to the midpoint of the upper left point and the lower left point of the current text box; the left midpoint of the candidate text box refers to the midpoint of the upper left point and the lower left point of the candidate text box, and the first distance refers to the distance between the left midpoint of the current text box and the left midpoint of the candidate text box. For example, the coordinates of the upper left point and the lower left point of the current text box are: (2, 4) and (2, 6), respectively, then the left midpoint of the current text box is expressed as: (2, 5); the coordinates of the upper left point and the lower left point of the candidate text box are: (3, 8) and (3, 10), respectively, then the left midpoint of the candidate text box is expressed as: (3, 9); the first distance between the left midpoint of the selected text box and the left midpoint of the current text box calculated from the left midpoint of the current text box (2, 5) and the left midpoint of the candidate text box (3, 9) is

[0066] S302: Determine a second distance between the left midpoint of the candidate text box and the midline of the current text box.

[0067] Exemplarily, the center line of the current text box refers to the line connecting the midpoint of the upper left point and the lower left point of the current text box and the midpoint of the upper right point and the lower right point. The second distance refers to the distance between the left midpoint of the candidate text box and the center line of the current text box. For example, the coordinates of the left midpoint of the candidate text box are (3, 9), and the center line of the current text box is y=5, then the second distance between the left midpoint of the candidate text box and the center line of the current text box is 4.

[0068] S303: Calculate according to the first distance and the second distance to obtain an initial offset.

[0069] For example, the initial offset of the left side of the candidate text box relative to the left side of the current text box can be calculated according to the Pythagorean theorem, and the formula is expressed as follows: Where d31 represents the first distance, d32 represents the second distance, and left_offset represents the initial offset. When d32=4, the calculated initial offset left_offset=1.

[0070] S304: Determine the first height of the current text box.

[0071] Exemplarily, the first height refers to the height of the current text box. For example, the coordinates of the upper left point and the lower left point of the current text box are (2, 4) and (2, 6) respectively, then the first height of the current text box h=2.

[0072] S305: Divide the initial offset by the first height to obtain a first offset.

[0073] Exemplarily, the first offset left_offset_standard is determined according to an offset standardization formula, which is expressed as: left_offset_standard=left_offset / h. For example, if the first height h=2 and the initial offset left_offset=1, then the first offset left_offset_standard=1 / 2.

[0074] It should be noted that according to the offset normalization formula, the calculated first offset can be positive or negative. Figure 5 As shown. The following method can be used to determine the positive or negative of the first offset. The line connecting the lower left point and the upper left point of the candidate text box is used as vector c, and the line connecting the lower left point and the upper left point of the current text box is used as vector d. When cx*dy-cy*dx<0, the first offset is a negative number, and a negative sign needs to be added in front of the calculated first offset. Where c=(cx,cy,0), d=(dx,dy,0).

[0075] Similarly, the second offset right_offset_standard can be calculated based on the above method and the positive or negative sign of the second offset can be determined.

[0076] In the above embodiment, the first distance between the left midpoint of the candidate text box and the left midpoint of the current text box is determined, the second distance between the left midpoint of the candidate text box and the center line of the current text box is determined, and the first distance and the second distance are calculated to obtain an initial offset, and further, the first height of the current text box is calculated, and the initial offset is divided by the first height to finally calculate the first offset. The method provides a basis for judging whether the current text box and the candidate text box meet the preset splicing conditions through the first offset finally calculated.

[0077] Specifically, taking the processing of 4 text lines detected in the text detection model of the text image input as an example, the implementation process of the text splicing method is specifically analyzed:

[0078] The text image is input into a preset text detection model to perform text line detection to obtain a first text box set. Assume that the first text box set is {“A,3Related Work”, “B,In this section we review the contributions of works”, “C,tions[6,10,17,29,60,61]”, “D,most related to ours.These focus solely on text detec-”}, where the coordinates of the upper left point, lower left point, upper right point, and lower right point of “A,3Related Work” are: (1, 1), (1, 2), (2, 1), (2, 2); the coordinates of the upper left point, lower left point, upper right point, and lower right point of “B,In this section wereview the contributions of works” are: (1, 4), (1, 6), (4, 4), (4, 6); the coordinates of the upper left point, lower left point, upper right point, and lower right point of “C,tions[6,10,17,29,60,61],text recognition.” are: (1, 10), (1, 12), (3, 10), (3, 12); The coordinates of the upper left point, lower left point, upper right point, and lower right point are: (1, 7), (1, 9), (4, 7), and (4, 9), respectively. Then, all the text boxes in the first text box set are sorted in ascending order according to the size of the ordinate of the upper left point. If the ordinates are the same, they are sorted in ascending order according to the abscissa of the upper left point to obtain the sorted second text box set. The sorted set of the second text boxes is: {“A,3Related Work”, “B,In this section we review the contributions of works”, “D,most related to ours. These focus solely on text detec-”, “C,tions[6,10,17,29,60,61]”}, and the set of text boxes in the second text box except the current text box is: {“B,In this section we review the contributions of works”, “D,most related to ours. These focus solely on text detec-”, “C,tions[6,10,17,29,60,61]”}.

[0079] Each text box in the second text box set is determined as the current text box in turn, and a candidate text box corresponding to the current text box is determined. This can be achieved by using two for loop statements, wherein the outer loop is used to determine the current text box and the inner loop is used to determine the candidate text box. The text boxes whose distance value between the center point of each text box in the second text box set except the current text box and the midline of the current text box is less than a preset distance threshold are determined in turn.

[0080] The first step is to determine the first text box corresponding to the current text box from the second text box set. The first text box is a text box whose distance value between the center point of the second text box set and the midline of the current text box is less than a preset distance threshold. For example, the current text box is: "A, 3Related Work", and one of the text boxes in the second text box set is: "B, In this section we review the contributions of works". After calculation, the distance value between the center point of text box B and the midline of the current text box A is d=3. According to the formula: d<2h, because the height of the current text box h=1, the current text box A and text box B do not meet the constraint condition of the above formula d<2h, so the text line in the current text box A and the text line in text box B cannot be spliced. Similarly, if text box C or D in the second text box set and the current text box A do not meet the constraint condition of d<2h, the text line in the current text box A and the text line in text box C or D cannot be spliced.

[0081] If the current text box is text box B, the set of the second text box set except text box B is: {"D,most related to ours.These focus solely on text detec-","C,tions[6,10,17,29,60,61]"}.

[0082] When the text box in the second text box set is text box D, after calculation, the distance between the center point of text box D and the midline of the current text box is d=3. According to the formula: d<2h, because the height of the current text box B is h=2, the current text box B and text box D meet the constraint condition of the above formula d<2h, and text box D is determined as the first text box. Then, the second text box corresponding to the current text box B is determined from the first text box D, and the second text box is the text box below the current text box.

[0083] In the second step, determine the midline vector a = (ax, ay, 0) of the current text box B, and determine the midpoint vector b = (bx, by, 0) between the first text box D and the current text box B. According to the vector cross - product method, when ax * by - ay * bx > 0, the text line of the first text box D is below the text line of the current text box B. At this time, the first text box D is determined as the second text box.

[0084] In the third step, judge the coincidence degree between the current text box B and the second text box D. Based on the preset coincidence - degree calculation formula, calculate the coincidence degree between the current text box B and the second text box D. If the coincidence degree meets the coincidence - degree condition, then determine the second text box D as the third text box.

[0085] In the fourth step, calculate the height ratio of the current text box B to the third text box D. If the height of the third text box and the first height of the current text box meet the preset height condition, that is, the preset height condition is 2 / 3 < h2 / h < 3 / 2 or 1 / 2 < h2 / h < 2, or any other appropriate range, which can be set according to needs. Then determine the third text box D as the candidate text box. Similarly, in turn, take the text box D and the text box C in the second - text - box set as the current text box respectively, and judge whether the other text boxes in the second - text - box set are candidate text boxes, which will not be elaborated here.

[0086] In the fifth step, determine the position offset between the current text box B and the candidate text box D, which can include the first offset and the second offset. First, calculate the initial offset of the left side of the candidate text box D relative to the left side of the current text box B, d 31 = 3, d 32 = 3, from Calculation shows that the initial offset of the left side of the candidate text box D relative to the left side of the current text box B is 0. Divide the initial offset by the first height to get the first offset left_offset_standard as 0; then, calculate the initial offset of the right side of the candidate text box D relative to the right side of the current text box B. d 42 = 3, by right_ Calculation shows that the initial right offset of the right side of the candidate text box D relative to the right side of the current text box B is 1. Dividing the initial right offset by the first height gives the second offset right_offset_standard = 1 / 2. Comparing and analyzing the obtained first offset and second offset with the preset splicing condition -1 < left_offset_standard < 4 and right_offset_standard < 2, since the calculated position offset between the candidate text box D and the current text box B is within the preset splicing condition range, the text lines in the current text box B and the text lines in the candidate text box D can be spliced. Similarly, the position offsets of the remaining candidate text boxes in the second text box set that meet the splicing condition and the current text box are calculated in turn to determine whether the other candidate text boxes in the second text box set and the current text box meet the splicing condition, which will not be elaborated here.

[0087] Please refer to Figure 7 , Figure 7 which is a schematic diagram of a text splicing device provided by an embodiment of the present application. The text splicing device is used to execute the foregoing text splicing method. Among them, the text splicing device can be configured in a server.

[0088] As Figure 7 shown, the text splicing device 700 includes: a first text box set acquisition module 701, a second text box set determination module 702, a candidate text box determination module 703, a position offset determination module 704, and a text line splicing module 705.

[0089] The first text box set acquisition module 701 is used to acquire a first text box set corresponding to a text image, and each text box in the first text box set corresponds to a text line.

[0090] The second text box set determination module 702 is used to sort the text boxes in the first text box set to obtain a second text box set;

[0091] The candidate text box determination module 703 is used to sequentially determine each text box in the second text box set as the current text box and determine the candidate text box corresponding to the current text box;

[0092] The position offset determination module 704 is used to determine the position offset between the current text box and the candidate text box;

[0093] The text line splicing module 705 is used to splice the text lines in the current text box and the text lines in the candidate text box if the position offset meets the preset splicing condition.

[0094] In some embodiments, the first text box set acquisition module 701 is specifically used to:

[0095] A first text box set corresponding to the text image is obtained, where each text box in the first text box set corresponds to a text line.

[0096] In some embodiments, the second text box set determination module 702 is specifically configured to:

[0097] The text boxes in the first text box set are sorted to obtain a second text box set.

[0098] In some embodiments, the candidate text box determination module 703 is specifically used to:

[0099] From the second text box set, determine the first text box corresponding to the current text box, the first text box being a text box whose distance value between the center point of the remaining text boxes and the midline of the current text box is less than a preset distance threshold. Determine the midline vector of the current text box; determine the midpoint vector of the first text box; calculate the vector value between the midline vector and the midpoint vector; if the vector value satisfies the preset vector condition, determine the first text box as the second text box. , the second text box is the text box below the current text box. Based on a preset overlap calculation formula, calculate the overlap between the current text box and the second text box; if the overlap meets the overlap condition, determine the second text box as the third text box; the third text box is a text box whose overlap with the current text box is greater than the preset overlap. If the height of the third text box and the first height of the current text box meet the preset height condition, determine the third text box as the candidate text box.

[0100] In some embodiments, the position offset determination module 704 is specifically configured to:

[0101] Determine a first offset of the left side of the candidate text box relative to the left side of the current text box; determine a first distance between the left midpoint of the candidate text box and the left midpoint of the current text box;

[0102] Determine a second distance between the left midpoint of the candidate text box and the midline of the current text box; calculate based on the first distance and the second distance to obtain an initial offset; determine a first height of the current text box; divide the initial offset by the first height to obtain the first offset. Determine a second offset of the right side of the candidate text box relative to the right side of the current text box.

[0103] In some embodiments, the text line splicing module 705 is specifically used to:

[0104] If the position offset satisfies a preset splicing condition, the text line in the current text box and the text line in the candidate text box are spliced.

[0105] It should be noted that the above text splicing method can also be implemented by training a detection model, which uses paragraphs as detection targets instead of detecting text lines in existing text splicing methods.

[0106] It should be noted that the above-mentioned text splicing method can also input the results of text detection and text recognition in the above-mentioned text splicing method into chatGPT or other NLP large models, so that the large model can splice the text lines according to semantic relationships to achieve the purpose of merging them into paragraphs.

[0107] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described device and each module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0108] The above-mentioned device can be implemented in the form of a computer program. Figure 8 Runs on the computer device shown.

[0109] See also Figure 8 , Figure 8 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device may be a server.

[0110] See also Figure 8 The computer device includes a processor, a memory and a network interface connected through a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.

[0111] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any text splicing method.

[0112] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.

[0113] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any text splicing method.

[0114] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 8The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0115] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0116] In one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:

[0117] Acquire a text image, input the text image into a preset text detection model to perform text line detection, and obtain a first text box set, each text box in the first text box set corresponds to a text line; sort the text boxes in the first text box set to obtain a second text box set; determine each text box in the second text box set as a current text box in turn, and determine the candidate text box corresponding to the current text box; determine the position offset between the current text box and the candidate text box; if the position offset meets the preset splicing condition, splice the text line in the current text box with the text line in the candidate text box.

[0118] In one embodiment, the position offset includes a first offset and a second offset; when the processor determines the position offset between the current text box and the candidate text box, it is used to implement:

[0119] A first offset of the left side of the candidate text box relative to the left side of the current text box is determined; and a second offset of the right side of the candidate text box relative to the right side of the current text box is determined.

[0120] In one embodiment, when determining the first offset of the left side of the candidate text box relative to the left side of the current text box, the processor is configured to implement:

[0121] Determine a first distance between the left midpoint of the candidate text box and the left midpoint of the current text box; determine a second distance between the left midpoint of the candidate text box and the center line of the current text box; calculate based on the first distance and the second distance to obtain an initial offset; determine a first offset based on the initial offset.

[0122] In one embodiment, when determining the first offset according to the initial offset, the processor is configured to implement:

[0123] Determine the first height of the current text box; divide the initial offset by the first height to obtain the first offset.

[0124] In one embodiment, when determining the candidate text box corresponding to the current text box, the processor is used to implement:

[0125] From the second text box set, determine the first text box corresponding to the current text box, the first text box is a text box whose distance value between the center point of the remaining text boxes and the midline of the current text box is less than a preset distance threshold; determine the second text box corresponding to the current text box from the first text box, the second text box is a text box below the current text box; determine the third text box from the second text box, the third text box is a text box whose overlap with the current text box is greater than a preset overlap; if the height of the third text box and the first height of the current text box meet the preset height condition, then determine the third text box as a candidate text box. The preset distance threshold is a preset multiple of the first height.

[0126] In one embodiment, when the processor implements determining the second text box corresponding to the current text box from the first text box, it is used to implement:

[0127] Determine the centerline vector of the current text box; determine the midpoint vector of the first text box; calculate the vector value between the centerline vector and the midpoint vector; if the vector value meets the preset vector condition, determine the first text box as the second text box.

[0128] In one embodiment, when the processor implements determining the third text box from the second text box, it is configured to implement:

[0129] Based on a preset overlap calculation formula, the overlap between the current text box and the second text box is calculated; if the overlap meets the overlap condition, the second text box is determined as the third text box.

[0130] A computer-readable storage medium is also provided in an embodiment of the present application. The computer-readable storage medium stores a computer program. The computer program includes program instructions. The processor executes the program instructions to implement any text splicing method provided in the embodiment of the present application.

[0131] The computer-readable storage medium may be an internal storage unit of the computer device of the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device.

[0132] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A text splicing method, characterized in that: include: Get the first text box set corresponding to the text image , Each text box in the first text box set corresponds to a text line; Sorting the text boxes in the first text box set to obtain a second text box set; sequentially determining each text box in the second text box set as a current text box, and determining a candidate text box corresponding to the current text box; Determine the position offset between the current text box and the candidate text box; If the position offset satisfies a preset splicing condition, the text line in the current text box and the text line in the candidate text box are spliced.

2. The text splicing method according to claim 1, characterized in that: The position offset includes a first offset and a second offset; and determining the position offset between the current text box and the candidate text box includes: Determine a first offset of the left side of the candidate text box relative to the left side of the current text box; A second offset of the right side of the candidate text box relative to the right side of the current text box is determined.

3. The text splicing method according to claim 2, characterized in that: The determining a first offset of the left side of the candidate text box relative to the left side of the current text box includes: Determine a first distance between a left midpoint of the candidate text box and a left midpoint of the current text box; Determining a second distance between the left midpoint of the candidate text box and the midline of the current text box; Calculate the first distance and the second distance to obtain an initial offset; The first offset is determined according to the initial offset.

4. The text splicing method according to claim 3, characterized in that: The determining the first offset according to the initial offset includes: Determine a first height of the current text box; The initial offset is divided by the first height to obtain the first offset.

5. The text splicing method according to claim 1, characterized in that: The determining a candidate text box corresponding to the current text box includes: Determine, from the second text box set, a first text box corresponding to the current text box, the first text box being a text box for which a distance value between a center point of the remaining text boxes and a midline of the current text box is less than a preset distance threshold; Determine a second text box corresponding to the current text box from the first text box, where the second text box is a text box below the current text box; Determining a third text box from the second text box, the third text box being a text box whose overlap with the current text box is greater than a preset overlap; If the height of the third text box and the first height of the current text box meet a preset height condition, the third text box is determined as the candidate text box.

6. The text splicing method according to claim 5, characterized in that: The determining, from the first text box, a second text box corresponding to the current text box includes: Determining the centerline vector of the current text box; Determine the midpoint vector of the first text box; Calculating a vector value between the midline vector and the midpoint vector; If the vector value satisfies a preset vector condition, the first text box is determined as the second text box.

7. The text splicing method according to claim 5, characterized in that: The determining the third text box from the second text box comprises: Calculating the degree of overlap between the current text box and the second text box based on a preset degree of overlap calculation formula; If the overlap degree satisfies the overlap degree condition, the second text box is determined as the third text box.

8. A text splicing device, characterized in that: include: A first text box set acquisition module is used to acquire a first text box set corresponding to a text image. , Each text box in the first text box set corresponds to a text line; A second text box set determination module, which sorts the text boxes in the first text box set to obtain a second text box set; A candidate text box determination module determines each text box in the second text box set as a current text box in turn, and determines a candidate text box corresponding to the current text box; A position offset determination module is used to determine the position offset between the current text box and the candidate text box; The text line splicing module splices the text line in the current text box with the text line in the candidate text box if the position offset satisfies a preset splicing condition.

9. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program and implement the text splicing method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the text splicing method according to any one of claims 1 to 7.