Rectangular pattern-based label sample generation method, image processing method and device

CN118898850BActive Publication Date: 2026-09-08SAIC GM WULING AUTOMOBILE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410964087.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2026-09-08
Estimated Expiration
2044-07-18

AI Technical Summary

Technical Problem

[0003]现有的文字标注的方法,一种是使用水平矩形框标注文字的位置,再手动拖动旋转矩形框至合适的旋转角度,得到旋转矩形框的标注;这种文字标注的方法需要先标注水平矩形再手动进行旋转,若目标文字区域与水平矩形区域偏差很大,则需要手动一点点拖动水平矩形框至合适的旋转角度,标注效率低

Benefits of technology

[0209] The method for generating labeled samples based on rectangular patterns according to this application has the following advantages: This application uses the coordinates of three scene points at preset locations to determine the positions of the four vertices of four rectangular text boxes, thereby automatically generating text boxes with the shape of rotated rectangles, improving the efficiency of image annotation. Furthermore, this application determines the top and bottom edges of the rectangular text boxes, thus determining the orientation of the text within the rectangular text boxes during the generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118898850B_ABST
    Figure CN118898850B_ABST
Patent Text Reader

Abstract

The application discloses a rectangular pattern-based label sample generation method, an image processing method and device, and belongs to the technical field of image recognition. The method comprises the following steps: in response to a user operation on a target image, the coordinates of three scene points on the target image are determined; according to the positional relationship among a first scene point, a second scene point and a third scene point, a first straight line where the top edge of a rectangular text box is located and a second straight line where the bottom edge of the rectangular text box is located are determined; according to the first straight line and the second straight line, the foot points of the three scene points corresponding to the first straight line are respectively determined in combination with a preset rectangular coordinate system, and the coordinates of four vertices of the rectangular text box are determined according to the positional relationship among the foot points, the first straight line and the second straight line, so as to generate the rectangular text box; and a label sample is generated according to the rectangular text box and the target image, so as to automatically generate a character boundary rectangular frame, improve the labeling efficiency of the image, determine the direction of characters in the image, and improve the quality of text region slicing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and in particular to a method for generating labeled samples based on rectangular patterns, an image processing method, an apparatus, a terminal, and a storage medium. Background Technology

[0002] Text annotation is the process of drawing bounding boxes for text in an image and adding text content based on text detection and recognition technologies. It then uses perspective transformation to extract the text regions from the original image, using these text region slices as the training dataset for a text recognition model. Text detection aims to accurately locate the text in the image, while text recognition refers to identifying the text information from the located text. The text in the text region slices must be upright.

[0003] One existing method for text annotation is to use a horizontal rectangle to mark the position of the text, and then manually drag and rotate the rectangle to the appropriate rotation angle to obtain the annotation of the rotated rectangle. This method requires marking the horizontal rectangle first and then manually rotating it. If the target text area deviates greatly from the horizontal rectangle area, it is necessary to manually drag the horizontal rectangle to the appropriate rotation angle little by little, which results in low annotation efficiency.

[0004] Another existing method for text annotation is to set the four vertices of the box and generate corresponding quadrilaterals around the text area to generate the annotation of the quadrilateral box. However, when the quadrilateral area is subsequently transformed into a horizontal rectangle through perspective transformation, it will cause the text in the area to be distorted, resulting in poor quality of the text area slice. In addition, this text annotation method cannot determine the orientation of the text area during the annotation process, and therefore cannot guarantee the orientation of the text after the text area is cut out from the image, resulting in poor quality of the text area slice. Summary of the Invention

[0005] This application provides a method for generating labeled samples based on rectangular patterns, an image processing method, an apparatus, a terminal, and a storage medium, which realizes the automatic generation of text boundary rectangles, improving the efficiency of image labeling; and determines the direction of text in the image, improving the quality of text region slicing.

[0006] This application provides a method for generating labeled samples based on rectangular patterns, including:

[0007] In response to user actions on the target image, the coordinates of three scene points on the target image are determined; wherein, the first scene point and the second scene point are located above the text content area in the target image; the third scene point is used to constrain the bottom edge boundary, or to constrain the bottom edge boundary and correct the left and right edge boundaries.

[0008] Based on the positional relationship between the first scene point, the second scene point, and the third scene point, determine the first straight line containing the top edge of the rectangular text box and the second straight line containing the bottom edge;

[0009] Based on the first line and the second line, and in conjunction with a preset rectangular coordinate system, the perpendicular feet of the three scene points to the first line are determined respectively. Based on the positional relationship between the perpendicular feet, the first line, and the second line, the coordinates of the four vertices of the rectangular text box are determined to generate the rectangular text box.

[0010] Based on the rectangular text box and the target image, annotated samples are generated.

[0011] The rectangular pattern-based annotation sample generation method of this application has the following effects: This application uses the coordinates of three scene points at preset positions to determine the positions of the four vertices of four rectangular text boxes, thereby automatically generating text boxes with the shape of rotated rectangles (including the special case of horizontal rectangles), improving the efficiency of image annotation. In addition, this application determines the top and bottom edges of the rectangular text boxes, thereby determining the direction of the text within the rectangular text boxes during the generation process.

[0012] Furthermore, determining the coordinates of three scene points on the target image also includes:

[0013] If the user inputs the coordinates of more than three scene points, then the coordinates of the first three scene points input by the user will be obtained according to the order of input.

[0014] Wherein, the first scene point and the second scene point are located on the top edge of the initial rectangle;

[0015] The initial rectangle is set by the user based on the target text of the target image;

[0016] The top edge of the initial rectangle is the edge above the target text when the target text is viewed directly.

[0017] The distance between the first scene point and the vertex at one end of the top edge is not greater than the first distance;

[0018] The third scene point is located on the bottom edge of the initial rectangle;

[0019] The position of the third scene point is set according to the distance between the second scene point and the vertex at the other end of the top edge;

[0020] The bottom edge of the initial rectangle is the edge below the target text when the target text is viewed directly.

[0021] Furthermore, the position of the third scene point is set based on the distance between the second scene point and the vertex at the other end of the top edge, specifically:

[0022] If the distance between the second scene point and the vertex at the other end of the top edge is not greater than the first distance, then the setting range of the third scene point is: any point on the target projection line segment located on the bottom edge of the initial rectangle; the target projection line segment is the line segment between the projection points of the first scene point and the second scene point on the bottom edge;

[0023] If the distance between the second scene point and the vertex at the other end of the top edge is greater than the first distance, then the setting range of the third scene point is: located on the bottom edge of the initial rectangle, and the distance between it and the target vertex is not greater than the second distance; the target vertex and the vertex at one end of the top edge are the diagonal vertices of the initial rectangle.

[0024] This application uses three scene points set by the user according to the above conditions. The first and second scene points jointly constrain the top boundary of the rectangular text box, while the third scene point constrains the bottom boundary. The first, second, and third scene points together constrain the left and right boundaries of the rectangular text box. Therefore, based on the three scene points set by the user, the four vertices of the rectangular text box can be completed, automatically generating a rotated rectangular text box that meets the text region requirements. This eliminates the need to manually drag the horizontal rectangle to the appropriate rotation angle, improving the efficiency of target image annotation.

[0025] Furthermore, based on the positional relationship between the first scene point, the second scene point, and the third scene point, the first straight line containing the top edge of the rectangular text box and the second straight line containing the bottom edge are determined, specifically as follows:

[0026] Connect the first scene point and the second scene point to obtain the first straight line containing the top edge of the rectangular text box;

[0027] Draw a line parallel to the first line through the third scene point to obtain the second line containing the bottom edge of the rectangular text box.

[0028] Furthermore, based on the first straight line and the second straight line, and in conjunction with a preset rectangular coordinate system, the perpendicular feet of the three scene points to the first straight line are determined respectively, specifically as follows:

[0029] Draw the first perpendicular line from the first scene point to the first straight line, with the foot of the perpendicular being the first foot of the perpendicular.

[0030] Draw the second perpendicular line from the second scene point to the first straight line, with the foot of the perpendicular being the second foot of the perpendicular.

[0031] Draw the third perpendicular line from the third scene point to the first straight line, with the foot of the perpendicular being the third foot of the perpendicular.

[0032] Furthermore, based on the positional relationship between the perpendicular feet, the first line, and the second line, the coordinates of the four vertices of the rectangular text box are determined, specifically as follows:

[0033] Based on the positional relationship between each perpendicular foot and the first straight line, determine the first vertex and the second vertex on the top edge of the rectangular text box;

[0034] Based on the positional relationship between the first vertex, the second vertex, and the second straight line, determine the third and fourth vertices on the bottom edge of the rectangular text box.

[0035] Furthermore, based on the positional relationship between each perpendicular foot and the first straight line, the first vertex and the second vertex on the top edge of the rectangular text box are determined, specifically as follows:

[0036] If the first line is parallel to the vertical axis of the preset rectangular coordinate system, then the vertical coordinate values ​​of the first foot, the second foot, and the third foot are compared to determine the first vertex and the second vertex, and the coordinates of the first vertex and the second vertex are obtained.

[0037] Wherein, the first vertex is the foot of the perpendicular with the largest ordinate value, and the second vertex is the foot of the perpendicular with the smallest ordinate value;

[0038] Alternatively, the first vertex is the foot of the perpendicular with the smallest ordinate value, and the second vertex is the foot of the perpendicular with the largest ordinate value;

[0039] If the first straight line is not parallel to the vertical axis of the preset rectangular coordinate system, then the horizontal coordinates of the first perpendicular foot, the second perpendicular foot, and the third perpendicular foot are compared to determine the first vertex and the second vertex, and the coordinates of the first vertex and the second vertex are obtained.

[0040] Wherein, the first vertex is the foot of the perpendicular with the largest x-coordinate value, and the second vertex is the foot of the perpendicular with the smallest x-coordinate value;

[0041] Alternatively, the first vertex is the foot of the perpendicular with the smallest x-coordinate value, and the second vertex is the foot of the perpendicular with the largest x-coordinate value.

[0042] Furthermore, based on the positional relationship between the first vertex, the second vertex, and the second straight line, the third and fourth vertices on the bottom edge of the rectangular text box are determined, specifically as follows:

[0043] The perpendicular line from the first vertex to the first line is taken as the third line;

[0044] The perpendicular line from the second vertex to the first line is taken as the fourth line;

[0045] The intersection of the second line and the third line is taken as the third vertex, and the coordinates of the third vertex are obtained;

[0046] The intersection of the second line and the fourth line is taken as the fourth vertex, and the coordinates of the fourth vertex are obtained.

[0047] The annotation sample generation method based on rectangular patterns provided in this application first determines the first straight line containing the top edge of the rectangular text box and the second straight line containing the bottom edge, based on three scene points set by the user; then, combined with a preset rectangular coordinate system, the perpendicular feet of the three scene points to the first straight line are determined respectively; according to the positional relationship between the perpendicular feet, the first straight line, and the second straight line, the four vertices of the rectangular text box are completed to generate the rectangular text box; then, a rotating rectangular text box that meets the requirements of the text area is automatically generated, and the direction of the text inside the rectangular text box is determined according to the top and bottom edges of the rectangular text box.

[0048] Furthermore, after determining the third and fourth vertices on the bottom edge of the rectangular text box, the method further includes:

[0049] Use the distance between the first vertex and the second vertex as the length of the rectangular text box;

[0050] Use the distance between the first and third vertices as the width of the rectangular text box.

[0051] Further, the rectangular text box is generated as follows:

[0052] Based on the positional relationship of the first vertex, second vertex, third vertex and fourth vertex of the rectangular text box, the first vertex, second vertex, third vertex and fourth vertex are sorted to obtain the first coordinate point, second coordinate point, third coordinate point and fourth coordinate point in sequence;

[0053] A path is generated by connecting the first coordinate point, the second coordinate point, the third coordinate point, and the fourth coordinate point using a preset component;

[0054] Instantiate the path to obtain a rectangular text box.

[0055] Furthermore, after instantiating the path to obtain a rectangular text box, the method further includes:

[0056] In response to the user's dragging and modification of the rectangular text box, the coordinates of the first coordinate point, the second coordinate point, the third coordinate point, and the fourth coordinate point, as well as the length and width of the rectangular text box, are updated.

[0057] Furthermore, after instantiating the path to obtain a rectangular text box, the method further includes:

[0058] Starting from the first coordinate point, the first straight line generates a text editing box of a preset size;

[0059] The text editing box is instantiated, and the changes in the text content within the text editing box are monitored; the text content within the text editing box is entered by the user.

[0060] If a change occurs, the latest entered text content is retrieved and displayed in the text editing box;

[0061] Save the latest input text content as text annotation content.

[0062] This application allows the generated rectangular text boxes to be dragged and modified with the mouse, enabling manual fine-tuning based on the automatically generated rectangular text boxes and updating the coordinate and size data of the rectangular text boxes, thereby improving the efficiency and accuracy of target image annotation.

[0063] Accordingly, this application also provides an image processing method based on rectangular patterns, comprising: acquiring a target image, and acquiring a labeled sample corresponding to the target image according to the labeled sample generation method based on rectangular patterns as described in this application;

[0064] Based on the labeled samples, a perspective transformation is performed on the text region within the rectangular text box on the target image to obtain an image slice.

[0065] Furthermore, based on the labeled samples, a perspective transformation is performed on the text region within the rectangular text box on the target image to obtain image slices, specifically:

[0066] Use the height of the rectangular text box as the height of the target rectangle, and the width of the rectangular text box as the width of the target rectangle to generate a list of coordinates for the target rectangle;

[0067] The coordinate list includes the coordinates of the four vertices of the target rectangle; the coordinates of the vertices of the target rectangle are preset to be the origin coordinates; the order of the four vertices is strongly correlated with the order of the four vertices of the rectangular text box.

[0068] Based on the coordinate list and the coordinates of the first, second, third, and fourth coordinate points, a perspective transformation matrix is ​​calculated to perform perspective transformation on the target image, thereby obtaining the text region image within the rectangular text box.

[0069] The text region image is used as an image slice.

[0070] Furthermore, after performing perspective transformation on the text region within the rectangular text box on the target image to obtain image slices, the process further includes:

[0071] Get the text annotations corresponding to the image slices;

[0072] Associate and save text annotations with image slices.

[0073] The image processing method based on rectangular patterns proposed in this application has the following advantages: Since the text box generated by this application is rectangular, the region within the rectangular text box is perspective-transformed into a horizontal rectangle, allowing the text in the generated image slices to retain its original shape. Furthermore, because this application determines the top and bottom edges of the rectangular text box, when performing perspective transformation slicing on the region within the rectangular text box in the target image, the text in the image slices can maintain the correct orientation, improving the quality of the text region slices. This facilitates using the obtained text region slices and text annotations as text recognition data to create a training dataset for the text recognition model.

[0074] Accordingly, this application also provides a device for generating labeled samples based on rectangular patterns, including: a scene point data acquisition module, a data processing module, and a labeling module;

[0075] The scene point data acquisition module is used to respond to the user's operation on the target image and determine the coordinates of three scene points on the target image; wherein, the first scene point and the second scene point of the three scene points are located above the area of ​​text content in the target image; the third scene point is used to constrain the bottom edge boundary, or to constrain the bottom edge boundary and correct the left and right edge boundaries.

[0076] The data processing module is used to determine the first straight line where the top edge of the rectangular text box is located and the second straight line where the bottom edge is located, based on the positional relationship between the first scene point, the second scene point and the third scene point.

[0077] The annotation module is used to determine the perpendicular foot of each of the three scene points to the first line based on the first line and the second line, combined with a preset rectangular coordinate system, and to determine the coordinates of the four vertices of the rectangular text box based on the positional relationship between each perpendicular foot, the first line and the second line, so as to generate the rectangular text box.

[0078] Based on the rectangular text box and the target image, annotated samples are generated.

[0079] The rectangular pattern-based annotation sample generation device of this application has the following advantages: The device uses the coordinates of three scene points at preset locations to determine the positions of the four vertices of four rectangular text boxes, thereby automatically generating text boxes with the shape of rotated rectangles, improving the efficiency of image annotation. Furthermore, the device determines the top and bottom edges of the rectangular text boxes, thus determining the orientation of the text within the rectangular text boxes during the generation process.

[0080] Accordingly, this application also provides an image processing device based on rectangular patterns, including: a labeled data acquisition module and a perspective transformation module;

[0081] The annotation data acquisition module is used to acquire the target image and, according to the annotation sample generation method based on rectangular patterns as described in this invention, acquire the annotation sample corresponding to the target image.

[0082] The perspective transformation module is used to perform perspective transformation on the text region within the rectangular text box on the target image based on the labeled sample, so as to obtain image slices.

[0083] The image processing apparatus based on rectangular patterns of this application has the following advantages: Since the text boxes in the labeled samples acquired by the annotation data acquisition module of this application are rectangular, the perspective transformation module transforms the region within the rectangular text box into a horizontal rectangle, and the text in the generated image slices retains its original shape. Furthermore, since this application determines the top and bottom edges of the rectangular text boxes, when the perspective transformation module performs perspective transformation slicing on the region within the rectangular text boxes in the target image, the text in the image slices can maintain the correct orientation, improving the quality of the text region slices. This facilitates using the obtained text region slices and text annotations as text recognition data to create a training dataset for the text recognition model.

[0084] Accordingly, this application also provides a terminal device, comprising: at least one processor and a memory; the memory being used to store program instructions; the processor being used to call and execute the program instructions stored in the memory, so that the terminal device performs the rectangular pattern-based annotation sample generation method as described in this invention, or performs the rectangular pattern-based image processing method as described in this invention.

[0085] Accordingly, this application also provides a computer-readable storage medium, characterized in that the computer-readable storage medium includes a stored computer program; wherein, when the computer program is running, it controls the device where the computer-readable storage medium is located to execute the rectangular pattern-based annotation sample generation method as described in the present invention, or to execute the rectangular pattern-based image processing method as described in the present invention. Attached Figure Description

[0086] Figure 1 This is a flowchart illustrating an embodiment of the method for generating labeled samples based on rectangular patterns provided in this application.

[0087] Figure 2 This is a schematic diagram (I) showing the setup of the three scene points provided in this application;

[0088] Figure 3 This is a schematic diagram (II) showing the setup of the three scene points provided in this application;

[0089] Figure 4 This is a schematic diagram (III) showing the setup of the three scene points provided in this application;

[0090] Figure 5 This is a schematic diagram illustrating the calculation of the vertices of the rectangular text box provided in this application;

[0091] Figure 6 This is a schematic diagram (I) of an existing rectangular pattern annotation sample;

[0092] Figure 7 This is a schematic diagram (II) of an existing rectangular pattern annotation sample;

[0093] Figure 8 This is a schematic diagram (III) of an existing rectangular pattern annotation sample;

[0094] Figure 9 IV is a schematic diagram of an existing rectangular pattern annotation sample;

[0095] Figure 10 This is a schematic diagram (I) of the annotation software interface provided in this application;

[0096] Figure 11 This is a schematic diagram (II) of the annotation software interface provided in this application;

[0097] Figure 12 This is a flowchart illustrating an embodiment of the image processing method based on rectangular patterns provided in this application;

[0098] Figure 13 This is a schematic diagram of the structure of an embodiment of the rectangular pattern-based annotation sample generation device provided in this application;

[0099] Figure 14 This is a schematic diagram of an embodiment of the image processing device based on rectangular patterns provided in this application. Detailed Implementation

[0100] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0101] Example 1

[0102] The method for generating labeled samples based on rectangular patterns provided in this application is applied in labeling software in some preferred embodiments. The image region in the labeling software is divided into an inner layer and a surface layer. The inner layer stores the target image to be labeled, and the surface layer is used to observe the image within the image region. The image region in the labeling software is configured with a preset coordinate system.

[0103] In some preferred implementations, the annotation software is implemented using PySide6; the inner layer is created using the QGraphicsScene scene, and the outer layer is created using the QGraphicsView view class; the annotated image is displayed using QGraphicsPixmapItem.

[0104] In some preferred embodiments, the coordinates used for calculation within the annotation software are all scene coordinates, while the saved annotation data is converted into annotation coordinates. If the annotation coordinates of the target image are in a different coordinate system than the scene coordinates of the target image, the mapToItem method is used to calculate the annotation coordinates of the scene coordinates relative to the target image. The coordinate system in which the annotation coordinates are located takes the vertex of the upper left corner of the target image as the origin, the horizontal rightward direction is the positive x-direction, and the vertical downward direction is the positive y-direction.

[0105] Please refer to Figure 1 The present application provides a method for generating labeled samples based on rectangular patterns, comprising steps S101-S104:

[0106] Step S101: In response to the user's operation on the target image, determine the coordinates of three scene points on the target image; wherein, the first scene point and the second scene point of the three scene points are located above the area of ​​text content in the target image; the third scene point is used to constrain the bottom edge boundary, or to constrain the bottom edge boundary and correct the left and right edge boundaries.

[0107] Furthermore, determining the coordinates of three scene points on the target image further includes:

[0108] If the user inputs the coordinates of more than three scene points, then the coordinates of the first three scene points input by the user will be obtained according to the order of input.

[0109] Wherein, the first scene point and the second scene point are located on the top edge of the initial rectangle;

[0110] The initial rectangle is set by the user based on the target text of the target image;

[0111] The top edge of the initial rectangle is the edge above the target text when the target text is viewed directly.

[0112] The distance between the first scene point and the vertex at one end of the top edge is not greater than the first distance;

[0113] The third scene point is located on the bottom edge of the initial rectangle;

[0114] The position of the third scene point is set according to the distance between the second scene point and the vertex at the other end of the top edge;

[0115] The bottom edge of the initial rectangle is the edge below the target text when the target text is viewed directly.

[0116] Furthermore, the position of the third scene point is set based on the distance between the second scene point and the vertex at the other end of the top edge, specifically:

[0117] If the distance between the second scene point and the vertex at the other end of the top edge is not greater than the first distance, then the setting range of the third scene point is: any point on the target projection line segment located on the bottom edge of the initial rectangle; the target projection line segment is the line segment between the projection points of the first scene point and the second scene point on the bottom edge;

[0118] If the distance between the second scene point and the vertex at the other end of the top edge is greater than the first distance, then the setting range of the third scene point is: located on the bottom edge of the initial rectangle, and the distance between it and the target vertex is not greater than the second distance; the target vertex and the vertex at one end of the top edge are the diagonal vertices of the initial rectangle.

[0119] In some preferred embodiments, the user's operation on the target image is that the user clicks on the target image with the mouse, and the scene coordinates of the click position are used as the scene coordinates input by the user. The scene coordinates of the user's click position are obtained by the scenePos method of the mouse event in PySide6.

[0120] The annotation software acquires the mouse click events of the first three users and obtains three scene coordinates generated by these mouse click events. If a user clicks the target image more than three times, that is, the user inputs more than three scene coordinates, then the first three scene coordinates are selected.

[0121] The coordinates and input order of the three scene points set by the user in this application need to meet preset requirements. In some preferred embodiments, before the user clicks on the target image with the mouse, an imaginary rectangular area, i.e., an initial rectangle, is defined by the user based on the target text in the target image and is not constructed in the annotation software. Specifically, the edge closest to the top of the text when viewed directly is the top edge of the initial rectangle, and the edge closest to the bottom of the text is the bottom edge of the initial rectangle.

[0122] The three scene points set by the user are the first scene point, the second scene point, and the third scene point, respectively.

[0123] The first scene point and the second scene point are located on the straight line where the top edge of the initial rectangle lies, and are used to constrain the top edge boundary of the rectangular text box; the order of the first scene point and the second scene point can be interchanged.

[0124] The third scene point is located on the straight line where the bottom edge of the initial rectangle lies, and is used to constrain the bottom boundary of the rectangular text box and correct the left and right boundaries of the rectangular text box.

[0125] In some preferred embodiments, the first scene point and the second scene point lie on the straight line containing the top edge of the initial rectangle. The distance between the first scene point and a vertex at one end of the top edge is no greater than a first distance, and the distance between the second scene point and a vertex at the other end of the top edge is no greater than the first distance; that is, the first scene point and the second scene point are respectively close to two vertices of the top edge. In this case, any point can be selected as the third scene point on the target projection line segment of the bottom edge of the initial rectangle; the target projection line segment is the line segment between the projection points of the first scene point and the second scene point on the bottom edge.

[0126] For example, please refer to Figure 2 and Figure 3 , Figure 2 and Figure 3 Here are specific examples of setting the first scene point a, the second scene point b, and the third scene point c, respectively, where the first scene point a is close to the initial rectangle. Figure 2 and Figure 3 The left vertex of the top edge of the first scene point (the initial rectangle is not drawn), the right vertex of the second scene point b is close to the top edge; the third scene point c is one of the points on the line segment between the projection points of the first scene point a and the second scene point b on the bottom edge.

[0127] As one preferred implementation, the first scene point and the second scene point jointly constrain the top boundary of the rectangular text box, and at the same time, constrain the left and right boundaries of the rectangular text box; the third scene point constrains the bottom boundary of the rectangular text box.

[0128] In some preferred embodiments, the first scene point and the second scene point lie on the straight line containing the top edge of the initial rectangle. The distance between the first scene point and a vertex at one end of the top edge is no greater than a first distance, while the distance between the second scene point and a vertex at the other end of the top edge is greater than the first distance. That is, the first scene point is close to one of the vertices of the top edge, while the second scene point is not close to the other vertex. It should be noted that at least one of the first and second scene points must be close to one of the two vertices of the top edge. In this case, the distance between the third scene point and the target vertex must be no greater than a second distance; the target vertex and the vertex at one end of the top edge are diagonal vertices of the initial rectangle.

[0129] For example, please refer to Figure 4 ( Figure 4 (The initial rectangle has not been drawn.) Figure 4 Another specific implementation scheme is provided, where the first scene point a, the second scene point b, and the third scene point c are set. The first scene point a is close to the left vertex of the top edge, but the second scene point b is not close to the right vertex of the top edge. Therefore, the third scene point c needs to be as close as possible to the right vertex of the bottom edge to constrain the bottom boundary of the rectangular text box while correcting its right boundary.

[0130] In one preferred embodiment, the first scene point and the second scene point jointly constrain the top boundary of the rectangular text box, while the first scene point constrains the left boundary of the rectangular text box; the third scene point constrains the bottom boundary of the rectangular text box, while the third scene point corrects the right boundary of the rectangular text box.

[0131] This application uses three scene points set by the user according to the above conditions. The first and second scene points jointly constrain the top boundary of the rectangular text box, while the third scene point constrains the bottom boundary. The first, second, and third scene points together constrain the left and right boundaries of the rectangular text box. Therefore, based on the three scene points set by the user, the four vertices of the rectangular text box can be completed, automatically generating a rotated rectangular text box that meets the text region requirements. This eliminates the need to manually drag the horizontal rectangle to the appropriate rotation angle, improving the efficiency of target image annotation.

[0132] Step S102: Based on the positional relationship between the first scene point, the second scene point, and the third scene point, determine the first straight line where the top edge of the rectangular text box is located and the second straight line where the bottom edge is located;

[0133] Furthermore, based on the positional relationship between the first scene point, the second scene point, and the third scene point, the first straight line containing the top edge of the rectangular text box and the second straight line containing the bottom edge are determined, specifically as follows:

[0134] Connect the first scene point and the second scene point to obtain the first straight line containing the top edge of the rectangular text box;

[0135] Draw a line parallel to the first line through the third scene point to obtain the second line containing the bottom edge of the rectangular text box.

[0136] Step S103: Based on the first straight line and the second straight line, and in conjunction with the preset rectangular coordinate system, determine the perpendicular foot of each of the three scene points to the first straight line, and determine the coordinates of the four vertices of the rectangular text box according to the positional relationship between each perpendicular foot, the first straight line and the second straight line, so as to generate the rectangular text box;

[0137] Furthermore, based on the first straight line and the second straight line, and in conjunction with a preset rectangular coordinate system, the perpendicular feet of the three scene points to the first straight line are determined respectively, specifically as follows:

[0138] Draw the first perpendicular line from the first scene point to the first straight line, with the foot of the perpendicular being the first foot of the perpendicular.

[0139] Draw the second perpendicular line from the second scene point to the first straight line, with the foot of the perpendicular being the second foot of the perpendicular.

[0140] Draw the third perpendicular line from the third scene point to the first straight line, with the foot of the perpendicular being the third foot of the perpendicular.

[0141] Furthermore, based on the positional relationship between the perpendicular feet, the first line, and the second line, the coordinates of the four vertices of the rectangular text box are determined, specifically as follows:

[0142] Based on the positional relationship between each perpendicular foot and the first straight line, determine the first vertex and the second vertex on the top edge of the rectangular text box;

[0143] Based on the positional relationship between the first vertex, the second vertex, and the second straight line, determine the third and fourth vertices on the bottom edge of the rectangular text box.

[0144] Furthermore, based on the positional relationship between each perpendicular foot and the first straight line, the first vertex and the second vertex on the top edge of the rectangular text box are determined, specifically as follows:

[0145] If the first line is parallel to the vertical axis of the preset rectangular coordinate system, then the vertical coordinate values ​​of the first foot, the second foot, and the third foot are compared to determine the first vertex and the second vertex, and the coordinates of the first vertex and the second vertex are obtained.

[0146] Wherein, the first vertex is the foot of the perpendicular with the largest ordinate value, and the second vertex is the foot of the perpendicular with the smallest ordinate value;

[0147] Alternatively, the first vertex is the foot of the perpendicular with the smallest ordinate value, and the second vertex is the foot of the perpendicular with the largest ordinate value;

[0148] If the first straight line is not parallel to the vertical axis of the preset rectangular coordinate system, then the horizontal coordinates of the first perpendicular foot, the second perpendicular foot, and the third perpendicular foot are compared to determine the first vertex and the second vertex, and the coordinates of the first vertex and the second vertex are obtained.

[0149] Wherein, the first vertex is the foot of the perpendicular with the largest x-coordinate value, and the second vertex is the foot of the perpendicular with the smallest x-coordinate value;

[0150] Alternatively, the first vertex is the foot of the perpendicular with the smallest x-coordinate value, and the second vertex is the foot of the perpendicular with the largest x-coordinate value.

[0151] Furthermore, based on the positional relationship between the first vertex, the second vertex, and the second straight line, the third and fourth vertices on the bottom edge of the rectangular text box are determined, specifically as follows:

[0152] The perpendicular line from the first vertex to the first line is taken as the third line;

[0153] The perpendicular line from the second vertex to the first line is taken as the fourth line;

[0154] The intersection of the second line and the third line is taken as the third vertex, and the coordinates of the third vertex are obtained;

[0155] The intersection of the second line and the fourth line is taken as the fourth vertex, and the coordinates of the fourth vertex are obtained.

[0156] Furthermore, after determining the third and fourth vertices on the bottom edge of the rectangular text box, the method further includes:

[0157] Use the distance between the first vertex and the second vertex as the length of the rectangular text box;

[0158] Use the distance between the first and third vertices as the width of the rectangular text box.

[0159] In some preferred embodiments, please refer to Figure 5 The preset rectangular coordinate system takes the vertex at the top left corner of the target image as the origin, with the horizontal axis pointing horizontally to the right and the vertical axis pointing vertically downwards.

[0160] Given the first scene point a, the second scene point b, and the third scene point c, connect the first scene point a and the second scene point b to obtain the first straight line containing the top edge of the rectangular text box.

[0161] Draw a line parallel to the first line through the third scene point c, and obtain the second line containing the bottom edge of the rectangular text box.

[0162] Draw the first perpendicular line from the first scene point a to the first straight line, with the foot of the perpendicular being the first foot of the perpendicular P1.

[0163] Draw the second perpendicular line from the second scene point b to the first straight line, with the foot of the perpendicular being the second perpendicular foot P2.

[0164] Draw the third perpendicular line from point c in the third scene to the first straight line, with the foot of the perpendicular being P3.

[0165] If the first line is parallel to the vertical axis, then among the first, second, and third perpendiculars, the perpendicular with the largest vertical coordinate value is taken as the first vertex, and the perpendicular with the smallest vertical coordinate value is taken as the second vertex.

[0166] If the first straight line is not parallel to the vertical axis, then among the first, second, and third perpendiculars, the perpendicular with the largest x-coordinate value is taken as the first vertex, and the perpendicular with the smallest x-coordinate value is taken as the second vertex.

[0167] exist Figure 5 In the diagram, the first straight line containing the top edge is not parallel to the vertical axis. Among the first perpendicular foot P1, the second perpendicular foot P2, and the third perpendicular foot P3, the vertical coordinate value of the first perpendicular foot P1 is the largest. Therefore, the first perpendicular foot P1 is taken as the first vertex P_MAX, i.e. Figure 5 Point b in the middle; the ordinate value of the second perpendicular foot P2 is the smallest, so the second perpendicular foot P2 is taken as the second vertex P_MIN, that is... Figure 5 Point a in the middle.

[0168] The perpendicular line from the first foot of the perpendicular P1 to the first line is the third line, and the perpendicular line from the second foot of the perpendicular P2 to the first line is the fourth line.

[0169] This application uses the intersection of the second and third lines containing the bottom edge as the third vertex P_PA_MAX, that is... Figure 5 Point f in the middle; this application takes the intersection of the second line containing the bottom edge and the fourth line as the fourth vertex P_PA_MIN, that is Figure 5 Point e in the middle.

[0170] The annotation sample generation method based on rectangular patterns provided in this application first determines the first straight line containing the top edge of the rectangular text box and the second straight line containing the bottom edge, based on three scene points set by the user; then, combined with a preset rectangular coordinate system, the perpendicular feet of the three scene points to the first straight line are determined respectively; according to the positional relationship between the perpendicular feet, the first straight line, and the second straight line, the four vertices of the rectangular text box are completed to generate the rectangular text box; then, a rotating rectangular text box that meets the requirements of the text area is automatically generated, and the direction of the text inside the rectangular text box is determined according to the top and bottom edges of the rectangular text box.

[0171] Further, the rectangular text box is generated as follows:

[0172] Based on the positional relationship of the first vertex, second vertex, third vertex and fourth vertex of the rectangular text box, the first vertex, second vertex, third vertex and fourth vertex are sorted to obtain the first coordinate point, second coordinate point, third coordinate point and fourth coordinate point in sequence;

[0173] A path is generated by connecting the first coordinate point, the second coordinate point, the third coordinate point, and the fourth coordinate point using a preset component;

[0174] Instantiate the path to obtain a rectangular text box.

[0175] In some preferred embodiments, QPainterPath is used to load the first, second, third, and fourth coordinate points. When instantiating QPainterPath, the startPoint parameter is passed as the fourth coordinate point. The lineTo method is used to connect the four coordinate points sequentially, starting from the first coordinate point, to generate the path. The generated path is then loaded using the setPath method of QGraphicsPathItem.

[0176] Besides the QGraphicsPathItem component in the PySide software interface library, tools for loading a path formed by four coordinate points include, but are not limited to, loading tools in PyQt. Both PySide and PyQt are Python implementations of Qt, a cross-platform C++ library used for developing GUI applications.

[0177] In some preferred embodiments, the preset rectangular coordinate system takes the vertex of the upper left corner of the target image as the origin, with the horizontal axis pointing horizontally to the right and the vertical axis pointing vertically downward.

[0178] Based on the positional relationship of the first, second, third, and fourth vertices of the rectangular text box, the first, second, third, and fourth vertices are sorted to obtain the first, second, third, and fourth coordinate points in sequence, specifically:

[0179] If the first straight line is parallel to the vertical axis of the preset rectangular coordinate system, and the x-coordinate value of the second vertex is greater than the x-coordinate value of the fourth vertex, then the first coordinate point, the second coordinate point, the third coordinate point, and the fourth coordinate point are, in order: the second vertex, the fourth vertex, the third vertex, and the first vertex.

[0180] If the first straight line is parallel to the vertical axis of the preset rectangular coordinate system, and the x-coordinate value of the second vertex is less than the x-coordinate value of the fourth vertex, then the first coordinate point, the second coordinate point, the third coordinate point, and the fourth coordinate point are, in order: the first vertex, the third vertex, the fourth vertex, and the second vertex.

[0181] If the first straight line is not parallel to the vertical axis of the preset rectangular coordinate system, and the vertical coordinate value of the second vertex is less than the vertical coordinate value of the fourth vertex, then the first coordinate point, the second coordinate point, the third coordinate point, and the fourth coordinate point are, in order: the second vertex, the fourth vertex, the third vertex, and the first vertex.

[0182] If the first straight line is not parallel to the vertical axis of the preset rectangular coordinate system, and the vertical coordinate value of the second vertex is greater than the vertical coordinate value of the fourth vertex, then the first coordinate point, the second coordinate point, the third coordinate point, and the fourth coordinate point are, in order: the second vertex, the fourth vertex, the third vertex, and the first vertex.

[0183] For example, please refer to Figure 5 , Figure 5 The four vertices of the provided rectangular text box are point a, point b, point e, and point f.

[0184] The order of the four vertices of the rectangular text box, namely the first, second, third, and fourth coordinate points of the rectangular text box, is: point a, point e, point f, and point b.

[0185] It should be noted that the above-described vertex sorting method is only a sorting method in some preferred embodiments. This application can set different vertex sorting methods, such as a vertex sorting method that starts from one vertex and sorts in a clockwise or counterclockwise direction.

[0186] This application sorts the four vertices of a rectangular text box so that when the rectangular text box is instantiated, the four vertices can be connected in the order they appear.

[0187] In addition, the list of coordinates of the target rectangle generated in the image processing method based on rectangular patterns provided in this application includes the coordinates of the four vertices of the target rectangle.

[0188] This application requires establishing a one-to-one correspondence between the four vertices of a rectangular text box and the four vertices of a target rectangle. Therefore, the four vertices of the rectangular text box need to be sorted to correspond to the order of the four vertices of the target rectangle in the coordinate list. Furthermore, the image processing method based on rectangular patterns provided in this application performs perspective transformation on the area within the rectangular text box according to the order of the four vertices of the rectangular text box and the coordinate list of the target rectangle. This ensures that the text in the generated image slice remains undistorted and guarantees the correct orientation of the text.

[0189] Furthermore, after instantiating the path to obtain a rectangular text box, the method further includes:

[0190] In response to the user's dragging and modification of the rectangular text box, the coordinates of the first coordinate point, the second coordinate point, the third coordinate point, and the fourth coordinate point, as well as the length and width of the rectangular text box, are updated.

[0191] In some preferred embodiments, this application allows the generated rectangular text box to be modified by dragging the mouse, and manual fine-tuning can be performed on the basis of the automatically generated rectangular text box. Specifically, the mouse drag modification of the rectangular text box is implemented by QGraphicsItem.

[0192] After the rectangular text box is modified by dragging the mouse, a new rectangular text box is created, along with four new scene coordinates. The annotation software needs to convert these four new scene coordinates into annotation coordinates to obtain the four annotation coordinates of the rectangular text box. Based on the four annotation coordinates of the rectangular text box, the annotation software updates the coordinates of the first, second, third, and fourth coordinate points, as well as the length and width of the rectangular text box.

[0193] Furthermore, after instantiating the path to obtain a rectangular text box, the method further includes:

[0194] Starting from the first coordinate point, the first straight line generates a text editing box of a preset size;

[0195] The text editing box is instantiated, and the changes in the text content within the text editing box are monitored; the text content within the text editing box is entered by the user.

[0196] If a change occurs, the latest entered text content is retrieved and displayed in the text editing box;

[0197] Save the latest input text content as text annotation content.

[0198] In some preferred embodiments, this application uses its setPos method, which takes the first coordinate point of four scene coordinates as the input parameter, and instantiates a text edit box QTextDocument starting from the position of the first coordinate point. After instantiation, it listens for changes in the text content in QTextDocument. If a change occurs, the new text content is retrieved. The default value is empty. This listening method is provided by the contentsChanged signal of QTextDocument.

[0199] Step S104: Generate labeled samples based on the rectangular text box and the target image.

[0200] Please refer to Figure 6 The existing rectangular pattern annotation sample uses a horizontal rectangle to annotate the sample and records text content above the horizontal rectangle. However, the text in the sample is slanted, and there is a large deviation between the horizontal rectangular area and the text area.

[0201] Please refer to Figure 7 The existing rectangular pattern annotation sample first marks a horizontal rectangle, then gradually drags and rotates the rectangle to a suitable rotation angle to obtain a rotated rectangle annotation. However, the text in the sample is slanted, and this existing technique requires manual adjustment of the rectangle's rotation angle, resulting in low annotation efficiency.

[0202] Please refer to Figure 8 The existing rectangular pattern annotation sample uses a technique that sets four points around the text area to obtain a quadrilateral text box. If the resulting quadrilateral text box is not rectangular, performing a perspective transformation on the text area within the quadrilateral text box will cause the quadrilateral to be perceived as a horizontal rectangle, resulting in text distortion.

[0203] Please refer to Figure 9 The existing rectangular pattern annotation sample first sets four points around the text area to obtain a quadrilateral text box; then it finds the minimum bounding rectangle of the quadrilateral as the rectangular text box. This existing technology cannot guarantee the orientation of the text in the text area within the text box.

[0204] This application generates rectangular text boxes at the location of text areas in the target image within the annotation software interface, used to identify the regions where text is located in the image. This allows the image processing method based on rectangular patterns provided in this application to perform perspective transformation on the regions within the rectangular text boxes, generating image slices.

[0205] In addition, this application sets a text editing box at the top left corner of the rectangular text box in the annotation software interface, and listens to the text content in the text editing box to obtain the text annotation content corresponding to the image slice, so as to use the obtained text region slice as text recognition data to create a training dataset for the text recognition model.

[0206] Please refer to Figure 10 If the annotation software detects that the text content in the text editing box is "panel height 320mm", the text editing box will display this text content in the text editing box, and the text in the text editing box will not be empty.

[0207] Please refer to Figure 11 If the annotation software does not detect any text content in the text editing box, the text in the text editing box will be empty.

[0208] Implementing the embodiments of this application has the following effects:

[0209] The method for generating labeled samples based on rectangular patterns according to this application has the following advantages: This application uses the coordinates of three scene points at preset locations to determine the positions of the four vertices of four rectangular text boxes, thereby automatically generating text boxes with the shape of rotated rectangles, improving the efficiency of image annotation. Furthermore, this application determines the top and bottom edges of the rectangular text boxes, thus determining the orientation of the text within the rectangular text boxes during the generation process.

[0210] Example 2

[0211] Please refer to Figure 12 The present application provides an image processing method based on a rectangular pattern, comprising steps S201-S202:

[0212] Step S201: Obtain the target image, and obtain the annotation sample corresponding to the target image according to the annotation sample generation method based on rectangular pattern as described in Embodiment 1;

[0213] Step S202: Based on the labeled sample, perform perspective transformation on the text region within the rectangular text box on the target image to obtain an image slice.

[0214] Furthermore, based on the labeled samples, a perspective transformation is performed on the text region within the rectangular text box on the target image to obtain image slices, specifically:

[0215] Use the height of the rectangular text box as the height of the target rectangle, and the width of the rectangular text box as the width of the target rectangle to generate a list of coordinates for the target rectangle;

[0216] The coordinate list includes the coordinates of the four vertices of the target rectangle; the coordinates of the vertices of the target rectangle are preset to be the origin coordinates; the order of the four vertices is strongly correlated with the order of the four vertices of the rectangular text box.

[0217] Based on the coordinate list and the coordinates of the first, second, third, and fourth coordinate points, a perspective transformation matrix is ​​calculated to perform perspective transformation on the target image, thereby obtaining the text region image within the rectangular text box.

[0218] The text region image is used as an image slice.

[0219] In some preferred embodiments, this application uses the getPerspectiveTransform method of opencv-python to calculate the perspective transformation matrix.

[0220] The order of the coordinate points in the coordinate list is strongly correlated with the order of the coordinates of the four vertices of the rectangular text box obtained in Example 1.

[0221] For example, please refer to Figure 5 , Figure 5 The four vertices of the provided rectangular text box are in the following order: point a, point e, point f, and point b.

[0222] The list of the four coordinate points of the target rectangle is [[0,0],[0,rect_height],[rect_width,rect_height],[rect_width,0]].

[0223] Where rect_height is the height of the rectangular text box and rect_width is the width of the rectangular text box.

[0224] Furthermore, after performing perspective transformation on the text region within the rectangular text box on the target image to obtain image slices, the process further includes:

[0225] Get the text annotations corresponding to the image slices;

[0226] Associate and save text annotations with image slices.

[0227] Implementing the embodiments of this application has the following effects:

[0228] The image processing method based on rectangular patterns proposed in this application has the following advantages: Since the text box generated by this application is rectangular, the region within the rectangular text box is perspective-transformed into a horizontal rectangle, allowing the text in the generated image slices to retain its original shape. Furthermore, because this application determines the top and bottom edges of the rectangular text box, when performing perspective transformation slicing on the region within the rectangular text box in the target image, the text in the image slices can maintain the correct orientation, improving the quality of the text region slices. This facilitates using the obtained text region slices as text recognition data to create a training dataset for a text recognition model.

[0229] Example 3

[0230] Please refer to Figure 13 The present application provides a device for generating labeled samples based on rectangular patterns, comprising: a scene point data acquisition module 301, a data processing module 302, and a labeling module 303.

[0231] The scene point data acquisition module 301 is used to respond to the user's operation on the target image and determine the coordinates of three scene points on the target image; wherein, the first scene point and the second scene point of the three scene points are located above the area of ​​text content in the target image; the third scene point is used to constrain the bottom edge boundary, or to constrain the bottom edge boundary and correct the left and right edge boundaries.

[0232] The data processing module 302 is used to determine the first straight line where the top edge of the rectangular text box is located and the second straight line where the bottom edge is located based on the positional relationship between the first scene point, the second scene point and the third scene point;

[0233] The annotation module 303 is used to determine the perpendicular feet of the three scene points to the first line based on the first line and the second line, combined with a preset rectangular coordinate system, and to determine the coordinates of the four vertices of the rectangular text box based on the positional relationship between the perpendicular feet, the first line and the second line, so as to generate the rectangular text box.

[0234] Based on the rectangular text box and the target image, annotated samples are generated.

[0235] Furthermore, determining the coordinates of three scene points on the target image further includes:

[0236] If the user inputs the coordinates of more than three scene points, then the coordinates of the first three scene points input by the user will be obtained according to the order of input.

[0237] Wherein, the first scene point and the second scene point are located on the top edge of the initial rectangle;

[0238] The initial rectangle is set by the user based on the target text of the target image;

[0239] The top edge of the initial rectangle is the edge above the target text when the target text is viewed directly.

[0240] The distance between the first scene point and the vertex at one end of the top edge is not greater than the first distance;

[0241] The third scene point is located on the bottom edge of the initial rectangle;

[0242] The position of the third scene point is set according to the distance between the second scene point and the vertex at the other end of the top edge;

[0243] The bottom edge of the initial rectangle is the edge below the target text when the target text is viewed directly.

[0244] Furthermore, the position of the third scene point is set based on the distance between the second scene point and the vertex at the other end of the top edge, specifically:

[0245] If the distance between the second scene point and the vertex at the other end of the top edge is not greater than the first distance, then the setting range of the third scene point is: any point on the target projection line segment located on the bottom edge of the initial rectangle; the target projection line segment is the line segment between the projection points of the first scene point and the second scene point on the bottom edge;

[0246] If the distance between the second scene point and the vertex at the other end of the top edge is greater than the first distance, then the setting range of the third scene point is: located on the bottom edge of the initial rectangle, and the distance between it and the target vertex is not greater than the second distance; the target vertex and the vertex at one end of the top edge are the diagonal vertices of the initial rectangle.

[0247] Furthermore, based on the positional relationship between the first scene point, the second scene point, and the third scene point, the first straight line containing the top edge of the rectangular text box and the second straight line containing the bottom edge are determined, specifically as follows:

[0248] Connect the first scene point and the second scene point to obtain the first straight line containing the top edge of the rectangular text box;

[0249] Draw a line parallel to the first line through the third scene point to obtain the second line containing the bottom edge of the rectangular text box.

[0250] Furthermore, based on the first straight line and the second straight line, and in conjunction with a preset rectangular coordinate system, the perpendicular feet of the three scene points to the first straight line are determined respectively, specifically as follows:

[0251] Draw the first perpendicular line from the first scene point to the first straight line, with the foot of the perpendicular being the first foot of the perpendicular.

[0252] Draw the second perpendicular line from the second scene point to the first straight line, with the foot of the perpendicular being the second foot of the perpendicular.

[0253] Draw the third perpendicular line from the third scene point to the first straight line, with the foot of the perpendicular being the third foot of the perpendicular.

[0254] Furthermore, based on the positional relationship between the perpendicular feet, the first line, and the second line, the coordinates of the four vertices of the rectangular text box are determined, specifically as follows:

[0255] Based on the positional relationship between each perpendicular foot and the first straight line, determine the first vertex and the second vertex on the top edge of the rectangular text box;

[0256] Based on the positional relationship between the first vertex, the second vertex, and the second straight line, determine the third and fourth vertices on the bottom edge of the rectangular text box.

[0257] Furthermore, based on the positional relationship between each perpendicular foot and the first straight line, the first vertex and the second vertex on the top edge of the rectangular text box are determined, specifically as follows:

[0258] If the first line is parallel to the vertical axis of the preset rectangular coordinate system, then the vertical coordinate values ​​of the first foot, the second foot, and the third foot are compared to determine the first vertex and the second vertex, and the coordinates of the first vertex and the second vertex are obtained.

[0259] Wherein, the first vertex is the foot of the perpendicular with the largest ordinate value, and the second vertex is the foot of the perpendicular with the smallest ordinate value;

[0260] Alternatively, the first vertex is the foot of the perpendicular with the smallest ordinate value, and the second vertex is the foot of the perpendicular with the largest ordinate value;

[0261] If the first straight line is not parallel to the vertical axis of the preset rectangular coordinate system, then the horizontal coordinates of the first perpendicular foot, the second perpendicular foot, and the third perpendicular foot are compared to determine the first vertex and the second vertex, and the coordinates of the first vertex and the second vertex are obtained.

[0262] Wherein, the first vertex is the foot of the perpendicular with the largest x-coordinate value, and the second vertex is the foot of the perpendicular with the smallest x-coordinate value;

[0263] Alternatively, the first vertex is the foot of the perpendicular with the smallest x-coordinate value, and the second vertex is the foot of the perpendicular with the largest x-coordinate value.

[0264] Furthermore, based on the positional relationship between the first vertex, the second vertex, and the second straight line, the third and fourth vertices on the bottom edge of the rectangular text box are determined, specifically as follows:

[0265] The perpendicular line from the first vertex to the first line is taken as the third line;

[0266] The perpendicular line from the second vertex to the first line is taken as the fourth line;

[0267] The intersection of the second line and the third line is taken as the third vertex, and the coordinates of the third vertex are obtained;

[0268] The intersection of the second line and the fourth line is taken as the fourth vertex, and the coordinates of the fourth vertex are obtained.

[0269] Furthermore, after determining the third and fourth vertices on the bottom edge of the rectangular text box, the method further includes:

[0270] Use the distance between the first vertex and the second vertex as the length of the rectangular text box;

[0271] Use the distance between the first and third vertices as the width of the rectangular text box.

[0272] Further, the rectangular text box is generated as follows:

[0273] Based on the positional relationship of the first vertex, second vertex, third vertex and fourth vertex of the rectangular text box, the first vertex, second vertex, third vertex and fourth vertex are sorted to obtain the first coordinate point, second coordinate point, third coordinate point and fourth coordinate point in sequence;

[0274] A path is generated by connecting the first coordinate point, the second coordinate point, the third coordinate point, and the fourth coordinate point using a preset component;

[0275] Instantiate the path to obtain a rectangular text box.

[0276] Furthermore, after instantiating the path to obtain a rectangular text box, the method further includes:

[0277] In response to the user's dragging and modification of the rectangular text box, the coordinates of the first coordinate point, the second coordinate point, the third coordinate point, and the fourth coordinate point, as well as the length and width of the rectangular text box, are updated.

[0278] Furthermore, after instantiating the path to obtain a rectangular text box, the method further includes:

[0279] Starting from the first coordinate point, the first straight line generates a text editing box of a preset size;

[0280] The text editing box is instantiated, and the changes in the text content within the text editing box are monitored; the text content within the text editing box is entered by the user.

[0281] If a change occurs, the latest entered text content is retrieved and displayed in the text editing box;

[0282] Save the latest input text content as text annotation content.

[0283] The aforementioned rectangular pattern-based annotation sample generation device can implement the rectangular pattern-based annotation sample generation method described in the above method embodiments. The options in the above method embodiments are also applicable to this embodiment, and will not be detailed here. The remaining content of this application's embodiments can be referred to the content of the above method embodiments, and will not be repeated in this embodiment.

[0284] The rectangular pattern-based annotation sample generation device of this application has the following advantages: The device uses the coordinates of three scene points at preset locations to determine the positions of the four vertices of four rectangular text boxes, thereby automatically generating text boxes with the shape of rotated rectangles, improving the efficiency of image annotation. Furthermore, the device determines the top and bottom edges of the rectangular text boxes, thus determining the orientation of the text within the rectangular text boxes during the generation process.

[0285] Example 4

[0286] Please refer to Figure 14 An image processing device based on a rectangular pattern provided in this application includes: a labeling data acquisition module 401 and a perspective transformation module 402;

[0287] The annotation data acquisition module 401 is used to acquire the target image and, according to the annotation sample generation method based on rectangular pattern as described in Embodiment 1, acquire the annotation sample corresponding to the target image.

[0288] The perspective transformation module 402 is used to perform perspective transformation on the text region within the rectangular text box on the target image based on the labeled sample, so as to obtain an image slice.

[0289] Furthermore, based on the labeled samples, a perspective transformation is performed on the text region within the rectangular text box on the target image to obtain image slices, specifically:

[0290] Use the height of the rectangular text box as the height of the target rectangle, and the width of the rectangular text box as the width of the target rectangle to generate a list of coordinates for the target rectangle;

[0291] The coordinate list includes the coordinates of the four vertices of the target rectangle; the coordinates of the vertices of the target rectangle are preset to be the origin coordinates; the order of the four vertices is strongly correlated with the order of the four vertices of the rectangular text box.

[0292] Based on the coordinate list and the coordinates of the first, second, third, and fourth coordinate points, a perspective transformation matrix is ​​calculated to perform perspective transformation on the target image, thereby obtaining the text region image within the rectangular text box.

[0293] The text region image is used as an image slice.

[0294] Further, based on the labeled samples, after performing perspective transformation on the text region within the rectangular text box on the target image to obtain image slices, the process also includes:

[0295] Get the text annotations corresponding to the image slices;

[0296] Associate and save text annotations with image slices.

[0297] The image processing apparatus based on rectangular patterns described above can implement the image processing method based on rectangular patterns described in the above method embodiments. The options in the above method embodiments are also applicable to this embodiment, and will not be detailed here. The remaining contents of this application's embodiments can be referred to the contents of the above method embodiments, and will not be repeated in this embodiment.

[0298] The image processing apparatus based on rectangular patterns of this application has the following advantages: Since the text boxes in the labeled samples acquired by the annotation data acquisition module of this application are rectangular, the perspective transformation module transforms the region within the rectangular text box into a horizontal rectangle, and the text in the generated image slices retains its original shape. Furthermore, since this application determines the top and bottom edges of the rectangular text boxes, when the perspective transformation module performs perspective transformation slicing on the region within the rectangular text boxes in the target image, the text in the image slices can maintain the correct orientation, improving the quality of the text region slices. This facilitates using the obtained text region slices and text annotations as text recognition data to create a training dataset for the text recognition model.

[0299] Example 5

[0300] Accordingly, this application also provides a terminal device comprising: at least one processor and a memory; wherein the memory is used to store program instructions; and the processor is used to call and execute the program instructions stored in the memory to cause the terminal device to perform the rectangular pattern-based annotation sample generation method as described in Embodiment 1, or to perform the rectangular pattern-based image processing method as described in Embodiment 2.

[0301] In some embodiments, the terminal device may include various personal computers, laptops, smartphones, tablets, Internet of Things devices, and portable wearable devices.

[0302] This application runs the rectangular pattern-based annotation sample generation method described in Embodiment 1 as a program on a terminal device, or runs the rectangular pattern-based image processing method described in Embodiment 2 as a program on a terminal device. The terminal device may include various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices, and can run on different systems and application platforms to perform text annotation in different scenarios, with stronger scalability.

[0303] Example 6

[0304] Accordingly, this application also provides a computer-readable storage medium, the computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the rectangular pattern-based annotation sample generation method or the rectangular pattern-based image processing method as described in any of the above embodiments.

[0305] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete this application. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the terminal device.

[0306] The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.

[0307] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.

[0308] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created based on the use of the mobile terminal, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0309] Wherein, if the modules / units integrated in the terminal device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. Wherein, the computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0310] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application for those skilled in the art.

Claims

1. A method for generating labeled samples based on rectangular patterns, characterized in that, include: In response to user actions on the target image, the coordinates of three scene points on the target image are determined; wherein, the first scene point and the second scene point are located above the text content area in the target image, and the third scene point is used to constrain the bottom edge boundary, or to constrain the bottom edge boundary and correct the left and right edge boundaries. Based on the positional relationship between the first scene point, the second scene point, and the third scene point, determine the first straight line containing the top edge of the rectangular text box and the second straight line containing the bottom edge; Based on the first line and the second line, and in conjunction with a preset rectangular coordinate system, the perpendicular feet of the three scene points to the first line are determined respectively. Based on the positional relationship between the perpendicular feet, the first line, and the second line, the coordinates of the four vertices of the rectangular text box are determined to generate the rectangular text box. Based on the rectangular text box and the target image, generate labeled samples; The determination of the coordinates of three scene points on the target image further includes: If the user inputs the coordinates of more than three scene points, then the coordinates of the first three scene points input by the user will be obtained according to the order of input. Wherein, the first scene point and the second scene point are located on the top edge of the initial rectangle; The initial rectangle is set by the user based on the target text of the target image; The top edge of the initial rectangle is the edge above the target text when the target text is viewed directly. The distance between the first scene point and the vertex at one end of the top edge is not greater than the first distance; The third scene point is located on the bottom edge of the initial rectangle; The position of the third scene point is set according to the distance between the second scene point and the vertex at the other end of the top edge; The bottom edge of the initial rectangle is the edge below the target text when the target text is viewed directly.

2. The method for generating labeled samples based on rectangular patterns as described in claim 1, characterized in that, The position of the third scene point is set based on the distance between the second scene point and the vertex at the other end of the top edge, specifically: If the distance between the second scene point and the vertex at the other end of the top edge is not greater than the first distance, then the setting range of the third scene point is: any point on the target projection line segment located on the bottom edge of the initial rectangle; the target projection line segment is the line segment between the projection points of the first scene point and the second scene point on the bottom edge; If the distance between the second scene point and the vertex at the other end of the top edge is greater than the first distance, then the setting range of the third scene point is: located on the bottom edge of the initial rectangle, and the distance between it and the target vertex is not greater than the second distance; the target vertex and the vertex at one end of the top edge are the diagonal vertices of the initial rectangle.

3. The method for generating labeled samples based on rectangular patterns as described in claim 1, characterized in that, Based on the positional relationship between the first scene point, the second scene point, and the third scene point, the first straight line containing the top edge of the rectangular text box and the second straight line containing the bottom edge are determined, specifically as follows: Connect the first scene point and the second scene point to obtain the first straight line containing the top edge of the rectangular text box; Draw a line parallel to the first line through the third scene point to obtain the second line containing the bottom edge of the rectangular text box.

4. The method for generating labeled samples based on rectangular patterns as described in claim 3, characterized in that, Based on the first line and the second line, and using a preset rectangular coordinate system, the perpendicular feet of the three scene points to the first line are determined respectively, specifically as follows: Draw the first perpendicular line from the first scene point to the first straight line, with the foot of the perpendicular being the first foot of the perpendicular. Draw the second perpendicular line from the second scene point to the first straight line, with the foot of the perpendicular being the second foot of the perpendicular. Draw the third perpendicular line from the third scene point to the first straight line, with the foot of the perpendicular being the third foot of the perpendicular.

5. The method for generating labeled samples based on rectangular patterns as described in claim 4, characterized in that, Based on the positional relationship between the perpendicular feet, the first line, and the second line, the coordinates of the four vertices of the rectangular text box are determined, specifically as follows: Based on the positional relationship between each perpendicular foot and the first straight line, determine the first vertex and the second vertex on the top edge of the rectangular text box; Based on the positional relationship between the first vertex, the second vertex, and the second straight line, determine the third and fourth vertices on the bottom edge of the rectangular text box.

6. The method for generating labeled samples based on rectangular patterns as described in claim 5, characterized in that, Based on the positional relationship between each perpendicular foot and the first straight line, the first vertex and the second vertex on the top edge of the rectangular text box are determined, specifically as follows: If the first line is parallel to the vertical axis of the preset rectangular coordinate system, then the vertical coordinate values ​​of the first foot, the second foot, and the third foot are compared to determine the first vertex and the second vertex, and the coordinates of the first vertex and the second vertex are obtained. Wherein, the first vertex is the foot of the perpendicular with the largest ordinate value, and the second vertex is the foot of the perpendicular with the smallest ordinate value; Alternatively, the first vertex is the foot of the perpendicular with the smallest ordinate value, and the second vertex is the foot of the perpendicular with the largest ordinate value; If the first straight line is not parallel to the vertical axis of the preset rectangular coordinate system, then the horizontal coordinates of the first perpendicular foot, the second perpendicular foot, and the third perpendicular foot are compared to determine the first vertex and the second vertex, and the coordinates of the first vertex and the second vertex are obtained. Wherein, the first vertex is the foot of the perpendicular with the largest x-coordinate value, and the second vertex is the foot of the perpendicular with the smallest x-coordinate value; Alternatively, the first vertex is the foot of the perpendicular with the smallest x-coordinate value, and the second vertex is the foot of the perpendicular with the largest x-coordinate value.

7. The method for generating labeled samples based on rectangular patterns as described in claim 5, characterized in that, Based on the positional relationship between the first vertex, the second vertex, and the second straight line, the third and fourth vertices on the bottom edge of the rectangular text box are determined, specifically as follows: The perpendicular line from the first vertex to the first line is taken as the third line; The perpendicular line from the second vertex to the first line is taken as the fourth line; The intersection of the second line and the third line is taken as the third vertex, and the coordinates of the third vertex are obtained; The intersection of the second line and the fourth line is taken as the fourth vertex, and the coordinates of the fourth vertex are obtained.

8. The method for generating labeled samples based on rectangular patterns as described in claim 7, characterized in that, After determining the third and fourth vertices on the bottom edge of the rectangular text box, the method further includes: Use the distance between the first vertex and the second vertex as the length of the rectangular text box; Use the distance between the first and third vertices as the width of the rectangular text box.

9. The method for generating labeled samples based on rectangular patterns as described in claim 5, characterized in that, The rectangular text box is generated as follows: Based on the positional relationship of the first vertex, second vertex, third vertex and fourth vertex of the rectangular text box, the first vertex, second vertex, third vertex and fourth vertex are sorted to obtain the first coordinate point, second coordinate point, third coordinate point and fourth coordinate point in sequence; A path is generated by connecting the first coordinate point, the second coordinate point, the third coordinate point, and the fourth coordinate point using a preset component; Instantiate the path to obtain a rectangular text box.

10. The method for generating labeled samples based on rectangular patterns as described in claim 9, characterized in that, After instantiating the path to obtain a rectangular text box, the process also includes: In response to the user's dragging and modification of the rectangular text box, the coordinates of the first coordinate point, the second coordinate point, the third coordinate point, and the fourth coordinate point, as well as the length and width of the rectangular text box, are updated.

11. The method for generating labeled samples based on rectangular patterns as described in claim 9, characterized in that, After instantiating the path to obtain a rectangular text box, the process also includes: Starting from the first coordinate point, the first straight line generates a text editing box of a preset size; The text editing box is instantiated, and the changes in the text content within the text editing box are monitored; the text content within the text editing box is entered by the user. If a change occurs, the latest entered text content is retrieved and displayed in the text editing box; Save the latest input text content as text annotation content.

12. An image processing method based on rectangular patterns, characterized in that, include: Obtain the target image, and obtain the annotation sample corresponding to the target image according to the annotation sample generation method based on rectangular pattern as described in any one of claims 1-11; Based on the labeled samples, a perspective transformation is performed on the text region within the rectangular text box on the target image to obtain an image slice.

13. The image processing method based on rectangular patterns as described in claim 12, characterized in that, Based on the labeled samples, a perspective transformation is performed on the text region within the rectangular text box on the target image to obtain image slices, specifically: Use the height of the rectangular text box as the height of the target rectangle, and the width of the rectangular text box as the width of the target rectangle to generate a list of coordinates for the target rectangle; The coordinate list includes the coordinates of the four vertices of the target rectangle; the coordinates of the vertices of the target rectangle are preset to be the origin coordinates; the order of the four vertices is strongly correlated with the order of the four vertices of the rectangular text box. Based on the coordinate list and the coordinates of the first, second, third, and fourth coordinate points, a perspective transformation matrix is ​​calculated to perform a perspective transformation on the target image, thereby obtaining the text region image within the rectangular text box. The text region image is used as an image slice.

14. The image processing method based on rectangular patterns as described in claim 12, characterized in that, Based on the labeled samples, after performing perspective transformation on the text region within the rectangular text box on the target image to obtain image slices, the process further includes: Get the text annotations corresponding to the image slices; Associate and save text annotations with image slices.

15. A device for generating labeled samples based on rectangular patterns, characterized in that, include: The module includes a scene point data acquisition module, a data processing module, and a labeling module. The scene point data acquisition module is used to respond to the user's operation on the target image and determine the coordinates of three scene points on the target image; wherein, the first scene point and the second scene point of the three scene points are located above the area of ​​text content in the target image, and the third scene point is used to constrain the bottom edge boundary, or to constrain the bottom edge boundary and correct the left and right edge boundaries. The data processing module is used to determine the first straight line where the top edge of the rectangular text box is located and the second straight line where the bottom edge is located, based on the positional relationship between the first scene point, the second scene point and the third scene point. The annotation module is used to determine the perpendicular foot of each of the three scene points to the first line based on the first line and the second line, combined with a preset rectangular coordinate system, and to determine the coordinates of the four vertices of the rectangular text box based on the positional relationship between each perpendicular foot, the first line and the second line, so as to generate the rectangular text box. Based on the rectangular text box and the target image, generate labeled samples; The determination of the coordinates of three scene points on the target image further includes: If the user inputs the coordinates of more than three scene points, then the coordinates of the first three scene points input by the user will be obtained according to the order of input. Wherein, the first scene point and the second scene point are located on the top edge of the initial rectangle; The initial rectangle is set by the user based on the target text of the target image; The top edge of the initial rectangle is the edge above the target text when the target text is viewed directly. The distance between the first scene point and the vertex at one end of the top edge is not greater than the first distance; The third scene point is located on the bottom edge of the initial rectangle; The position of the third scene point is set according to the distance between the second scene point and the vertex at the other end of the top edge; The bottom edge of the initial rectangle is the edge below the target text when the target text is viewed directly.

16. An image processing device based on rectangular patterns, characterized in that, include: Annotated data acquisition module and perspective transformation module; The annotation data acquisition module is used to acquire the target image and, according to the annotation sample generation method based on rectangular pattern as described in any one of claims 1-11, acquire the annotation sample corresponding to the target image; The perspective transformation module is used to perform perspective transformation on the text region within the rectangular text box on the target image based on the labeled sample, so as to obtain image slices.

17. A terminal device, characterized in that, include: At least one processor and memory; The memory is used to store program instructions; The processor is configured to call and execute program instructions stored in the memory, so that the terminal device executes the rectangular pattern-based annotation sample generation method as described in any one of claims 1-11, or executes the rectangular pattern-based image processing method as described in any one of claims 12-14.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program; wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the method for generating labeled samples based on rectangular patterns as described in any one of claims 1-11, or to perform the image processing method based on rectangular patterns as described in any one of claims 12-14.

Citation Information

Patent Citations

  • All-round view system automatic calibration method, automobile, calibration device and storage medium

    CN107993263A

  • Chinese character sign identification method and device

    CN108629238A