A picture conversion method, apparatus, device and storage medium
By recognizing and transforming lines and text in lab reports from the medical or chemical industries, converted images are generated, solving the problem of image distortion and improving viewing convenience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DATAGRAND TECH INC
- Filing Date
- 2022-08-26
- Publication Date
- 2026-06-23
AI Technical Summary
In existing technologies, the image is skewed and distorted due to the non-perpendicularity of the camera's optical axis to the plane being photographed, which affects the visual experience, especially in the processing of laboratory reports in the medical and chemical industries, making them difficult to view.
By obtaining the line list and text information of the original image, a coordinate system is established, lines and text are identified, the first region is selected and expanded into the second region, and perspective transformation is performed to generate the transformed image.
It effectively improved the skew and distortion of the images, making them easier for relevant personnel to view.
Smart Images

Figure CN115393853B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image conversion technology, and in particular to an image conversion method, apparatus, device, and storage medium. Background Technology
[0002] When the optical axis of a camera is not perpendicular to the plane being photographed, the shape of the image will be distorted. The distorted image will affect the human visual experience. For example, companies in the medical or chemical industries have a huge number of test reports to process every day, and the current technology directly displays the original images uploaded by users to the relevant personnel.
[0003] However, the way existing technology directly displays the original image is greatly affected by the user's shooting angle, which can lead to various skews and distortions in the images uploaded by users, making it inconvenient for relevant personnel to view them. Summary of the Invention
[0004] This invention provides an image conversion method, apparatus, device, and storage medium to solve the problem of skewed and distorted original images.
[0005] According to one aspect of the present invention, an image conversion method is provided, comprising:
[0006] Retrieve the line list and original text information of the original image;
[0007] Get the first region of the original image based on the list of lines;
[0008] A second region for the original image is obtained based on the first region and the original text information, wherein the second region is larger than the first region;
[0009] The transformed image is generated by performing perspective transformation on the original image based on the second region.
[0010] Preferably, obtaining the line list and original text information of the original image includes: establishing a coordinate system for the original image; performing line detection on the original image in the coordinate system to obtain a line list, wherein the line list contains information about the original lines; and performing text recognition on the original image in the coordinate system to obtain the original text information.
[0011] Preferably, the process of performing line detection on the original image in a coordinate system to obtain a line list includes: performing line detection on the original image in a coordinate system to obtain the original lines contained in the original image; obtaining the endpoint coordinates and type of the original lines, and using the endpoint coordinates and type as information of the original lines; and obtaining a line list based on the information of the original lines.
[0012] Preferably, the original text information is obtained by performing text recognition on the original image in a coordinate system, including: performing text recognition on the original image in a coordinate system to obtain the text content contained in the original image; determining the text coordinates of the text content in the coordinate system; and using the text content and text coordinates as the original text information.
[0013] Preferably, obtaining the first region of the original image based on the line list includes: filtering target lines from the line list, wherein the target lines are located at a specified position in the original image; and taking the closed region formed by the target lines as the first region.
[0014] Preferably, obtaining a second region for the original image based on the first region and the original text information includes: filtering edge text information from the original text information according to the text coordinates and specified rules, wherein the number of edge text information is four; determining the corresponding lines of each edge text information in the first region; obtaining four construction lines based on the corresponding lines and the edge text information; and taking the closed region formed by the construction lines as the second region.
[0015] Preferably, the process of generating a transformed image by performing perspective transformation on the original image based on the second region includes: obtaining vertex information and border information of the second region; obtaining a first matrix based on the vertex information and a second matrix based on the border information, wherein the border information includes the width and height of the second region; generating a transformation matrix based on the first matrix and the second matrix; and performing coordinate transformation on the original text information in the original image based on the transformation matrix to generate a transformed image.
[0016] According to another aspect of the present invention, an image conversion apparatus is provided, comprising:
[0017] The line list and text information acquisition module is used to acquire the line list and original text information of the original image;
[0018] The first region acquisition module is used to obtain the first region of the original image based on the line list;
[0019] The second region acquisition module is used to acquire a second region for the original image based on the first region and the original text information, wherein the second region is larger than the first region.
[0020] The image conversion generation module is used to perform perspective transformation on the original image based on the second region to generate a converted image.
[0021] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0022] At least one processor; and
[0023] A memory communicatively connected to the at least one processor; wherein,
[0024] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform an image conversion method according to any embodiment of the present invention.
[0025] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement an image conversion method according to any embodiment of the present invention.
[0026] The technical solution of this invention obtains a list of lines and original text information by recognizing the original image, determines a first region of the original image based on the list of lines, and then obtains a second region of the original image based on the first region and the original text information. This can accurately determine the effective region in the original image. Then, perspective transformation is performed on the original image through the second region to generate a transformed image, which can effectively improve the distortion and warping of the image and facilitate viewing by relevant personnel.
[0027] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 This is a flowchart of an image conversion method provided in Embodiment 1 of the present invention;
[0030] Figure 2 This is a flowchart of another image conversion method provided in Embodiment 1 of the present invention;
[0031] Figure 3 This is a schematic diagram of the first region provided according to Embodiment 1 of the present invention;
[0032] Figure 4 This is a schematic diagram of the second region provided in Embodiment 1 of the present invention;
[0033] Figure 5 This is a flowchart of another image conversion method provided in Embodiment 2 of the present invention;
[0034] Figure 6 This is a schematic diagram of the structure of an image conversion device according to Embodiment 3 of the present invention;
[0035] Figure 7 This is a schematic diagram of the structure of an electronic device that implements an image conversion method according to an embodiment of the present invention. Detailed Implementation
[0036] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0037] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0038] Example 1
[0039] Figure 1 This document provides a flowchart of an image conversion method according to Embodiment 1 of the present invention. This embodiment is applicable to processing laboratory reports in the medical or chemical industries. The method can be executed by an image conversion device, which can be implemented in hardware and / or software and can be configured in a computer. Figure 1 As shown, the method includes:
[0040] S110. Obtain the line list and original text information of the original image.
[0041] The original image refers to an image that has not been processed by the controller. The original image can be a lab report or other data from the medical or chemical industries, input by the user. Input can be via a terminal device connected to the controller, including but not limited to mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), tablets, and PMPs (Portable Multimedia Players). The line list is an information list containing the original lines from the original image, generated by the controller after extracting the information from the original lines. The original text information refers to the text content and position within the original image. When a user inputs an original image via a terminal device, the controller can process the acquired image to obtain the line list and the original text information.
[0042] Figure 2 A flowchart of an image conversion method is provided for Embodiment 1 of the present invention. Step S110 mainly includes the following steps S111 to S113:
[0043] S111. Establish a coordinate system for the original image.
[0044] Specifically, when a user inputs an original image through a terminal device, the controller can acquire the original image and establish a coordinate system for it. This coordinate system is a two-dimensional Cartesian coordinate system, including an x-axis, a y-axis, and an origin. Establishing the coordinate system allows the controller to represent the specific location of the content in the original image using coordinates, facilitating the positioning of the content within the image. For example, the top-left corner of the original image is used as the origin, the positive x-axis extends to the right from the origin, and the positive y-axis extends downwards from the origin. Intervals of 0.2 cm are used, where each 0.2 cm represents the number 1. This embodiment only uses the top-left corner of the original image as the origin for illustration and does not limit the specific method of establishing the coordinate system.
[0045] S112. Perform line detection on the original image in the coordinate system to obtain a list of lines.
[0046] The line list contains information about the original lines.
[0047] Preferably, the process of performing line detection on the original image in a coordinate system to obtain a line list includes: performing line detection on the original image in a coordinate system to obtain the original lines contained in the original image; obtaining the endpoint coordinates and type of the original lines, and using the endpoint coordinates and type as information of the original lines; and obtaining a line list based on the information of the original lines.
[0048] Specifically, after establishing a coordinate system for the original image, the controller performs deep learning based on a large amount of data input by the R&D personnel and establishes a line detection model. The line detection model contains a line recognition method corresponding to the image, so the controller can perform line detection on the original image in the coordinate system. When the controller receives the original image input by the user, it can perform line detection through the line detection model to identify the endpoint coordinates and type of the original lines. The controller will use the endpoint coordinates and type as the information of the original lines and generate a number for each original line to create a line list.
[0049] For example, when a user inputs original image A through a terminal device, the controller can obtain the original image and establish a coordinate system for it. After establishing the coordinate system, the controller will use a line detection model to detect lines in original image A, detecting original line 1, original line 2, original line 3, and original line 4. After identifying the original lines, the controller will also obtain information about each line, namely the endpoint coordinates and type. The controller can obtain the endpoint coordinates of original line 1 as (2,3) and (8,3) and its type as horizontal line; the endpoint coordinates of original line 2 as (3,4) and (7,4) and its type as horizontal line; the endpoint coordinates of original line 3 as (2,6) and (8,6) and its type as horizontal line; and the endpoint coordinates of original line 4 as (5,3) and (5,6) and its type as vertical line. After obtaining information about all the original lines in original image A, the controller will number them to generate a line list. Taking original image A as an example, Table 1 below shows a schematic of the line list:
[0050] Table 1
[0051]
[0052]
[0053] The line list includes line number, endpoint coordinates and type. Taking the original line 1 as an example, as shown in Table 1, the line number corresponding to the original line 1 is 01, the endpoint coordinates are (2,3) and (8,3), and the type is horizontal line.
[0054] S113. Perform text recognition on the original image in the coordinate system to obtain the original text information.
[0055] Specifically, text recognition can be Optical Character Recognition (OCR). The controller can use OCR to recognize the text information in the original image and obtain the recognized original text information, which includes the coordinates and content of the original text.
[0056] Preferably, the original text information is obtained by performing text recognition on the original image in a coordinate system, including: performing text recognition on the original image in a coordinate system to obtain the text content contained in the original image; determining the text coordinates of the text content in the coordinate system; and using the text content and text coordinates as the original text information.
[0057] Specifically, the controller performs OCR recognition on the original image in a coordinate system. After recognition, it can obtain the text content contained in the original image and determine the text coordinates corresponding to the text content in the coordinate system. The controller can use the obtained text content and text coordinates as the original text information of the original image.
[0058] S120. Obtain the first region of the original image based on the line list.
[0059] Preferably, target lines are selected from the line list, wherein the target lines are located at a specified position in the original image; the closed area formed by the target lines is taken as the first area.
[0060] Specifically, as shown in Table 1, the line list of the original image A contains multiple original lines. Therefore, the controller needs to filter the target lines from the line list according to the filtering conditions, and then take the closed area formed by the target lines as the first area. The filtering conditions are set in advance by the developers within the controller. For example, the developers set the filtering conditions for the first target line to be in the upper half of the original image, with a length greater than 40% of the width of the original image, and the line type to be a horizontal line. When the y-coordinate value of a certain original line is less than half of the height of the original image, it can be determined that the original line is in the upper half of the original image. The length of the first target line = the x-coordinate of the right endpoint - the x-coordinate of the left endpoint. When there are multiple target lines that meet the preset conditions, the controller will perform a secondary filtering, selecting only the original line with the smallest y-coordinate value as the first target line, that is, the first target line is located at the top of the original image.
[0061] For example, when the height y-coordinate of the original image is 10 and the width y-coordinate is 10, the original line 1 corresponding to line number 01 and the original line 2 corresponding to line number 02 that meet the preset conditions can be filtered from Table 1 using the above filtering conditions. Then the controller will perform a second filtering, selecting only the original line with the smallest y-coordinate value as the first target line. As can be seen from Table 1, the y-coordinate of line number 01 is... 01 The y-coordinate of line number 02 is 3. 02 The coordinate value is 4, that is, y 01 <y 02 Therefore, the controller will use the original line 1 with line number 01 as the first target line. For example, the researchers set the following criteria for the second target line: it must be in the lower half of the original image, its length must be greater than 40% of the original image width, and its line type must be horizontal. When the y-coordinate value of an original line is greater than half the height of the original image, it can be determined that the original line is in the lower half of the original image. The length of the second target line = the x-coordinate of the right endpoint - the x-coordinate of the left endpoint. When multiple target lines meet the preset conditions, the controller will perform a secondary screening, selecting only the original line with the largest y-coordinate value as the second target line. That is, the second target line is located at the bottom of the original image. The second target line that meets the preset conditions can be selected from Table 1 using the above screening criteria as the original line 3 corresponding to line number 03.
[0062] Furthermore, after the controller determines the first target line and the second target line based on the screening conditions set by the R&D personnel, the closed area formed by connecting the endpoints of the two target lines can be used as the first area. Figure 3 As shown, Figure 3 The closed area in the diagram represents the first region. Figure 3 In the diagram, O is the origin of the coordinate system, x and y are the coordinate axes of the original image, L1 is the first target line, L2 is the second target line, and the coordinates of the four vertices of the first region are p1(2,3), p2(8,3), p3(2,6) and p4(8,6).
[0063] S130. Obtain the second region for the original image based on the first region and the original text information.
[0064] Specifically, the controller can expand the first region of the original image to obtain a second region based on the first region and the original text information, wherein the second region is larger than the first region.
[0065] Preferably, edge text information is filtered from the original text information according to the text coordinates and specified rules, wherein the number of edge text information is four; the corresponding lines of each edge text information in the first region are determined; four construction lines are obtained according to the corresponding lines and edge text information; and the closed region formed by the construction lines is taken as the second region.
[0066] Specifically, the controller will filter out edge text information from the original text information according to the text coordinates in the original text information according to the specified rules. The specified rules are set in advance by the R&D personnel in the controller, that is, four coordinate points in the original image with the maximum x value, minimum x value, maximum y value and minimum y value, and these four coordinate points are used as edge text information. That is, the edge text information is the four original text information located at the top, bottom, left and right ends of the original image. When there are multiple coordinate points under a certain specified rule, the controller will select any one of the coordinate points as edge text information. After determining the edge text information, the controller can determine the corresponding line of each edge text information in the first area. The corresponding line refers to the first area border closest to each edge text information, that is, one of the four borders in the first area. Each border has a corresponding slope. According to the corresponding line and the edge text information, four construction lines can be obtained. The construction lines are lines that pass through the edge text information and have the same slope as the corresponding line of the edge text. That is, the construction lines can be calculated according to the following formula (1):
[0067]
[0068] Where y represents the vertical coordinate of the constructed line, and x represents the horizontal coordinate of the constructed line. p1 The y-coordinate represents the first endpoint of the corresponding line. p2 The x-coordinate represents the ordinate of the second endpoint of the corresponding line. p1 The x-coordinate represents the x-coordinate of the first endpoint of the corresponding line. p2 The x-coordinate represents the x-coordinate of the first endpoint of the corresponding line. z The x-coordinate of the edge text information, y z The vertical coordinate represents the edge text information. Furthermore, after obtaining four construction lines through the above formula (1), the controller can use the closed area formed by the four construction lines as the second area.
[0069] For example, the controller can obtain four edge text information that meet the specified rules from the original image A, namely Z1(3,2), Z2(3,7), Z3(2,5) and Z4(8,4); and the coordinates of the four vertices of the first region are known to be (2,3), (8,3), (2,6) and (8,6), so the coordinates of the upper border endpoint of the first region corresponding to point Z1 are (2,3) and (8,3). Through the above formula (1), the construction line parallel to the upper border of point Z1 can be calculated as: y = 2. Similarly, the other three construction lines can be calculated as: y = 7, x = 2 and x = 8. After obtaining the four construction lines, the controller can use the closed region formed by the four construction lines as the second region. Figure 4 This is a schematic diagram of the second region. Figure 4 The dashed area in the diagram represents the second region. The coordinates of the four vertices of the second region are j1(2,2), j2(8,2), j3(2,7) and j4(8,7). Z1, Z2, Z3 and Z4 represent four edge text information.
[0070] S140. Perform perspective transformation on the original image based on the second region to generate a transformed image.
[0071] Specifically, after the controller determines the second region, it can perform perspective transformation on the original image based on the second region. That is, it performs perspective transformation on each pixel in the second region to generate a transformed image. A pixel is the smallest unit in an image. The transformed image generated by the perspective transformation contains all the original text information of the original image.
[0072] The technical solution of this invention obtains a list of lines and original text information by recognizing the original image, determines a first region of the original image based on the list of lines, and then obtains a second region of the original image based on the first region and the original text information. This can accurately determine the effective region in the original image. Then, perspective transformation is performed on the original image through the second region to generate a transformed image, which can effectively improve the distortion and warping of the image and facilitate viewing by relevant personnel.
[0073] Example 2
[0074] Figure 5 This is a flowchart of an image conversion method provided in Embodiment 2 of the present invention. Based on Embodiment 1, this embodiment specifically describes how to generate a converted image by performing perspective transformation on the original image according to the second region. Figure 5 As shown, the main steps include the following:
[0075] S210. Obtain the vertex information and border information of the second region.
[0076] Specifically, after determining the second region of the original image, the controller needs to further determine the vertex information and border information of the second region. The vertex information refers to the coordinates of the four vertices of the second region, and the border information refers to the height and width of the second region.
[0077] S220. Obtain the first matrix based on the vertex information and the second matrix based on the bounding box information.
[0078] Specifically, the controller can substitute the information j of the four vertices of the second region into matrix A to obtain the first matrix. Substitute the height w and width h of the second region into the rotation matrix model to obtain the second matrix.
[0079] S230. Generate the transformation matrix based on the first matrix and the second matrix.
[0080] Specifically, the controller can generate a transformation matrix M based on the first matrix A and the second matrix B. The specific process is as follows: first, homogenize the first matrix A and the second matrix B to obtain... and Where z is an arbitrary variable that can be eliminated in the subsequent solution process; then the transformation matrix M is obtained according to the following formula (2):
[0081] A * ·M=B * (2)
[0082] Specifically, the transformation matrix
[0083] S240. Based on the transformation matrix, perform coordinate transformation on the original text information in the original image to generate a transformed image.
[0084] Specifically, after obtaining the transformation matrix M, the controller can perform coordinate transformation on the original text information contained in the original image using the following formula (3) to generate a transformed image:
[0085]
[0086] Wherein, dst(x,y) represents the transformed coordinates, x represents the horizontal coordinate of the original text information, and y represents the vertical coordinate of the original text information. The above formula (3) can be used to transform the coordinates of the pixels containing the original text information to generate a transformed image. Since perspective transformation is an existing technical means of image processing, the specific generation process of the transformed image will not be described in detail in this embodiment.
[0087] The technical solution of this invention obtains a list of lines and original text information by recognizing the original image, determines a first region of the original image based on the line list, and then obtains a second region of the original image based on the first region and the original text information. This can accurately determine the effective region in the original image. Then, a first matrix and a second matrix are determined by the vertex information and border information of the second region. Finally, a transformation matrix is obtained to perform perspective transformation on the original image to generate a transformed image. This method is highly adaptable and can effectively improve the skewness and distortion of the image, making it easier for relevant personnel to view.
[0088] Example 3
[0089] Figure 6 This is a schematic diagram of the structure of an image conversion device provided in Embodiment 3 of the present invention. Figure 6 As shown, the device includes: a line list and text information acquisition module 310, used to acquire a line list and original text information of the original image; a first region acquisition module 320, used to acquire a first region of the original image based on the line list; a second region acquisition module 330, used to acquire a second region of the original image based on the first region and the original text information, wherein the second region is larger than the first region; and a converted image generation module 340, used to perform perspective transformation on the original image based on the second region to generate a converted image.
[0090] Preferably, the line list and text information acquisition module 310 specifically includes: a coordinate system establishment unit, used to establish a coordinate system for the original image; a line list acquisition unit, used to perform line detection on the original image under the coordinate system to acquire a line list, wherein the line list contains information about the original lines; and an original text information acquisition unit, used to perform text recognition on the original image under the coordinate system to acquire original text information.
[0091] Preferably, the line list acquisition unit is specifically used for: performing line detection on the original image in a coordinate system to obtain the original lines contained in the original image; obtaining the endpoint coordinates and type of the original lines, and using the endpoint coordinates and type as the information of the original lines; and obtaining a line list based on the information of the original lines.
[0092] Preferably, the original text information acquisition unit is specifically used for: performing text recognition on the original image in a coordinate system to obtain the text content contained in the original image; determining the text coordinates of the text content in the coordinate system; and using the text content and text coordinates as original text information.
[0093] Preferably, the first region acquisition module 320 is specifically used for: filtering target lines from the line list, wherein the target lines are located at a specified position in the original image; and taking the closed region formed by the target lines as the first region.
[0094] Preferably, the second region acquisition module 330 is specifically used for: filtering edge text information from the original text information according to the text coordinates and specified rules, wherein the number of edge text information is four; determining the corresponding lines of each edge text information in the first region; acquiring four construction lines according to the corresponding lines and edge text information; and taking the closed region formed by the construction lines as the second region.
[0095] Preferably, the image conversion generation module 340 is specifically used for: obtaining vertex information and border information of the second region; obtaining a first matrix based on the vertex information and a second matrix based on the border information, wherein the border information includes the width and height of the second region; generating a transformation matrix based on the first matrix and the second matrix; and performing coordinate transformation on the original text information in the original image based on the transformation matrix to generate a converted image.
[0096] The technical solution of this invention obtains a list of lines and original text information by recognizing the original image, determines a first region of the original image based on the list of lines, and then obtains a second region of the original image based on the first region and the original text information. This can accurately determine the effective region in the original image. Then, perspective transformation is performed on the original image through the second region to generate a transformed image, which can effectively improve the distortion and warping of the image and facilitate viewing by relevant personnel.
[0097] The image conversion device provided in this embodiment of the invention can execute an image conversion method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0098] Example 4
[0099] Figure 7 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0100] like Figure 7As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0101] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0102] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as an image conversion method.
[0103] In some embodiments, an image conversion method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the image conversion method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform an image conversion method by any other suitable means (e.g., by means of firmware).
[0104] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0105] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0106] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0107] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0108] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0109] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0110] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0111] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. An image conversion method, characterized in that, include: Retrieve the line list and original text information of the original image; Obtain the first region of the original image based on the line list; A second region for the original image is obtained based on the first region and the original text information, wherein the second region is larger than the first region; A transformed image is generated by performing a perspective transformation on the original image based on the second region; The step of obtaining the line list and original text information of the original image includes: Establish a coordinate system for the original image; Line detection is performed on the original image in the coordinate system to obtain the line list, wherein the line list contains information about the original lines; The original image is subjected to text recognition in the coordinate system to obtain the original text information; The step of performing text recognition on the original image in the coordinate system to obtain the original text information includes: The original image is subjected to text recognition in the coordinate system to obtain the text content contained in the original image; Determine the text coordinates of the text content in the coordinate system; The text content and the text coordinates are used as the original text information; The step of obtaining a second region for the original image based on the first region and the original text information includes: Based on the text coordinates, edge text information is filtered from the original text information according to a specified rule, wherein the number of edge text information is four; Determine the corresponding lines of each of the aforementioned edge text information in the first region; Four construction lines are obtained based on the corresponding lines and the edge text information; The closed area formed by the constructed lines is designated as the second region.
2. The method according to claim 1, characterized in that, The step of performing line detection on the original image in the coordinate system to obtain the line list includes: Line detection is performed on the original image in the coordinate system to obtain the original lines contained in the original image; Obtain the endpoint coordinates and type of the original line, and use the endpoint coordinates and type as information of the original line; The line list is obtained based on the information of the original lines.
3. The method according to claim 2, characterized in that, The step of obtaining the first region for the original image based on the line list includes: Select a target line from the list of lines, wherein the target line is located at a specified position in the original image; The closed area formed by the target lines is taken as the first area.
4. The method according to claim 1, characterized in that, The step of performing perspective transformation on the original image based on the second region to generate a transformed image includes: Obtain the vertex and bounding box information of the second region; A first matrix is obtained based on the vertex information, and a second matrix is obtained based on the border information, wherein the border information includes the width and height of the second region; Generate a transformation matrix based on the first matrix and the second matrix; The transformed image is generated by performing coordinate transformation on the original text information in the original image based on the transformation matrix.
5. An image conversion device, characterized in that, include: The line list and text information acquisition module is used to acquire the line list and original text information of the original image; The first region acquisition module is used to acquire a first region for the original image based on the line list; The second region acquisition module is used to acquire a second region for the original image based on the first region and the original text information, wherein the second region is larger than the first region; The image conversion generation module is used to perform perspective transformation on the original image based on the second region to generate a converted image; The line list and text information acquisition module specifically includes: a coordinate system establishment unit, used to establish a coordinate system for the original image; a line list acquisition unit, used to perform line detection on the original image under the coordinate system to acquire the line list, wherein the line list contains information about the original lines; and an original text information acquisition unit, used to perform text recognition on the original image under the coordinate system to acquire the original text information. Specifically, the original text information acquisition unit is used to: perform text recognition on the original image in the coordinate system to obtain the text content contained in the original image; determine the text coordinates of the text content in the coordinate system; and use the text content and the text coordinates as the original text information. The second region acquisition module is specifically used for: filtering edge text information from the original text information according to the text coordinates and a specified rule, wherein the number of edge text information is four; determining the corresponding lines of each edge text information in the first region; acquiring four construction lines according to the corresponding lines and the edge text information; and taking the closed region formed by the construction lines as the second region.
6. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-4.
7. A computer storage medium, characterized in that, The computer storage medium stores computer instructions that are used to cause a processor to execute the method of any one of claims 1-4.
Citation Information
Patent Citations
Curved surface image correction method, device and electronic equipment
CN113012029A
Image document correction method and system, terminal and medium
CN113808033A