Image processing method and device, storage medium and electronic equipment
By performing image distortion correction, character recognition, and layout reconstruction on the text image to be translated, the problem of visual mismatch between the translation and the original text was solved, and the visual display effect of the translation in the original image was improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-01
- Publication Date
- 2026-03-24
AI Technical Summary
In existing photo translation technologies, the translated text is difficult to visually match the original text, resulting in poor visual display effects.
By performing image distortion correction on the text image to be translated, recognizing characters and text layout, reconstructing the text layout, translating the text and restoring the background, the translated text is presented naturally in the original image.
It achieves a match between the layout of the translated text and the original image, thus improving the visual display effect.
Smart Images

Figure CN121725486A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an image processing method, apparatus, storage medium, and electronic device. Background Technology
[0002] The development of photo translation technology benefits from continuous breakthroughs in fields such as computer vision, natural language processing (NLP), and deep learning. With the widespread use of smartphones, tablets, and other terminal devices, users can capture images anytime, anywhere and translate them instantly, greatly facilitating cross-language communication and the global flow of information. Traditional translation methods typically rely on manual input and sentence-by-sentence translation, while photo translation, through its automated process, not only improves translation efficiency but also reduces the user's operational burden.
[0003] In photo translation technology, text in the original image is detected, recognized, and translated, and the translated text is then written back to the original image. However, most related photo translation technologies focus on text translation, making it difficult for the translated text to visually integrate into the original image. This results in a mismatch between the translated text's layout and the original text, leading to poor visual display of the translated text. Summary of the Invention
[0004] This application provides an image processing method, apparatus, computer storage medium, and electronic device. The technical solution is as follows: In a first aspect, embodiments of this application provide an image processing method, the method comprising: Obtain the text image to be translated, and perform image distortion correction processing on the text image to be translated to obtain a distortion-corrected text image; The distortion-corrected text image is subjected to character recognition processing to obtain character recognition information, and the distortion-corrected text image is subjected to text layout recognition processing to obtain text layout structure information. Based on the character recognition information and the text layout structure information, text layout reconstruction processing is performed to obtain text layout reconstruction information. Based on the text layout reconstruction information, text translation processing is performed to obtain the text translation result, and based on the text layout reconstruction information, text background repair processing is performed to obtain the translated text rendering image; Based on the text layout reconstruction information, the translated text rendering image, and the text translation result, the target translated text image is obtained through translated text backfilling and image distortion restoration.
[0005] In some possible implementations, the step of performing image distortion correction processing on the text image to be translated to obtain a distortion-corrected text image includes: The text image to be translated is segmented to obtain a text region image, and a perspective transformation is performed on the text region image to obtain an initial distortion-corrected text image. The initial distortion-corrected text image is subjected to image dedistortion processing to obtain the distortion-corrected text image.
[0006] In some possible implementations, the text layout reconstruction process based on the character recognition information and the text layout structure information to obtain text layout reconstruction information includes: Based on the text layout structure information, each candidate text layout box is determined, and based on the character recognition information, each character line box is determined. Determine the overlap ratio between the character line box and the candidate text layout box, and perform text layout reconstruction processing based on the overlap ratio and preset layout structure segmentation rules to obtain text layout reconstruction information.
[0007] In some possible implementations, the text layout reconstruction process based on the overlapping area ratio and preset page layout segmentation rules to obtain text layout reconstruction information includes: Obtain a preset area ratio threshold, and perform matching processing on the character line box and the candidate text layout box based on the overlapping area ratio and the preset area ratio threshold to obtain the matching character line box and the unmatched character line box corresponding to the candidate text layout box; Based on the preset layout structure segmentation rules, the matching character line boxes are segmented within the layout box to obtain reference layout box structure information. Based on the reference layout box structure information, the target layout box structure information corresponding to the candidate text layout box is determined. Based on the target layout box structure information and the unmatched character line boxes, text layout reconstruction information is obtained.
[0008] In some possible implementations, the text background restoration process based on the text layout reconstruction information to obtain the translated rendered image includes: Based on the text layout reconstruction information, a reference mask image is determined, and the reference mask image is scaled to obtain a target mask image. An initial distortion-corrected text image is obtained, and the initial distortion-corrected text image is scaled to obtain a target text image. A target mask region is determined in the target mask image, a target mask mapping region corresponding to the target mask region is determined in the target text image, and background restoration processing is performed on the target mask mapping region in the target text image to obtain a restored text image. The repaired text image is subjected to image size restoration processing to obtain a candidate text image. A reference mask region is determined in the reference mask image. The pixel information of the first reference mask mapping region corresponding to the reference mask region is determined in the candidate text image. The pixel information of the second reference mask mapping region corresponding to the reference mask region is determined in the initial distortion-corrected text image. In the initial distortion-corrected text image, the pixel information of the second reference mask mapping region is replaced with the pixel information of the first reference mask mapping region to obtain the translated text rendering image.
[0009] In some possible implementations, the step of performing translated text backfilling and image distortion restoration processing based on the text layout reconstruction information, the translated text rendering image, and the text translation result to obtain the target translated text image includes: Determine the image distortion and restoration information corresponding to the distortion-corrected text image, and perform distortion restoration processing on the translated rendered image based on the image distortion and restoration information to obtain the distortion-restored text image; Based on the text layout reconstruction information, character style information is determined, and based on the character style information and the text translation result information, the distorted restored text image is processed to obtain the initial translated text image; The initial translated text image is subjected to inverse perspective transformation and segmentation to obtain the target translated text image.
[0010] In some possible implementations, determining the image distortion restoration information corresponding to the distortion-corrected text image includes: Obtain at least two rounds of image flattening deformation information corresponding to the distortion-corrected text image; Determine the first round of image flattening deformation information from the at least two rounds of image flattening vector information, and determine the image distortion and restoration deformation information corresponding to the first round of image flattening deformation information.
[0011] Secondly, embodiments of this application provide an image processing apparatus, the apparatus comprising: An image distortion correction module is used to acquire an image of the text to be translated and to perform image distortion correction processing on the image of the text to be translated to obtain a distortion-corrected text image. The image layout recognition module is used to perform character recognition processing on the distortion-corrected text image to obtain character recognition information, perform text layout recognition processing on the distortion-corrected text image to obtain text layout structure information, and perform text layout reconstruction processing based on the character recognition information and the text layout structure information to obtain text layout reconstruction information. The image-text translation module is used to perform text translation processing based on the text layout reconstruction information to obtain the text translation result, and to perform text background repair processing based on the text layout reconstruction information to obtain the translated text rendering image; The translation text synthesis module is used to perform translation text backfilling and image distortion restoration processing based on the text layout reconstruction information, the translated text rendering image, and the text translation result to obtain the target translated text image.
[0012] Optional, an image distortion correction module, specifically used for: The text image to be translated is segmented to obtain a text region image, and a perspective transformation is performed on the text region image to obtain an initial distortion-corrected text image. The initial distortion-corrected text image is subjected to image dedistortion processing to obtain the distortion-corrected text image.
[0013] Optional, the image layout recognition module includes: The layout detection unit is used to determine each candidate text layout box based on the text layout structure information and to determine each character line box based on the character recognition information. The layout reconstruction unit is used to determine the overlap area ratio between the character line box and the candidate text layout box, and to perform text layout reconstruction processing based on the overlap area ratio and the preset layout structure segmentation rules to obtain text layout reconstruction information.
[0014] Optional, layout reconfiguration unit, used for: Obtain a preset area ratio threshold, and perform matching processing on the character line box and the candidate text layout box based on the overlapping area ratio and the preset area ratio threshold to obtain the matching character line box and the unmatched character line box corresponding to the candidate text layout box; Based on the preset layout structure segmentation rules, the matching character line boxes are segmented within the layout box to obtain reference layout box structure information. Based on the reference layout box structure information, the target layout box structure information corresponding to the candidate text layout box is determined. Based on the target layout box structure information and the unmatched character line boxes, text layout reconstruction information is obtained.
[0015] Optional, image-to-text translation module, specifically used for: Based on the text layout reconstruction information, a reference mask image is determined, and the reference mask image is scaled to obtain a target mask image. An initial distortion-corrected text image is obtained, and the initial distortion-corrected text image is scaled to obtain a target text image. A target mask region is determined in the target mask image, a target mask mapping region corresponding to the target mask region is determined in the target text image, and background restoration processing is performed on the target mask mapping region in the target text image to obtain a restored text image. The repaired text image is subjected to image size restoration processing to obtain a candidate text image. A reference mask region is determined in the reference mask image. The pixel information of the first reference mask mapping region corresponding to the reference mask region is determined in the candidate text image. The pixel information of the second reference mask mapping region corresponding to the reference mask region is determined in the initial distortion-corrected text image. In the initial distortion-corrected text image, the pixel information of the second reference mask mapping region is replaced with the pixel information of the first reference mask mapping region to obtain the translated text rendering image.
[0016] Optional, the translation text synthesis module includes: The first image restoration unit is used to determine the image distortion restoration deformation information corresponding to the distortion-corrected text image, and to perform distortion restoration processing on the translated rendered image based on the image distortion restoration deformation information to obtain the distorted text image. The translation write-back unit is used to determine character style information based on the text layout reconstruction information, and to perform translation write-back processing on the distorted restored text image based on the character style information and the text translation result information to obtain an initial translated text image. The second image restoration unit is used to perform perspective transformation inverse processing and segmentation restoration on the initial translated text image to obtain the target translated text image.
[0017] Optionally, the first image restoration unit is specifically used for: Obtain at least two rounds of image flattening deformation information corresponding to the distortion-corrected text image; Determine the first round of image flattening deformation information from the at least two rounds of image flattening vector information, and determine the image distortion and restoration deformation information corresponding to the first round of image flattening deformation information.
[0018] Thirdly, embodiments of this application provide a computer storage medium having multiple instructions adapted for loading and executing the methods described above by a processor.
[0019] Fourthly, embodiments of this application provide an electronic device, which may include: a memory and a processor; wherein the memory stores a computer program adapted to be loaded by the memory and to execute the above-described method.
[0020] The beneficial effects of the technical solutions provided in this application include at least the following: The image processing method provided in this application obtains a text image to be translated. First, it performs image distortion correction processing on the text image to obtain a distortion-corrected text image, thereby eliminating the distortion problem of the text image to be translated. Then, it performs character recognition processing on the distortion-corrected text image to obtain character recognition information. It then performs text layout recognition processing on the distortion-corrected text image to obtain text layout structure information. Based on the character recognition information and text layout structure information, it performs text layout reconstruction processing to obtain text layout reconstruction information, thereby accurately identifying the layout in the text image. Finally, it performs text translation processing based on the text layout reconstruction information to obtain the text translation result. Based on the text layout reconstruction information, the text background restoration processing, it obtains a translated text rendering image. Thus, based on the text layout reconstruction information, the translated text rendering image, and the text translation result, it performs translated text backfilling processing and image distortion restoration processing to obtain the target translated text image. This realizes the writing of the text translation result back to the corresponding layout of the image, so that the translated text matches the layout of the original image, and also ensures that the translated text is presented naturally in the original distortion, improving the visual display effect of the translated text in the original image. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a scene diagram of an image processing system provided in an embodiment of this application; Figure 2 This is a schematic flowchart of an image processing method provided in an embodiment of this application; Figure 3 This is a flowchart illustrating another image processing method provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an image processing device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0023] To make the inventive objectives, features, and advantages of the embodiments of this application more apparent and understandable, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0024] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this application, it should be noted that, unless otherwise expressly specified and limited, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0025] The present application will now be described in detail with reference to specific embodiments.
[0026] like Figure 1 The image shown is a scene diagram of an image processing system provided in an embodiment of this application. Figure 1 As shown, the image processing system may include at least a client cluster and a service platform 100.
[0027] In some embodiments, the client cluster may include at least one client, such as Figure 1 As shown, it specifically includes client 1 corresponding to user 1, client 2 corresponding to user 2, ..., client n corresponding to user n, where n is an integer greater than 0.
[0028] Each client in a client cluster can be an electronic device with communication capabilities, including but not limited to: wearable devices, handheld devices, personal computers, tablets, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem. Electronic devices may have different names in different networks, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), and electronic devices in 5G networks or future evolved networks.
[0029] The service platform 100 can be a standalone server device, such as a rack-mounted, blade, tower, or cabinet-type server device, or a workstation, mainframe, or other hardware device with strong computing power; or it can be a server cluster composed of multiple servers. The servers in the service cluster can be composed in a symmetrical manner, wherein each server is functionally and hierarchically equivalent in the transaction chain, and each server can provide services to the outside world independently. The independent provision of services can be understood as not requiring the assistance of other servers.
[0030] In one or more embodiments of this application, the service platform 100 can establish a communication connection with at least one client in the client cluster, and complete data interaction during the image processing process based on the communication connection. For example, the client has a target application with a photo translation function installed. The client captures an image of text to be translated and uploads the image to the service platform through the target application. The service platform obtains the image of text to be translated, performs image distortion correction processing on the image to obtain a distortion-corrected text image, performs character recognition processing on the distortion-corrected text image to obtain character recognition information, performs text layout recognition processing on the distortion-corrected text image to obtain text layout structure information, performs text layout reconstruction processing based on the character recognition information and the text layout structure information to obtain text layout reconstruction information, performs text translation processing based on the text layout reconstruction information to obtain a text translation result, performs text background restoration processing based on the text layout reconstruction information to obtain a translated rendered image, and performs translated text backfilling processing and image distortion restoration processing based on the text layout reconstruction information, the translated rendered image, and the text translation result to obtain the target translated text image.
[0031] It should be noted that the service platform 100 establishes a communication connection with at least one client in the client cluster via a network for interactive communication. This network can be a wireless network or a wired network. Wireless networks include, but are not limited to, cellular networks, wireless LANs, infrared networks, or Bluetooth networks. Wired networks include, but are not limited to, Ethernet, universal serial bus (USB), or controller area networks. In one or more embodiments of the specification, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network (such as target compressed packets). Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0032] The image processing system embodiments provided in this specification and the image processing methods described in one or more embodiments belong to the same concept. The execution entity corresponding to the image processing methods involved in one or more embodiments of this specification can be an electronic device, and the electronic device can be the aforementioned service platform 100. The specific implementation process of the image processing system embodiments can be found in the following method embodiments, and will not be repeated here.
[0033] In one embodiment, such as Figure 2 As shown, an image processing method is proposed, which can be implemented using a computer program and can run on an image processing device based on the von Neumann architecture. This computer program can be integrated into applications or run as a standalone utility application.
[0034] Specifically, the image processing method includes: S201, Obtain the text image to be translated, and perform image distortion correction processing on the text image to obtain the distortion-corrected text image.
[0035] The text image to be translated refers to an image containing at least one language (such as Chinese characters, English, French, Japanese, Korean, German, etc.). Optionally, the text image to be translated can be an image obtained by the user through a smart device by taking or scanning paper documents, books, signs, posters, etc., or an image downloaded by the user from the Internet, or an image obtained by screenshotting.
[0036] In one embodiment, the text image to be translated is obtained from a user terminal, or from a preset database.
[0037] It is understandable that the user captures an image of the text to be translated through a user terminal, and the user terminal sends the image of the text to be translated to an electronic device. It is also understandable that the user captures an image of the text to be translated through a user terminal, the user terminal uploads the image of the text to be translated to a preset database, and the electronic device retrieves the image of the text to be translated from the preset database according to a preset period.
[0038] Among them, distortion-corrected text image refers to the text image obtained by pulling the text region containing the text content of the text image to be translated to a normal viewing angle and removing the distortion.
[0039] In one embodiment, the above-mentioned image distortion correction processing may include, but is not limited to, segmenting the text region image containing the text content from the text image to be translated, converting the text region image to the frontal view plane, and removing distortion from the text region image converted to the frontal view plane.
[0040] S202, perform character recognition processing on the distortion-corrected text image to obtain character recognition information, perform text layout recognition processing on the distortion-corrected text image to obtain text layout structure information, and perform text layout reconstruction processing based on character recognition information and text layout structure information to obtain text layout reconstruction information.
[0041] The character recognition information includes the coordinates of each character line box and the character content contained in each character line box. A character line box is a polygonal box used to enclose a single line of characters.
[0042] Understandably, the distortion-corrected text image can be input into a preset character recognition model for recognition to obtain character recognition information.
[0043] The text layout structure information includes the coordinates of each candidate text layout box and the character content contained within each candidate text layout box. A candidate text layout box is a polygonal box used to enclose the characters in the layout structure.
[0044] Understandably, the distortion-corrected text image can be input into a preset text layout recognition model for recognition to obtain text layout structure information.
[0045] The text layout reconstruction information includes the coordinate information of different text boxes obtained through layout reconstruction and the character content they contain.
[0046] In one embodiment, for each candidate text layout box, matching character line boxes that should be included in the internal structure of the candidate text layout box, and unmatched character line boxes that are not included in the internal structure of the candidate text layout box are filtered out. Text layout reconstruction information is obtained by performing text layout reconstruction processing through matching character line boxes and unmatched character line boxes.
[0047] S203, based on the text layout reconstruction information, perform text translation processing to obtain the text translation result, and based on the text layout reconstruction information, perform text background restoration processing to obtain the translated text rendering image.
[0048] The text translation result is the translation of the characters to be translated in the text layout reconstruction information.
[0049] In one embodiment, a large-scale text translation model is invoked to process the text layout reconstruction information to obtain the text translation result. A reference mask image is determined based on the text layout reconstruction information. Subsequently, the initial distortion-corrected text image obtained by converting the text region image to a frontal plane and the reference mask image are scaled to a preset size using a proportional strategy. The mask region in the scaled reference mask image is determined to be the mask mapping region in the scaled initial distortion-corrected text image. Texture and structure reconstruction is performed on the mask mapping region in a content-aware inpainting network to obtain a clean background region without text and with consistent lighting. This results in a repaired text image containing the clean background region. Finally, the repaired text image is scaled back to its original size, and only the clean background region in the repaired text image is used to replace the mask mapping region in the initial distortion-corrected text image pixel by pixel. The region outside the mask mapping region remains unchanged, resulting in the translated text rendering image.
[0050] S204. Based on the text layout reconstruction information, the translated text rendering image, and the text translation results, the translated text backfilling process and image distortion restoration process are performed to obtain the target translated text image.
[0051] In one embodiment, in the process of obtaining the target translated text image, the first step is to perform image distortion and restoration processing on the translated text rendering image to obtain a distorted and restored text image. Then, based on the text layout reconstruction information and the text translation result, the translated text is written back into the distorted and restored text image to obtain the initial translated text image. Finally, the initial translated text image is subjected to perspective transformation inverse processing and segmentation and restoration to obtain the target translated text image.
[0052] In the image processing method provided in this application embodiment, an image of text to be translated is obtained. First, the image of text to be translated is subjected to image distortion removal processing to obtain a distortion-corrected text image, so as to eliminate the distortion problem of the text image to be translated. Then, character recognition processing is performed on the distortion-corrected text image to obtain character recognition information. Text layout recognition processing is performed on the distortion-corrected text image to obtain text layout structure information. Based on the character recognition information and text layout structure information, text layout reconstruction processing is performed to obtain text layout reconstruction information, thereby accurately identifying the layout in the text image. Finally, text translation processing is performed based on the text layout reconstruction information to obtain the text translation result. Text background restoration processing is performed based on the text layout reconstruction information to obtain the translated text rendering image. Thus, based on the text layout reconstruction information, the translated text rendering image, and the text translation result, translated text backfilling processing and image distortion restoration processing are performed to obtain the target translated text image. This realizes that the text translation result is written back to the corresponding layout of the image, so that the translated text matches the layout of the original image, and can also ensure that the translated text is presented naturally in the original distortion, improving the visual display effect of the translated text in the original image.
[0053] Please see Figure 3 This is a flowchart illustrating another embodiment of an image processing method proposed in this application.
[0054] Specifically, the image processing method includes: S301, Obtain the image of the text to be translated.
[0055] For details on how step S301 is implemented, please refer to [link / reference]. Figure 2 The descriptions of the relevant steps in the illustrated embodiments will not be repeated here.
[0056] S302, perform image segmentation processing on the text image to be translated to obtain a text region image, and perform perspective transformation processing on the text region image to obtain an initial distortion-corrected text image.
[0057] Here, a text region image refers to a region image containing text content segmented from a text image to be translated. It can be understood that the text image to be translated may include regions containing text content and regions not containing text content; in this case, the text region image is obtained by segmenting the regions containing text content from the text region image. It can also be understood that the text image to be translated may not contain any regions not containing text content; in this case, the text image to be translated is still a text region image.
[0058] In one embodiment, obtaining a text region image can specifically involve: identifying an initial mask image corresponding to the text image to be translated using a preset segmentation model; extracting text region contour information from the text image to be translated based on the initial mask image; filtering at least one candidate quadrilateral based on the text region contour information; selecting the candidate quadrilateral with the largest area as the target quadrilateral; and segmenting the target quadrilateral from the text image to be translated to obtain the text region image.
[0059] It is understood that the initial mask image mentioned above is an image containing binary masks of the same size as the text image to be translated. These binary masks can be used to characterize each pixel in the text image to be translated as a text pixel or a non-text pixel. The text region contour mentioned above includes a set of text region boundary points, and at least one candidate quadrilateral with four vertices can be selected using a polygon approximation algorithm, such as the Douglas-Peucker algorithm.
[0060] For example, the above-mentioned preset segmentation model can be the YOLO-Seg model.
[0061] The initial distortion-corrected text image is the text image obtained by transforming the text region image to a frontal viewing angle.
[0062] In one embodiment, obtaining an initial distortion-corrected text image can specifically involve: determining each initial vertex in the text region image, and determining the target vertex corresponding to each initial vertex; wherein, the target vertex is the vertex of the initial vertex in the frontal plane after the text region image is perspective-expanded to the frontal plane; determining the homography matrix based on the target vertex and the initial vertex, and obtaining the initial distortion-corrected text image by perspective-expanding the text region image to the frontal plane based on the homography matrix.
[0063] S303, perform image dedistortion processing on the initial distortion-corrected text image to obtain the distortion-corrected text image.
[0064] In one embodiment, during the image dedistortion process, the image flattening deformation information of the initial distortion correction text model is predicted round by round by the distortion deformation correction model, and the cumulative image flattening deformation information is obtained based on at least one round of image flattening deformation information. The initial distortion correction text image is corrected into a distortion correction text image based on the cumulative image flattening deformation information.
[0065] Understandably, the image flattening deformation information in each round includes the displacement vectors of the pixels in the current text image that need to be moved to eliminate image distortion in this round. In the first round, the current text image is the initial distortion-corrected text image; in the second round, the current text image is the image obtained by correcting the distortion in the first round from the initial distortion-corrected text image; and so on. In the last round, the current text image is the image obtained by correcting the distortion in the penultimate round. The cumulative image flattening deformation information can include the cumulative displacement vectors of the pixels that need to be moved from the initial distortion-corrected text image to the final image distortion elimination.
[0066] Understandably, in the process of predicting the image flattening deformation information of the initial distortion-corrected text image through the distortion correction model round by round, the distortion correction model predicts the first image flattening deformation information based on the initial distortion-corrected text image. This first image flattening deformation information is used to eliminate global perspective and bending deformation. The distortion correction model determines the first distortion-corrected text image based on the initial distortion-corrected text image and the first image flattening deformation information. This first distortion-corrected text image is obtained by flattening the initial distortion-corrected text image using the first image flattening deformation information. The distortion correction model then predicts the second image flattening deformation information based on the first distortion-corrected text image. The first image is used to flatten deformation information, and the second image is used to eliminate mid-frequency ripple deformation. A second distortion-corrected text image is determined based on the first distortion-corrected text image and the second image's flattened deformation information using a distortion correction model. The second distortion-corrected text image is obtained by flattening the first distortion-corrected text image using the second image's flattened deformation information. A third image's flattened deformation information is predicted based on the second distortion-corrected text image using the distortion correction model. This third image's flattened deformation information is used to eliminate nonlinear distortion. Finally, the distortion-corrected text image is obtained by flattening the second distortion-corrected text image using the third image's flattened deformation information and the second image.
[0067] It is also understandable that, after obtaining at least one round of image flattening deformation information, the reference image distortion restoration deformation information corresponding to each round of image flattening deformation information can be determined, as can the cumulative image distortion restoration deformation information corresponding to the cumulative image flattening deformation information. The reference image distortion restoration deformation information can include the displacement vector required for each pixel to move when restoring the image from the current distortion-removed image to the image before distortion removal in the current round. The cumulative image distortion restoration deformation information can include the cumulative displacement vector required for each pixel to move when restoring the image from the final distortion-removed image to the initial distortion-corrected text image.
[0068] For example, the above-mentioned torsion deformation correction model can be obtained by constructing an initial torsion deformation correction model and training it based on the Thin Plate Spline (TPS) algorithm.
[0069] S304, Perform character recognition processing on the distortion-corrected text image to obtain character recognition information, and perform text layout recognition processing on the distortion-corrected text image to obtain text layout structure information.
[0070] In one embodiment, character recognition information is obtained by performing character recognition on the distortion-corrected text image using a preset character recognition model. The character recognition information includes the coordinate information of each character line frame and the character content contained in each character line frame. Here, a character line frame is a polygonal frame used to enclose a single line of characters.
[0071] By using a pre-defined text layout recognition model to identify the text layout structure of a distortion-corrected text image, text layout structure information is obtained. This information includes the coordinates of each candidate text layout box and the character content contained within each text layout box. The candidate text layout box is a polygonal box used to enclose the characters in the layout structure. For example, if the distortion-corrected text image includes a layout structure consisting of a title and body text, then the candidate text layout boxes could be polygonal boxes enclosing the title and polygonal boxes enclosing the body text.
[0072] For example, the above-mentioned preset text layout recognition model can adopt the DocLayout-YOLO model.
[0073] S305, determine the layout boxes of each candidate text based on text layout structure information, and determine the line boxes of each character based on character recognition information.
[0074] In one embodiment, candidate text layout boxes are extracted from text layout structure information, and character line boxes are extracted from character recognition information.
[0075] S306, determine the overlap ratio of the character line box and the candidate text layout box, and perform text layout reconstruction processing based on the overlap ratio and the preset layout structure segmentation rules to obtain text layout reconstruction information.
[0076] The overlapping area ratio is the ratio of the overlapping area of the character line box and the candidate text layout box to the area of the character line box.
[0077] In one embodiment, for each candidate text layout box, the percentage of overlap between the candidate text layout box and each character line box is calculated. It is understood that the percentage of overlap can be easily calculated based on the coordinate information of each character line box and the coordinate information of the candidate text layout box.
[0078] Among them, the preset page layout segmentation rules are the rules for splitting individual page structures configured for different page layout structures. For example, the preset page layout segmentation rules may include, but are not limited to, segmentation rules triggered by list numbers, segmentation rules based on multi-column discrimination, segmentation rules based on line length differences, segmentation rules based on vertical paragraphs, and so on.
[0079] In one embodiment, text layout reconstruction processing based on overlapping area ratio and preset page structure segmentation rules is performed to obtain text layout reconstruction information, which may specifically include the following steps: A1: Obtain the preset area ratio threshold, and perform matching processing on the character line box and the candidate text layout box based on the overlapping area ratio and the preset area ratio threshold to obtain the matching character line box and the unmatched character line box corresponding to the candidate text layout box; A2: Based on the preset layout structure segmentation rules, the matching character line boxes are segmented within the layout box to obtain the reference layout box structure information. Based on the reference layout box structure information, the target layout box structure information corresponding to the candidate text layout box is determined. Based on the target layout box structure information and the unmatched character line boxes, the text layout reconstruction information is obtained.
[0080] In step A1, for each candidate text layout box, if the overlap area between the character line box and the candidate text layout box is greater than or equal to a preset area percentage threshold, the character line box is determined to be a matching character line box corresponding to the candidate text layout box; if the overlap area between the character line box and the candidate text layout box is less than the preset area percentage threshold, the character line box is determined to be an unmatched character line box corresponding to the candidate text layout box. Here, a matching character line box indicates that the character line box should be included within the candidate text layout box, and an unmatched character line box indicates that the character line box should not be included within the candidate text layout box.
[0081] In step A2, each candidate text layout box corresponds to a reference layout box internal structure information. This reference layout box internal structure information includes the sub-layout structure information to which the matching character line box belongs for each candidate text layout box. Based on the character content of the matching character line box, a target layout structure segmentation rule is determined from the preset layout structure segmentation rules. The reference layout box internal structure information is then determined based on the target layout structure segmentation rule. For example, if the character content in different matching character line boxes includes different numbers, it can be determined that the target layout structure segmentation rule is a splitting rule triggered by list numbers. Therefore, the sub-layout structure information to which the matching character line box belongs can be determined to be a list line box, and reference layout box internal structure information indicating that the matching character line box is a list line box can be generated.
[0082] In step A2, each candidate text layout box corresponds to a target layout box's internal structural information. This target layout box's internal structural information includes the coordinates of the reference text layout box synthesized from the matching character line boxes and its layout description information. The reference text layout box is obtained by merging all the matching character line boxes corresponding to the candidate text layout box. The layout description information of the reference text layout box describes its layout structure. For example, the layout description information might indicate that the character content of the reference text layout box consists of line characters with different list numbers, or it might indicate that the character content of the reference text layout box consists of multiple vertical bar characters. It is understandable that some matching character line boxes have their entire boundaries within the candidate text layout box's boundaries, while some matching character line boxes have only part of their boundaries outside the candidate text layout box's boundaries. Therefore, it is necessary to redefine a new layout box, i.e., the reference text layout box, based on the matching character line boxes.
[0083] In step A2, the text layout reconstruction information includes the coordinate information of each reference text layout box, the character content contained in each reference text layout box, the layout description information of each reference text layout box, the coordinate information of each unmatched character line box, and the character content contained in each unmatched character line box.
[0084] Therefore, by matching the character line boxes in the character recognition information with the candidate text layout boxes in the text layout structure information to reconstruct the layout structure, the accuracy of the layout structure can be guaranteed, thereby ensuring that the subsequent translation typesetting can match the original image.
[0085] S307, Text translation results are obtained by processing text translation based on text layout reconstruction information.
[0086] In one embodiment, text layout reconstruction information is input into a large-scale text translation model to obtain an initial text translation result. The initial text translation result is then subjected to format validation processing to obtain the final text translation result. During the format validation process, if any abnormal format translation result exists in the initial text translation result, the abnormal format translation result is adjusted to a normal format translation result.
[0087] Understandably, a large text translation model can be obtained by fine-tuning a large multimodal model, or it can be directly adopted.
[0088] S308, determine a reference mask image based on text layout reconstruction information, perform image size scaling on the reference mask image to obtain a target mask image, determine an initial distortion-corrected text image, and perform image size scaling on the initial distortion-corrected text image to obtain a target text image.
[0089] In one embodiment, a target text box is determined based on text layout reconstruction information. The mask of pixels within the target text box in the initial distortion-corrected text image is determined as a first mask value, and the mask of pixels outside the target text box in the initial distortion-corrected text image is determined as a second mask value. A reference mask image is generated based on the first and second mask values. The first mask value indicates that pixels within the target text box are text pixels, and the second mask value indicates that pixels outside the target text box are non-text pixels.
[0090] Obtain the preset size, scale the reference mask image proportionally to the preset size to obtain the target mask image, and scale the initial distortion-corrected text image proportionally to the preset size to obtain the target text image.
[0091] S309, determine the target mask region in the target mask image, determine the target mask mapping region corresponding to the target mask region in the target text image, and perform background restoration processing on the target mask mapping region in the target text image to obtain the restored text image.
[0092] In one embodiment, the region containing pixels in the target mask image whose mask value is a first mask value is determined as the target mask region. The target mask image and the target text image are of the same size, and pixels at the same position in the target mask image and the target text image correspond one-to-one. Therefore, the target mask mapping region corresponding to the target mask region can be determined in the target text image.
[0093] Understandably, repairing a text image involves removing characters from the target mask mapping area to create a clean background for that area, while keeping other areas of the target text image unchanged.
[0094] S310, perform image size restoration processing on the repaired text image to obtain a candidate text image, determine the reference mask region in the reference mask image, determine the pixel information of the first reference mask mapping region corresponding to the reference mask region in the candidate text image, and determine the pixel information of the second reference mask mapping region corresponding to the reference mask region in the distortion-corrected text image.
[0095] In one embodiment, the repaired text image is proportionally restored to a candidate text image of the same size as the initial distortion-corrected text image. It is understood that the difference between the candidate text image and the initial distortion-corrected text image is that the pixel information of all pixels in the candidate text image except for the character regions is the same, while the character pixels within the character regions of the candidate text image are replaced with pixels representing a clean background without text.
[0096] In one embodiment, the region containing pixels in the reference mask image whose mask value is a first mask value is determined as the reference mask region. A first mapping region of the reference mask region in the candidate text image is determined, and pixel information of the first reference mask mapping region of the first mapping region is obtained. The pixel information of the first reference mask mapping region is the pixel information of the pixels in the first mapping region. Accumulated image flattening deformation information is obtained, and a second mapping region of the reference mask region in the initially distorted text image is determined based on the accumulated image flattening deformation information. Pixel information of the second reference mask mapping region of the second mapping region is obtained, and the pixel information of the pixels in the second mapping region is the pixel information of the pixels in the second mapping region. The accumulated image flattening deformation information may include the cumulative displacement vector required for the pixels to move from the initial distorted text image to the final elimination of image distortion.
[0097] S311, in the initial distortion-corrected text image, the pixel information of the second reference mask mapping region is replaced with the pixel information of the first reference mask mapping region to obtain the translated text rendering image.
[0098] It is understandable that the pixels at the same position in the initial distortion-corrected text image and the candidate text image are in one-to-one correspondence. The pixel information of the pixels in the second mapping region is replaced with the pixel information of the corresponding pixels in the first mapping region in turn to obtain the translated text rendering image.
[0099] Therefore, by mapping the mask area back to the initial distortion-corrected text image, the background of the text area in the initial distortion-corrected text is repaired to solve the problems of background destruction and floating of the translation, so that the translation and the original background are consistent in texture, noise and lighting, realizing seamless rewriting of the translation and improving the visual matching effect between the translation and the original background.
[0100] S312, determine the image distortion restoration information corresponding to the distortion-corrected text image, and perform distortion restoration processing on the translated rendering image based on the image distortion restoration information to obtain the distortion-restored text image.
[0101] In one embodiment, determining the image distortion restoration deformation information may specifically include: acquiring at least two rounds of image flattening deformation information corresponding to the distortion-corrected text image, determining the first round of image flattening deformation information from the at least two rounds of image flattening vector information, and determining the image distortion restoration deformation information corresponding to the first round of image flattening deformation information.
[0102] Understandably, the image flattening distortion information is obtained by progressively eliminating distortion from the initial distortion-corrected text image to the original distortion-corrected text image. Each round of image flattening distortion information includes the displacement vectors of the pixels in the current text image required to eliminate distortion in that round. In the first round, the current text image is the initial distortion-corrected text image; in the second round, the current text image is the image obtained after the first round of distortion correction from the initial distortion-corrected text image; and so on. In the final round, the current text image is the image obtained after the penultimate round of distortion correction. The image distortion restoration distortion information includes the displacement vectors of the pixels required to restore the image from the first round of distortion elimination to the initial distortion-corrected text image.
[0103] It can also be understood that the distorted restored text image is an image with the same distortion deformation as the initial distorted corrected text image.
[0104] S313, determine character style information based on text layout reconstruction information, and perform translation back-writing processing on the distorted restored text image based on character style information and text translation result information to obtain the initial translated text image.
[0105] The character style information includes the character size and character color for each target text box.
[0106] In one embodiment, the short side length of each target text box is determined based on text layout reconstruction information. Candidate character sizes corresponding to each target text box are then determined based on the short side length. All candidate character sizes are clustered to obtain reference character sizes of a preset category. The character size corresponding to each target text box is then determined from the preset category of reference character sizes based on the candidate character sizes. Here, the character size corresponding to each target text box is one of the preset category of reference character sizes. For example, if the preset category is 3, the reference character sizes include 12px, 18px, and 24px. A certain target text box may have a character size of 12px, some target text boxes may have a character size of 18px, and another group of target text boxes may have a character size of 24px.
[0107] In one embodiment, foreground pixel data corresponding to each target text box is sampled from the text image to be translated, and background pixel data corresponding to each target text box is sampled from the background restoration area of the translated image. The mean value of the background pixels is determined, the Euclidean distance between each foreground pixel data and the mean value of the background pixels is determined, the target foreground pixel data corresponding to the maximum Euclidean distance is determined, and the color indicated by the target foreground pixel data is determined as the character color corresponding to the target text box. The foreground pixel data, background pixel data, and mean value of the background pixels all include values for the RGB three color channels.
[0108] It is understandable that the background restoration area of the translated image is the mask area containing only the pure background obtained by removing the text.
[0109] In one embodiment, the text box to be written for each target text box is determined from the distorted and restored text image, the text to be written for each target text box is determined according to the text translation result, the text to be written for each target text box is adjusted according to the character style information to obtain the target text, and the target text is typed and written in the text box to obtain the initial translated text image.
[0110] Optionally, during the process of typesetting and writing the target text into the text box to obtain the initial translated text image, the target text can be rotated according to the average micro-tilt angle corresponding to the text box to make the text direction in the initial translated text image consistent with the perspective / tilt of the original image. Furthermore, the target text after the angle is rotated and the distorted restored text image are combined with transparency to achieve a pixel-level natural transition.
[0111] Therefore, by writing the translation back into the restored distorted image, the translation is avoided from being overcorrected, the naturalness of the translation is enhanced, and the translation looks like it matches the distortion of the image.
[0112] S314, perform perspective transformation inverse processing and segmentation restoration on the initial translated text image to obtain the target translated text image.
[0113] In one embodiment, the homography matrix used to obtain the initial distortion-corrected text image in step S302 is acquired, the inverse of the homography matrix is determined, and the initial translated text image is subjected to perspective transformation inverse processing based on the inverse of the homography matrix to obtain a candidate translated text image. When the text region image is segmented from the text image to be translated, the candidate translated text image is merged into the text image to be translated to obtain the target translated text image. When the text region image is the text image to be translated, the candidate translated text image is determined as the target translated text image.
[0114] In the image processing method provided in this application embodiment, an image of the text to be translated is acquired; image segmentation processing is performed on the image of the text to be translated to obtain a text region image; perspective transformation processing is performed on the text region image to obtain an initial distortion-corrected text image; image de-distortion processing is performed on the initial distortion-corrected text image to obtain a distortion-corrected text image; character recognition processing is performed on the distortion-corrected text image to obtain character recognition information; text layout recognition processing is performed on the distortion-corrected text image to obtain text layout structure information; candidate text layout frames are determined based on the text layout structure information; character line frames are determined based on the character recognition information; the overlap area ratio between the character line frames and the candidate text layout frames is determined; and the text is then processed based on the overlap area ratio and a preset layout structure segmentation rule. Text layout reconstruction processing yields text layout reconstruction information, which, by eliminating global perspective and local curvature, results in a flattened and straightened text image. Character recognition and layout recognition are then performed on this flattened and straightened text image, ensuring accuracy in both. Subsequently, text translation processing is performed based on the text layout reconstruction information to obtain the translation result. A reference mask image is determined based on the text layout reconstruction information, and its size is scaled to obtain a target mask image. An initial distortion-corrected text image is then determined, and its size is scaled to obtain the target text image. The target mask region is then identified within the target mask image, and the corresponding region is determined within the target text image. The target mask mapping region is used to perform background restoration processing on the target text image to obtain a restored text image. The restored text image is then subjected to image size restoration processing to obtain a candidate text image. A reference mask region is determined in the reference mask image. The pixel information of the first reference mask mapping region corresponding to the reference mask region is determined in the candidate text image. The pixel information of the second reference mask mapping region corresponding to the reference mask region is determined in the distortion-corrected text image. In the initial distortion-corrected text image, the pixel information of the second reference mask mapping region is replaced with the pixel information of the first reference mask mapping region to obtain the translated text rendering image. Thus, by restoring the background of the text region, the problems of background destruction, floating text, and edge distortion in the translated text are effectively solved. Issues such as misalignment are addressed; the distortion information corresponding to the distortion-corrected text image is determined, and the distortion-restored text image is obtained by performing distortion-restored processing on the translated text rendering image based on the distortion-restored distortion information. Character style information is determined based on text layout reconstruction information, and the initial translated text image is obtained by performing translation back-writing processing on the distorted text image based on the character style information and text translation result information. The initial translated text image is then subjected to perspective transformation inverse processing and segmentation restoration to obtain the target translated text image. Thus, by rendering along the original perspective / curvature direction after the translation is written back, the final translated text image is obtained, achieving seamless presentation with zero geometric changes, ensuring visual matching of the translation with the original image layout and distortion, and further improving the visual display effect of the translation.
[0115] The following will combine Figure 4 This application provides a detailed description of the image processing apparatus provided in its embodiments. It should be noted that... Figure 4 The image processing apparatus shown is used to execute this application. Figures 2-3 The methods shown in the embodiments are for illustrative purposes only, illustrating the parts relevant to the embodiments of this application. For specific technical details not disclosed, please refer to this application. Figures 2-3 The example shown.
[0116] Please see Figure 4 This diagram illustrates the structure of an image processing apparatus according to an embodiment of this application. The image processing apparatus 1 can be implemented as all or part of the apparatus through software, hardware, or a combination of both. According to some embodiments, the image processing apparatus 1 includes an image deformation correction module 11, an image layout recognition module 12, an image text translation module 13, and a translated text synthesis module 14, specifically used for: Image distortion correction module 11 is used to acquire the text image to be translated and perform image distortion correction processing on the text image to be translated to obtain a distortion-corrected text image. The image layout recognition module 12 is used to perform character recognition processing on the distortion-corrected text image to obtain character recognition information, perform text layout recognition processing on the distortion-corrected text image to obtain text layout structure information, and perform text layout reconstruction processing based on the character recognition information and the text layout structure information to obtain text layout reconstruction information. Image-text translation module 13 is used to perform text translation processing based on the text layout reconstruction information to obtain a text translation result, and to perform text background repair processing based on the text layout reconstruction information to obtain a translated text rendering image; The translation text synthesis module 14 is used to perform translation text backfilling and image distortion restoration processing based on the text layout reconstruction information, the translated text rendering image, and the text translation result to obtain the target translated text image.
[0117] Optionally, the image deformation correction module 11 is specifically used for: The text image to be translated is segmented to obtain a text region image, and a perspective transformation is performed on the text region image to obtain an initial distortion-corrected text image. The initial distortion-corrected text image is subjected to image dedistortion processing to obtain the distortion-corrected text image.
[0118] Optionally, the image layout recognition module 12 includes: The layout detection unit is used to determine each candidate text layout box based on the text layout structure information and to determine each character line box based on the character recognition information. The layout reconstruction unit is used to determine the overlap area ratio between the character line box and the candidate text layout box, and to perform text layout reconstruction processing based on the overlap area ratio and the preset layout structure segmentation rules to obtain text layout reconstruction information.
[0119] Optional, layout reconfiguration unit, used for: Obtain a preset area ratio threshold, and perform matching processing on the character line box and the candidate text layout box based on the overlapping area ratio and the preset area ratio threshold to obtain the matching character line box and the unmatched character line box corresponding to the candidate text layout box; Based on the preset layout structure segmentation rules, the matching character line boxes are segmented within the layout box to obtain reference layout box structure information. Based on the reference layout box structure information, the target layout box structure information corresponding to the candidate text layout box is determined. Based on the target layout box structure information and the unmatched character line boxes, text layout reconstruction information is obtained.
[0120] Optional, image-to-text translation module 13, specifically used for: Based on the text layout reconstruction information, a reference mask image is determined, and the reference mask image is scaled to obtain a target mask image. An initial distortion-corrected text image is obtained, and the initial distortion-corrected text image is scaled to obtain a target text image. A target mask region is determined in the target mask image, a target mask mapping region corresponding to the target mask region is determined in the target text image, and background restoration processing is performed on the target mask mapping region in the target text image to obtain a restored text image. The repaired text image is subjected to image size restoration processing to obtain a candidate text image. A reference mask region is determined in the reference mask image. The pixel information of the first reference mask mapping region corresponding to the reference mask region is determined in the candidate text image. The pixel information of the second reference mask mapping region corresponding to the reference mask region is determined in the initial distortion-corrected text image. In the initial distortion-corrected text image, the pixel information of the second reference mask mapping region is replaced with the pixel information of the first reference mask mapping region to obtain the translated text rendering image.
[0121] Optionally, the translated text synthesis module 14 includes: The first image restoration unit is used to determine the image distortion restoration deformation information corresponding to the distortion-corrected text image, and to perform distortion restoration processing on the translated rendered image based on the image distortion restoration deformation information to obtain the distorted text image. The translation write-back unit is used to determine character style information based on the text layout reconstruction information, and to perform translation write-back processing on the distorted restored text image based on the character style information and the text translation result information to obtain an initial translated text image. The second image restoration unit is used to perform perspective transformation inverse processing and segmentation restoration on the initial translated text image to obtain the target translated text image.
[0122] Optionally, the first image restoration unit is specifically used for: Obtain at least two rounds of image flattening deformation information corresponding to the distortion-corrected text image; Determine the first round of image flattening deformation information from the at least two rounds of image flattening vector information, and determine the image distortion and restoration deformation information corresponding to the first round of image flattening deformation information.
[0123] Please refer to Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include one or more of the following components: a processor 110, a memory 120, an input device 140, an output device 140, and a bus 150. The processor 110, memory 120, input device 140, and output device 140 can be connected via the bus 150.
[0124] Processor 110 may include one or more processing cores. Processor 110 connects to various parts of the electronic device using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 120, and by calling data stored in memory 120. Optionally, processor 110 may be implemented using at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). Processor 110 may integrate one or more of a central processing unit (CPU), graphics processing unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 110 and may be implemented separately using a communication chip.
[0125] The memory 120 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 120 may include a non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (e.g., touch function, sound playback function, image playback function, etc.), instructions for implementing the various method embodiments described below, etc. The operating system may be the Android system, including systems deeply developed based on the Android system, the iOS system developed by Apple Inc., including systems deeply developed based on the iOS system, or other systems.
[0126] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to establish data communication between the third-party applications and the operating system. This would allow the operating system to obtain the current scenario information of the third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.
[0127] The input device 130 is used to receive input instructions or data, and includes, but is not limited to, a keyboard, mouse, camera, microphone, or touch device. The output device 140 is used to output instructions or data, and includes, but is not limited to, a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 can be a touch display screen.
[0128] The touch display screen can be designed as a full-screen, curved screen, or irregularly shaped screen. It can also be designed as a combination of a full-screen and a curved screen, or a combination of an irregularly shaped screen and a curved screen; however, this application does not limit the specific design of the touch display screen.
[0129] In addition, those skilled in the art will understand that the structure of the electronic device shown in the above figures does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the electronic device may also include radio frequency circuits, input units, sensors, audio circuits, Wireless Fidelity (WiFi) modules, power supplies, Bluetooth modules, etc., which will not be described in detail here.
[0130] In some embodiments, Figure 5In the illustrated electronic device, the processor 110 can be used to call a program for an image processing method stored in the memory 120, and specifically perform the following operations: Obtain the text image to be translated, and perform image distortion correction processing on the text image to be translated to obtain a distortion-corrected text image; The distortion-corrected text image is subjected to character recognition processing to obtain character recognition information, and the distortion-corrected text image is subjected to text layout recognition processing to obtain text layout structure information. Based on the character recognition information and the text layout structure information, text layout reconstruction processing is performed to obtain text layout reconstruction information. Based on the text layout reconstruction information, text translation processing is performed to obtain the text translation result, and based on the text layout reconstruction information, text background repair processing is performed to obtain the translated text rendering image; Based on the text layout reconstruction information, the translated text rendering image, and the text translation result, the target translated text image is obtained through translated text backfilling and image distortion restoration.
[0131] In one embodiment, when the processor 110 performs image distortion correction processing on the text image to be translated to obtain a distortion-corrected text image, it specifically performs the following steps: The text image to be translated is segmented to obtain a text region image, and a perspective transformation is performed on the text region image to obtain an initial distortion-corrected text image. The initial distortion-corrected text image is subjected to image dedistortion processing to obtain the distortion-corrected text image.
[0132] In one embodiment, when the processor 110 performs the text layout reconstruction process based on the character recognition information and the text layout structure information to obtain text layout reconstruction information, it specifically performs the following operations: Based on the text layout structure information, each candidate text layout box is determined, and based on the character recognition information, each character line box is determined. Determine the overlap ratio between the character line box and the candidate text layout box, and perform text layout reconstruction processing based on the overlap ratio and preset layout structure segmentation rules to obtain text layout reconstruction information.
[0133] In one embodiment, when the processor 110 performs the text layout reconstruction process based on the overlapping area ratio and preset page structure segmentation rules to obtain text layout reconstruction information, it specifically performs the following operations: Obtain a preset area ratio threshold, and perform matching processing on the character line box and the candidate text layout box based on the overlapping area ratio and the preset area ratio threshold to obtain the matching character line box and the unmatched character line box corresponding to the candidate text layout box; Based on the preset layout structure segmentation rules, the matching character line boxes are segmented within the layout box to obtain reference layout box structure information. Based on the reference layout box structure information, the target layout box structure information corresponding to the candidate text layout box is determined. Based on the target layout box structure information and the unmatched character line boxes, text layout reconstruction information is obtained.
[0134] In one embodiment, when the processor 110 performs the text background restoration processing based on the text layout reconstruction information to obtain the translated text rendering image, it specifically performs the following operations: Based on the text layout reconstruction information, a reference mask image is determined, and the reference mask image is scaled to obtain a target mask image. An initial distortion-corrected text image is obtained, and the initial distortion-corrected text image is scaled to obtain a target text image. A target mask region is determined in the target mask image, a target mask mapping region corresponding to the target mask region is determined in the target text image, and background restoration processing is performed on the target mask mapping region in the target text image to obtain a restored text image. The repaired text image is subjected to image size restoration processing to obtain a candidate text image. A reference mask region is determined in the reference mask image. The pixel information of the first reference mask mapping region corresponding to the reference mask region is determined in the candidate text image. The pixel information of the second reference mask mapping region corresponding to the reference mask region is determined in the initial distortion-corrected text image. In the initial distortion-corrected text image, the pixel information of the second reference mask mapping region is replaced with the pixel information of the first reference mask mapping region to obtain the translated text rendering image.
[0135] In one embodiment, when the processor 110 performs the translation text backfilling process and image distortion restoration process based on the text layout reconstruction information, the translated text rendering image, and the text translation result to obtain the target translated text image, it specifically performs the following operations: Determine the image distortion and restoration information corresponding to the distortion-corrected text image, and perform distortion restoration processing on the translated rendered image based on the image distortion and restoration information to obtain the distortion-restored text image; Based on the text layout reconstruction information, character style information is determined, and based on the character style information and the text translation result information, the distorted restored text image is processed to obtain the initial translated text image; The initial translated text image is subjected to inverse perspective transformation and segmentation to obtain the target translated text image.
[0136] In one embodiment, when the processor 110 executes the step of determining the image distortion restoration information corresponding to the distortion-corrected text image, it specifically performs the following operations: Obtain at least two rounds of image flattening deformation information corresponding to the distortion-corrected text image; Determine the first round of image flattening deformation information from the at least two rounds of image flattening vector information, and determine the image distortion and restoration deformation information corresponding to the first round of image flattening deformation information.
[0137] This application also provides a computer-readable storage medium storing at least one instruction that is executed by a processor to implement the image processing method as described in the above embodiments.
[0138] This application also provides a computer program product that stores at least one instruction, which is loaded and executed by the processor to implement the image processing method described in the above embodiments.
[0139] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0140] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An image processing method, characterized in that, The method includes: Obtain the text image to be translated, and perform image distortion correction processing on the text image to be translated to obtain a distortion-corrected text image; The distortion-corrected text image is subjected to character recognition processing to obtain character recognition information, and the distortion-corrected text image is subjected to text layout recognition processing to obtain text layout structure information. Based on the character recognition information and the text layout structure information, text layout reconstruction processing is performed to obtain text layout reconstruction information. Based on the text layout reconstruction information, text translation processing is performed to obtain the text translation result, and based on the text layout reconstruction information, text background repair processing is performed to obtain the translated text rendering image; Based on the text layout reconstruction information, the translated text rendering image, and the text translation result, the target translated text image is obtained through translated text backfilling and image distortion restoration.
2. The method according to claim 1, characterized in that, The step of performing image distortion correction processing on the text image to be translated to obtain a distortion-corrected text image includes: The text image to be translated is segmented to obtain a text region image, and a perspective transformation is performed on the text region image to obtain an initial distortion-corrected text image. The initial distortion-corrected text image is subjected to image dedistortion processing to obtain the distortion-corrected text image.
3. The method according to claim 1, characterized in that, The text layout reconstruction process based on the character recognition information and the text layout structure information to obtain text layout reconstruction information includes: Based on the text layout structure information, each candidate text layout box is determined, and based on the character recognition information, each character line box is determined. Determine the overlap ratio between the character line box and the candidate text layout box, and perform text layout reconstruction processing based on the overlap ratio and preset layout structure segmentation rules to obtain text layout reconstruction information.
4. The method according to claim 3, characterized in that, The text layout reconstruction process based on the overlapping area ratio and preset page structure segmentation rules yields text layout reconstruction information, including: Obtain a preset area ratio threshold, and perform matching processing on the character line box and the candidate text layout box based on the overlapping area ratio and the preset area ratio threshold to obtain the matching character line box and the unmatched character line box corresponding to the candidate text layout box; Based on the preset layout structure segmentation rules, the matching character line boxes are segmented within the layout box to obtain reference layout box structure information. Based on the reference layout box structure information, the target layout box structure information corresponding to the candidate text layout box is determined. Based on the target layout box structure information and the unmatched character line boxes, text layout reconstruction information is obtained.
5. The method according to claim 1, characterized in that, The text background restoration process based on the text layout reconstruction information to obtain the translated text rendering image includes: Based on the text layout reconstruction information, a reference mask image is determined, and the reference mask image is scaled to obtain a target mask image. An initial distortion-corrected text image is obtained, and the initial distortion-corrected text image is scaled to obtain a target text image. A target mask region is determined in the target mask image, a target mask mapping region corresponding to the target mask region is determined in the target text image, and background restoration processing is performed on the target mask mapping region in the target text image to obtain a restored text image. The repaired text image is subjected to image size restoration processing to obtain a candidate text image. A reference mask region is determined in the reference mask image. The pixel information of the first reference mask mapping region corresponding to the reference mask region is determined in the candidate text image. The pixel information of the second reference mask mapping region corresponding to the reference mask region is determined in the initial distortion-corrected text image. In the initial distortion-corrected text image, the pixel information of the second reference mask mapping region is replaced with the pixel information of the first reference mask mapping region to obtain the translated text rendering image.
6. The method according to claim 1, characterized in that, The process of obtaining the target translated text image by performing translated text backfilling and image distortion restoration based on the text layout reconstruction information, the translated text rendering image, and the text translation result includes: Determine the image distortion and restoration information corresponding to the distortion-corrected text image, and perform distortion restoration processing on the translated rendered image based on the image distortion and restoration information to obtain the distortion-restored text image; Based on the text layout reconstruction information, character style information is determined, and based on the character style information and the text translation result information, the distorted restored text image is processed to obtain the initial translated text image; The initial translated text image is subjected to inverse perspective transformation and segmentation to obtain the target translated text image.
7. The method according to claim 6, characterized in that, The step of determining the image distortion and restoration information corresponding to the distortion-corrected text image includes: Obtain at least two rounds of image flattening deformation information corresponding to the distortion-corrected text image; Determine the first round of image flattening deformation information from the at least two rounds of image flattening vector information, and determine the image distortion and restoration deformation information corresponding to the first round of image flattening deformation information.
8. An image processing apparatus, characterized in that, The device includes: An image distortion correction module is used to acquire an image of the text to be translated and to perform image distortion correction processing on the image of the text to be translated to obtain a distortion-corrected text image. The image layout recognition module is used to perform character recognition processing on the distortion-corrected text image to obtain character recognition information, perform text layout recognition processing on the distortion-corrected text image to obtain text layout structure information, and perform text layout reconstruction processing based on the character recognition information and the text layout structure information to obtain text layout reconstruction information. The image-text translation module is used to perform text translation processing based on the text layout reconstruction information to obtain the text translation result, and to perform text background repair processing based on the text layout reconstruction information to obtain the translated text rendering image; The translation text synthesis module is used to perform translation text backfilling and image distortion restoration processing based on the text layout reconstruction information, the translated text rendering image, and the text translation result to obtain the target translated text image.
9. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions, which are adapted to be loaded by a processor and executed as described in any one of claims 1 to 7.
10. An electronic device, characterized in that, include: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed as described in any one of claims 1 to 7.