Method, device, equipment and medium for high-definition amplification of characters in main picture of commodity

By combining the HAT and lexica-DAT super-resolution algorithms with image fusion technology, the problem of high-definition enlargement of text and portraits in product main images is solved, and text clarity is improved while image quality is preserved. It is suitable for product main image processing systems.

CN120725874APending Publication Date: 2025-09-30ZIXUN TECHNOLOGY (FUJIAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510821665.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

When processing product main images, existing image super-resolution algorithms are unable to simultaneously meet the requirements of high-definition magnification of text and portraits, resulting in blurred text edges and loss of stroke details, affecting information communication and commercial display effects.

Method used

Combining the HAT super-resolution algorithm and the lexica-DAT super-resolution algorithm, we use text detection and image fusion technology to specifically optimize the text area in the main product image. We use the Gaussian blur algorithm to process the seams to achieve high-definition text enlargement while maintaining the image quality of the portrait and product parts.

Benefits of technology

It significantly improves text clarity, maintains overall image quality, improves processing efficiency, reduces costs, has good compatibility, and is suitable for existing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725874A_ABST
    Figure CN120725874A_ABST
Patent Text Reader

Abstract

The invention provides a high-definition amplification method, device, equipment and medium for characters of a commodity main picture, and the method comprises the steps: recognizing the characters of the commodity main picture, and obtaining a surrounding frame of the characters; filling the surrounding frame to generate a character mask; and an original image is processed by using an HAT super-division algorithm and a Lexica-DAT super-division algorithm. Carrying out super-division operation on the commodity main image by adopting an HAT super-division algorithm to obtain an HATimage after super-division amplification, and meanwhile, carrying out super-division operation on the commodity main image by adopting a Lexica-DAT super-division algorithm to obtain a Lexicimage; according to the character mask, the HATimage picture and the Lexicimage picture are combined, and a combined image is obtained; by integrating the advantages of an HAT super-division algorithm and a lexica-DAT super-division algorithm and combining character detection and image fusion technologies, high-definition amplification of characters is realized, and the image quality is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a method, device, equipment and medium for high-definition amplification of text in a main image of a product. Background Art

[0002] When processing product main images, image super-resolution algorithms are crucial for improving image clarity. The commonly used HAT super-resolution algorithm, while effective for portraits and products, is prone to blurring text edges and loss of stroke detail when processing text, resulting in unclear text and poor communication. The lexica-DAT super-resolution algorithm, while effective for text, falls short when applied to portraits and products. This makes it difficult to find a single algorithm that meets the requirements for high-definition magnification of all elements in product main images, limiting the quality improvement and commercial display of product main images. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method, device, equipment and medium for high-definition magnification of text in the main image of a product. By integrating the advantages of the HAT super-resolution algorithm and the lexica-DAT super-resolution algorithm, combined with text detection and image fusion technology, the text area in the main image of the product is specially optimized to achieve high-definition magnification of the text, while ensuring the image quality of the portrait and the product part, meeting the dual needs of the main image of the product for clear text and high-quality images in commercial displays.

[0004] In a first aspect, the present invention provides a method for high-definition amplification of text in a main image of a product, comprising the following steps:

[0005] Step 1: Recognize the text in the main image of the product and obtain the bounding box of the text;

[0006] Step 2: Fill the bounding box in the main product image to generate a text mask;

[0007] Step 3: Use the HAT super-resolution algorithm and the Lexica-DAT super-resolution algorithm to process the original image. Use the HAT super-resolution algorithm to super-resolve the main product image to obtain the super-resolved and enlarged HAT_image image. At the same time, use the Lexica-DAT super-resolution algorithm to super-resolve the main product image to obtain the Lexica_image image.

[0008] Step 4: Combine the HAT_image image and the Lexica_image image according to the text mask to obtain a merged image;

[0009] Step 5: Use Gaussian blur algorithm to process the seam produced by merging the images to obtain the desired magnified image.

[0010] In a second aspect, the present invention provides a device for high-definition amplification of text on a product main image, comprising:

[0011] The text area detection module recognizes the text in the main product image and obtains the text bounding box;

[0012] Generate a text mask module to fill the bounding box in the main product image and generate a text mask;

[0013] The image super-resolution processing module uses the HAT super-resolution algorithm and the Lexica-DAT super-resolution algorithm to process the original image. The HAT super-resolution algorithm is used to super-resolution the main product image to obtain the super-resolved and enlarged HAT_image image. At the same time, the Lexica-DAT super-resolution algorithm is used to super-resolution the main product image to obtain the Lexica_image image.

[0014] The image fusion module combines the HAT_image image and the Lexica_image image according to the text mask to obtain the merged image;

[0015] The seam processing module uses Gaussian blur algorithm to process the seam part generated by the merged image to obtain the required enlarged image.

[0016] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.

[0017] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the method described in the first aspect when the program is executed by a processor.

[0018] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:

[0019] 1. Significantly improved text clarity: This invention uses the lexica-DAT super-resolution algorithm to perform targeted processing on text areas, making text edges sharper and stroke details clearly visible. Compared with traditional single super-resolution algorithms, this improves text clarity and effectively solves the problem of blurred text, ensuring that the text information in the main product image can be accurately and clearly conveyed to consumers.

[0020] 2. Excellent overall image quality: While ensuring high-definition text magnification, the present invention retains the processing effects of the HAT super-resolution algorithm on portraits and products. The overall color, details and texture of the image are well displayed, avoiding the problem of quality degradation of some elements due to single algorithm processing, and improving the overall visual effect and commercial value of the main product image.

[0021] 3. High processing efficiency: Through the automated processing flow, the present invention can control the processing time of high-definition text enlargement of a single product main picture within 35 seconds. Compared with manual optimization or complex multi-step processing methods, it greatly improves the processing efficiency, meets the needs of batch processing of product pictures, and reduces labor and time costs.

[0022] 4. Strong compatibility: The present invention can be seamlessly integrated into existing commercial image processing software and pipelines, and has good compatibility with other image processing technologies and tools, making it convenient for users to directly apply it in the original workflow without the need for large-scale transformation of the existing system, thereby reducing the threshold and cost of technology application.

[0023] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0025] Figure 1 This is a flowchart of the method in Example 1 of the present invention;

[0026] Figure 2 This is a schematic diagram of the structure of the device in Example 2 of the present invention. DETAILED DESCRIPTION

[0027] The technical solutions in the embodiments of this application have the following general ideas:

[0028] HAT super-resolution algorithm: A super-resolution algorithm based on deep learning, which performs well in super-resolution processing of portraits and product images, but has defects in text processing.

[0029] Lexica-DAT super-resolution algorithm: This algorithm focuses on text super-resolution processing and can effectively improve text clarity, but it is not very effective for super-resolution of portraits and products.

[0030] PaddleOCR: A text recognition tool that can be used to detect text areas in images.

[0031] The main implementation is as follows:

[0032] Text area detection: Use paddleOCR to process the main product image. Through its built-in text detection model, all text areas are detected to obtain the text bounding box text_box, accurately marking the specific location of the text in the main image; the main product image is a picture that includes the product and text.

[0033] Generate a text mask: Fill the resulting text_box to generate a text_mask, making the text area stand out in the mask. To better cover any blurred areas around the text, dilate the text_mask using a 7*7 kernel size to further expand the text mask and obtain the final text_mask for image fusion.

[0034] Image super-resolution: The original image is processed using the HAT and lexica-DAT super-resolution algorithms. The HAT super-resolution algorithm performs super-resolution on the original image, generating the super-resolved image HAT_image, which shows good results for portraits and products. Meanwhile, the lexica-DAT super-resolution algorithm performs super-resolution on the original image, generating lexica_image, which shows better results for text.

[0035] Image fusion: Combine HAT_image and lexica_image using the generated text_mask. Use text_mask as a mask, and use the corresponding pixel values ​​in lexica_image for the areas covered by text_mask; and use the corresponding pixel values ​​in HAT_image for the areas not covered by text_mask, so as to merge the two super-resolution images into a combined image, giving full play to the advantages of the two super-resolution algorithms in different elements.

[0036] The merging formula is: imageA*maskA+imageB*maskB, maskA and maskB are 0 or 1; imageA is the pixel value of HAT_image; imageB is the pixel value of lexica_image;

[0037] Seam processing: The Gaussian blur algorithm is used to process the seams created by merging the images. This smooths the pixel transitions at the seams, effectively eliminating any seam artifacts. The result is a high-definition, enlarged text image with a visually appealing overall look.

[0038] Example 1

[0039] like Figure 1 As shown, this embodiment provides a method for high-definition amplification of the text of a product main image, including the following steps:

[0040] Step 1: Recognize the text in the main image of the product and obtain the bounding box of the text;

[0041] Step 2: Fill the bounding box in the main product image to generate a text mask;

[0042] Step 3: Use the HAT super-resolution algorithm and the Lexica-DAT super-resolution algorithm to process the original image. Use the HAT super-resolution algorithm to super-resolve the main product image to obtain the super-resolved and enlarged HAT_image image. At the same time, use the Lexica-DAT super-resolution algorithm to super-resolve the main product image to obtain the Lexica_image image.

[0043] Step 4: Combine the HAT_image image and the Lexica_image image according to the text mask to obtain a merged image;

[0044] Step 5: Use Gaussian blur algorithm to process the seam produced by merging the images to obtain the desired magnified image.

[0045] In this embodiment, preferably, step 1 specifically includes: using PaddleOCR to recognize the text in the main image of the product, and obtaining a bounding box of the text, which marks the position of the text in the main image of the product.

[0046] In this embodiment, preferably, step 2 specifically includes: filling the bounding box in the main image of the product to generate a text mask, and enlarging the area of ​​the text mask using a 7×7 kernel expansion algorithm to obtain an enlarged text mask.

[0047] In this embodiment, preferably, step 3 is specifically as follows: using the text mask as a mask, and for the area covered by the text mask, using the pixel values ​​of the corresponding area in the Lexica_image picture; for the area not covered by the text mask, using the pixel values ​​of the corresponding area in the HAT_image picture, and merging them to obtain a merged image.

[0048] Based on the same inventive concept, this application also provides a device corresponding to the method in Example 1, see Example 2 for details.

[0049] Example 2

[0050] like Figure 2 As shown, in this embodiment, a device for high-definition amplification of the main image and text of a product is provided, comprising:

[0051] The text area detection module recognizes the text in the main product image and obtains the text bounding box;

[0052] Generate a text mask module to fill the bounding box in the main product image and generate a text mask;

[0053] The image super-resolution processing module uses the HAT super-resolution algorithm and the Lexica-DAT super-resolution algorithm to process the original image. The HAT super-resolution algorithm is used to super-resolution the main product image to obtain the super-resolved and enlarged HAT_image image. At the same time, the Lexica-DAT super-resolution algorithm is used to super-resolution the main product image to obtain the Lexica_image image.

[0054] The image fusion module combines the HAT_image image and the Lexica_image image according to the text mask to obtain the merged image;

[0055] The seam processing module uses Gaussian blur algorithm to process the seam part generated by the merged image to obtain the required enlarged image.

[0056] In this embodiment, preferably, the text area detection module specifically uses PaddleOCR to recognize the text in the main image of the product to obtain a bounding box of the text, and the bounding box marks the position of the text in the main image of the product.

[0057] In this embodiment, preferably, the text mask generating module specifically fills the bounding box in the main product image to generate a text mask, and enlarges the area of ​​the text mask using a 7×7 kernel dilation algorithm to obtain an enlarged text mask.

[0058] In this embodiment, preferably, the image super-resolution processing module is specifically: using the text mask as a mask, for the area covered by the text mask, using the corresponding area pixel value in the Lexica_image picture; for the area not covered by the text mask, using the corresponding area pixel value in the HAT_image picture, and merging them to obtain a merged image.

[0059] Since the device described in the second embodiment of the present invention is used to implement the method of the first embodiment of the present invention, those skilled in the art will be able to understand the specific structure and variations of the device based on the method described in the first embodiment of the present invention, and therefore will not be described in detail here. All devices used in the method of the first embodiment of the present invention fall within the scope of protection of the present invention.

[0060] Based on the same inventive concept, this application provides an electronic device embodiment corresponding to the first embodiment, see the third embodiment for details.

[0061] Example 3

[0062] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, any implementation method in the first embodiment can be implemented.

[0063] Since the electronic device described in this embodiment is the device used to implement the method in Example 1 of this application, based on the method described in Example 1 of this application, those skilled in the art will be able to understand the specific implementation of the electronic device of this embodiment and its various variations. Therefore, how the electronic device implements the method in the embodiment of this application will not be described in detail here. As long as the device used by those skilled in the art to implement the method in the embodiment of this application falls within the scope of protection to be provided by this application.

[0064] Based on the same inventive concept, this application provides a storage medium corresponding to Example 1, see Example 4 for details.

[0065] Example 4

[0066] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, any implementation method in the first embodiment can be implemented.

[0067] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0068] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0069] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0070] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0071] Although the specific embodiments of the present invention are described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and are not intended to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A method for high-definition amplification of product main image text, characterized by: The steps include: Step 1: Recognize the text in the main image of the product and obtain the bounding box of the text; Step 2: Fill the bounding box in the main product image to generate a text mask; Step 3: Use the HAT super-resolution algorithm and the Lexica-DAT super-resolution algorithm to process the original image. Use the HAT super-resolution algorithm to super-resolve the main product image to obtain the super-resolved and enlarged HAT_image image. At the same time, use the Lexica-DAT super-resolution algorithm to super-resolve the main product image to obtain the Lexica_image image. Step 4: Combine the HAT_image image and the Lexica_image image according to the text mask to obtain a merged image; Step 5: Use Gaussian blur algorithm to process the seam produced by merging the images to obtain the desired magnified image.

2. The method for high-definition amplification of product main image text according to claim 1, characterized in that: The step 1 specifically includes: using PaddleOCR to recognize the text in the main image of the product, and obtaining a bounding box of the text, which marks the position of the text in the main image of the product.

3. The method for high-definition amplification of product main image text according to claim 1, characterized in that: The step 2 specifically includes: filling the bounding box in the main image of the product to generate a text mask, and enlarging the area of ​​the text mask using a 7×7 kernel dilation algorithm to obtain an enlarged text mask.

4. The method for high-definition amplification of product main image text according to claim 1, characterized in that: The step 3 is specifically as follows: using the text mask as a mask, for the area covered by the text mask, using the pixel values ​​of the corresponding area in the Lexica_image picture; for the area not covered by the text mask, using the pixel values ​​of the corresponding area in the HAT_image picture, and merging them to obtain a merged image.

5. A device for high-definition amplification of the main image and text of a product, characterized by: include: The text area detection module recognizes the text in the main product image and obtains the text bounding box; Generate a text mask module to fill the bounding box in the main product image and generate a text mask; The image super-resolution processing module uses the HAT super-resolution algorithm and the Lexica-DAT super-resolution algorithm to process the original image. The HAT super-resolution algorithm is used to super-resolution the main product image to obtain the super-resolved and enlarged HAT_image image. At the same time, the Lexica-DAT super-resolution algorithm is used to super-resolution the main product image to obtain the Lexica_image image. The image fusion module combines the HAT_image image and the Lexica_image image according to the text mask to obtain the merged image; The seam processing module uses Gaussian blur algorithm to process the seam part generated by the merged image to obtain the required enlarged image.

6. The device for high-definition magnification of product main images and text according to claim 5, characterized in that: The text area detection module specifically uses PaddleOCR to recognize the text in the main image of the product, and obtains a bounding box of the text, which marks the position of the text in the main image of the product.

7. The device for high-definition magnification of product main images and text according to claim 5, characterized in that: The text mask generation module specifically fills the bounding box in the main product image to generate a text mask, and enlarges the area of ​​the text mask using a 7×7 kernel expansion algorithm to obtain an enlarged text mask.

8. The device for high-definition magnification of product main images and text according to claim 5, characterized in that: The image super-resolution processing module specifically uses the text mask as a mask, and for the area covered by the text mask, uses the pixel values ​​of the corresponding area in the Lexica_image picture; for the area not covered by the text mask, uses the pixel values ​​of the corresponding area in the HAT_image picture, and merges them to obtain a merged image.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.