Image processing method and device, electronic equipment, storage medium and product

Through the image processing method, the text content in tobacco documents is separated and reconstructed using the pre-trained image reconstruction model, which solves the problem of low recognition accuracy caused by misalignment or overlap in the documents, and improves the readability and recognition accuracy of the documents.

CN120148050APending Publication Date: 2025-06-13YUNNAN TOBACCO LEAF
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510222390.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Due to poor accuracy of printing equipment, paper placement errors or copy scanning problems in tobacco documents, the text characters and background content are misaligned or overlapped, and existing recognition technology is difficult to achieve ideal accuracy, affecting the accuracy of subsequent business processes.

Method used

An image processing method is adopted to detect whether the content on the target document meets the preset conditions, obtain the target image and input it into the pre-trained image reconstruction model, and use the encoding module and the decoding module to perform feature extraction and content reconstruction, and separate and reconstruct the text content in the target document.

Benefits of technology

Effectively separate and reconstruct the text content in tobacco documents, improve the readability of the documents and the accuracy of subsequent content recognition, and solve the problem of low recognition accuracy under the problem of text overlap.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148050A_ABST
    Figure CN120148050A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device, electronic equipment, a storage medium and a product. According to the specific scheme, when it is detected that content on a target document meets a preset condition, a target image corresponding to the target document is obtained; inputting the target image into an image reconstruction model obtained by pre-training to obtain target content edited in the target document; wherein the image reconstruction model comprises a coding module and a decoding module, the coding module comprises at least one feature extraction sub-module, the feature extraction sub-module comprises a vector mapping layer and a feature conversion layer, and the feature conversion layer comprises a dynamic position coding sub-layer, a normalization sub-layer, a double-dynamic token mixer sub-layer and a multi-scale feed-forward network. The decoding module comprises an up-sampling layer and a feature conversion layer, and the target content is the content edited by the user in the target document. According to the invention, effective separation of the text content in the tobacco document is realized, and the readability of the tobacco document and the accuracy of subsequent target content identification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an image processing method, apparatus, electronic device, storage medium and product. Background Art

[0002] In the tobacco industry, tobacco documents usually include key information in the tobacco production and transportation processes. For example, tobacco documents may include information such as the type, quantity, production batch, unit price, and amount of tobacco materials. To ensure the smooth progress of the entire tobacco production process, accurate determination of document information is very crucial.

[0003] In actual scenarios, tobacco documents are usually based on pre-printed formats and use printing devices to output pre-edited text content. However, due to differences in the accuracy of printing devices, paper placement errors, or problems with multiple copying and scanning of text content, the finally generated tobacco documents may have the phenomenon that the text characters are misaligned or overlapped with the background content of the preset printing format.

[0004] Currently, tobacco documents are mainly recognized by manual methods or optical character recognition technologies. However, in the case of random overlapping and irregular distribution of text characters in tobacco documents, it is difficult to achieve ideal recognition accuracy through the above two methods, resulting in incorrect recognition of document information, thus affecting the accuracy of subsequent business processes. For example, it may lead to confusion in inventory data, delays in production plans, and even serious financial settlement problems. Summary of the Invention

[0005] The present invention provides an image processing method, apparatus, electronic device, storage medium and product, which realizes effective separation of text content in tobacco documents and improves the readability of tobacco documents and the accuracy of subsequent target content recognition.

[0006] According to one aspect of the present invention, there is provided an image processing method, which includes:

[0007] When it is detected that the content on the target document meets a preset condition, obtaining a target image corresponding to the target document;

[0008] Inputting the target image into a pre-trained image reconstruction model to obtain the target content edited in the target document;

[0009] Among them, the image reconstruction model includes an encoding module and a decoding module. The encoding module includes at least one feature extraction sub-module. The feature extraction sub-module includes a vector mapping layer and a feature transformation layer. The feature transformation layer includes a dynamic position encoding sub-layer, a normalization sub-layer, a double dynamic token mixer sub-layer, and a multi-scale feed-forward network. The decoding module includes an upsampling layer and a feature transformation layer. The target content is the content edited by the user in the target document.

[0010] According to another aspect of the present invention, there is provided an image processing apparatus, which includes:

[0011] A target image acquisition module, configured to acquire a target image corresponding to the target document when it is detected that the content on the target document meets a preset condition;

[0012] A target content determination module, configured to input the target image into a pre-trained image reconstruction model to obtain the target content edited in the target document;

[0013] Among them, the image reconstruction model includes an encoding module and a decoding module. The encoding module includes at least one feature extraction sub-module. The feature extraction sub-module includes a vector mapping layer and a feature transformation layer. The feature transformation layer includes a dynamic position encoding sub-layer, a normalization sub-layer, a double dynamic token mixer sub-layer, and a multi-scale feed-forward network. The decoding module includes an upsampling layer and a feature transformation layer. The target content is the content edited by the user in the target document.

[0014] According to another aspect of the present invention, there is provided an electronic device, which includes:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the image processing method according to any embodiment of the present invention.

[0018] According to another aspect of the present invention, there is provided a computer-readable storage medium, which stores computer instructions for causing a processor to implement the image processing method according to any embodiment of the present invention when executed.

[0019] According to another aspect of the present invention, there is provided a computer program product, including a computer program, characterized in that the computer program implements the image processing method according to any embodiment of the present invention when executed by a processor.

[0020] In the technical solution of the embodiment of the present invention, when it is detected that the content on the target document meets the preset conditions, the target image corresponding to the target document is obtained. The target image is input into the pre-trained image reconstruction model to obtain the target content edited in the target document. Based on this, the effective separation of the overlapping text content in the target document is realized, and the readability of the target document and the recognition accuracy of the subsequent target content are improved. Among them, the image reconstruction model includes an encoding module and a decoding module. The encoding module includes at least one feature extraction sub-module. The feature extraction sub-module includes a vector mapping layer and a feature transformation layer. The feature transformation layer includes a dynamic position encoding sub-layer, a normalization sub-layer, a double dynamic token mixer sub-layer, and a multi-scale feed-forward network. The decoding module includes an upsampling layer and a feature transformation layer. The target content is the content edited by the user in the target document. Through the above-created image reconstruction model, the problem that it is difficult to achieve ideal recognition accuracy by manual means or optical character recognition technology in the case of random overlap of text characters in tobacco documents is solved. The effective separation and reconstruction of the text content in the target document are realized, which not only improves the readability of the target document, but also improves the recognition accuracy of the subsequent optical character recognition technology for the target content.

[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 is a flowchart of an image processing method provided by an embodiment of the present invention;

[0024] Figure 2 is an example diagram of a target image provided by an embodiment of the present invention;

[0025] Figure 3 is an example diagram corresponding to the target content provided by an embodiment of the present invention;

[0026] Figure 4 is a flowchart of an image reconstruction model training method provided by an embodiment of the present invention;

[0027] Figure 5 is an example diagram of the first image of the first sample pair provided by an embodiment of the present invention;

[0028] Figure 6 It is an example diagram of the second image of the first sample pair provided by an embodiment of the present invention;

[0029] Figure 7 It is an example diagram of the structure of the image reconstruction model provided by an embodiment of the present invention;

[0030] Figure 8 It is an example diagram of the structure of the feature transformation layer provided by an embodiment of the present invention;

[0031] Figure 9 It is an example diagram of the dual-dynamic token mixer sublayer processing input features provided by an embodiment of the present invention;

[0032] Figure 10 It is a schematic structural diagram of an image processing device provided by an embodiment of the present invention;

[0033] Figure 11 It is a schematic structural diagram of an electronic device for implementing the image processing method according to an embodiment of the present invention. Detailed implementation manners

[0034] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0036] Embodiment 1

[0037] Figure 1The following is a flowchart of an image processing method provided in the first embodiment of the present invention. This embodiment is applicable to the situation of performing text separation processing on a target document that meets preset conditions based on a pre-trained image reconstruction model, and determining the target content edited in the target document. This method can be executed by an image processing device, which can be implemented in the form of hardware and / or software, and the image processing device can be configured in an electronic device such as a mobile phone, a computer, or a server. As Figure 1 shown, the method includes:

[0038] S110. When it is detected that the content on the target document meets the preset conditions, obtain a target image corresponding to the target document.

[0039] Among them, a document is usually used to record information in various operations. For example, taking a tobacco document as an example, various key information in tobacco production and transportation is usually recorded in the tobacco document. For example, the key information may include the type information of tobacco materials, the quantity information of tobacco materials, the production batch information, the amount information, etc. For each document, the processing method in the embodiment of the present invention is the same. Therefore, the current document is used as the target document for processing. The preset condition may be that there is text overlap in the content on the target document. That is to say, the object to be processed in the embodiment of the present invention in the following is: the target document with text overlap. It should be noted that the situation of text overlap in the target document may be caused by reasons such as paper placement error during document printing, multiple copying, or printing device failure. The target image may be a document image obtained after image acquisition of the target document.

[0040] Specifically, when it is detected that there is text overlap in the content on the target document, it is determined that the content on the target document meets the preset conditions. At this time, image acquisition is performed on the target document to obtain the target image of the target document. Through the above method, it is convenient to perform text separation processing on the target image with text overlap subsequently, thereby improving the readability of the target document.

[0041] Optionally, the target document is a tobacco document.

[0042] Specifically, in the case where the target document is a tobacco document, the content on the tobacco document is detected. When it is detected that there is text overlap in the content on the tobacco document, it is determined that the content on the tobacco document meets the preset conditions, and image acquisition is performed on the tobacco document to obtain the target image corresponding to the tobacco document. Based on this, it is convenient to perform efficient separation processing on the target document with text overlap through the target image subsequently.

[0043] S120. Input the target image into the pre-trained image reconstruction model to obtain the target content edited in the target document.

[0044] Among them, the image reconstruction model can be a model for separating text from overlapping content information in the target image. The image reconstruction model includes an encoding module and a decoding module. The encoding module includes at least one feature extraction sub-module. The feature extraction sub-module includes a vector mapping layer and a feature transformation layer. The feature transformation layer includes a dynamic position encoding sub-layer, a normalization sub-layer, a double dynamic token mixer sub-layer, and a multi-scale feed-forward network. The decoding module includes an upsampling layer and a feature transformation layer. The target content is the content edited by the user in the target document.

[0045] In the encoding module, the feature extraction sub-module is mainly used to perform feature extraction processing on the content information in the target image. The vector mapping layer in the feature extraction sub-module is used to convert the image data corresponding to the target image into vector-form data. Optionally, the Patch Embedding layer is used as an example for the vector mapping layer. The target image is sliced through the Patch Embedding layer to obtain multiple image patches corresponding to the target image. An embedding operation is performed on each image patch, that is, each image patch is mapped into a high-dimensional vector space. It should be noted that when performing the embedding operation on the image patches in the vector mapping layer, a convolutional kernel with the same size as the image patch can be used. Among them, the number of convolutional kernels is the dimension corresponding to the embedding operation. Since the above-mentioned slicing operation destroys the position information of the target image, position encoding can be introduced to determine the position information of each image patch in the target image. Among them, the position encoding can be obtained through trigonometric functions or optimized through the training process of the subsequent image reconstruction model.

[0046] The feature transformation layer in the feature extraction sub-module is usually used to extract the global features and local features corresponding to the image. The feature transformation layer includes: a dynamic position encoding sub-layer, a normalization sub-layer, a dual-dynamic token mixer sub-layer, and a multi-scale feed-forward network. The dynamic position encoding sub-layer in the feature transformation layer is used to add position information to the input feature map to determine the spatial relationship corresponding to the input feature map. The normalization sub-layer in the feature transformation layer is used to perform normalization processing on the input feature map. The dual-dynamic token mixer sub-layer further extracts local features and global features from the feature data input to this layer by introducing a dynamic token mixing mechanism. Optionally, the dual-dynamic token mixer sub-layer contains two parallel dynamic token mixing branches. One branch is an attention unit based on the Overlapping Spatial Reduction Attention (OSRA) mechanism, which is used to perform global modeling on the input feature map to extract global context information. The other branch is an Input-dependent Depthwise Convolution (IDConv) unit, which is used to extract the boundaries and detailed information of the text characters in the input feature map. Optionally, the input-dependent depthwise convolution unit contains at least four parallel convolution operations, and each convolution operation can be implemented using convolution kernels of different sizes. For example, taking the input-dependent depthwise convolution unit containing four parallel convolution operations as an example, the sizes of the convolution kernels can be 1×1, 3×3, 5×5, and 7×7 respectively. Optionally, in the attention unit based on the overlapping spatial reduction attention mechanism, different numbers of attention heads can be used to process the input data in parallel. For example, the number of attention heads of the attention unit based on the overlapping spatial reduction attention mechanism can be 1, 2, 4, and 8 respectively. The multi-scale feed-forward network in the feature transformation layer is used to capture the feature information of the input feature map at different scales. The multi-scale feed-forward network can include multiple convolutional layers and transformation layers to capture features at different scales.

[0047] The decoding module includes an upsampling layer and a feature transformation layer. The upsampling layer in the decoding module is used to perform upsampling processing on the input feature map. Optionally, the upsampling layer can include a convolutional sub-layer and a PixelShuffle layer. Among them, the convolutional sub-layer can be used to extract image features from the low-resolution feature map. The PixelShuffle layer upsamples the low-resolution feature map to a high-resolution feature map by means of pixel recombination.

[0048] Specifically, the target image is input into a pre-trained image reconstruction model, and the encoding module of the image reconstruction model performs encoding processing on the target image to obtain an encoded feature map. Among them, the specific encoding process can be: the input feature map corresponding to the target image is sequentially processed through the vector mapping layer and the feature transformation layer of at least one feature extraction sub-module in the encoding module to obtain the encoded feature map corresponding to the target image. Taking the encoded feature map output by the encoding module as the input of the decoding module of the image reconstruction model, the encoded feature map is upsampled through the upsampling layer of the decoding module, and the upsampled feature map is decoded through the feature transformation layer of the decoding module to obtain an output image, so as to determine the target content edited in the target document through the output image.

[0049] Optionally, after obtaining the target content edited in the target document, the target content can be recognized by optical character recognition technology to determine the content information of the target document.

[0050] Exemplarily, see Figure 2 , Figure 2 is an example diagram of the target image. Input Figure 2 into the pre-trained image reconstruction model, and perform feature extraction processing on the target image shown in Figure 2 through at least one feature extraction sub-module in the encoding module of the image reconstruction model to obtain an encoded feature map. The encoded feature map is upsampled through the upsampling layer in the decoding module of the image reconstruction model, and the upsampled feature map is decoded through the feature transformation layer to obtain the output image corresponding to the target document, so as to determine the target content edited in the target document through the output image. For example, Figure 3 is an example diagram of the output image. Figure 3 The content in black font shown is the target content.

[0051] In the technical solution of this embodiment, when it is detected that the content on the target document meets the preset conditions, the target image corresponding to the target document is obtained. The target image is input into the pre-trained image reconstruction model to obtain the target content edited in the target document. Based on this, the effective separation of the overlapping text content in the target document is realized, and the readability of the target document and the recognition accuracy of the subsequent target content are improved. Among them, the image reconstruction model includes an encoding module and a decoding module. The encoding module includes at least one feature extraction sub-module. The feature extraction sub-module includes a vector mapping layer and a feature transformation layer. The feature transformation layer includes a dynamic position encoding sub-layer, a normalization sub-layer, a double dynamic token mixer sub-layer, and a multi-scale feed-forward network. The decoding module includes an upsampling layer and a feature transformation layer. The target content is the content edited by the user in the target document. Through the above-created image reconstruction model, the problem that it is difficult to achieve ideal recognition accuracy by manual means or optical character recognition technology in the case of random overlap of text characters in tobacco documents is solved, and the effective separation and reconstruction of the text content in the target document are realized, which not only improves the readability of the target document, but also improves the accuracy of subsequent optical character recognition technology for recognizing the target content.

[0052] Embodiment 2

[0053] Figure 4 FIG. is a flowchart of a method for training an image reconstruction model provided in Embodiment 2 of the present invention. Based on the above embodiment, before processing the target image based on the pre-trained image reconstruction model, the image reconstruction model can also be trained based on the constructed sample set. The specific implementation manner can refer to the technical solution of this embodiment. Among them, the same or corresponding technical terms as those in the above embodiment will not be described in detail here. As Figure 4 shown, the method includes:

[0054] S210. Construct a sample set for training the image reconstruction model.

[0055] In the embodiment of the present invention, the specific way to construct the sample set can be: obtain a plurality of first sample pairs, where each first sample pair includes a first image and a second image. The first image is an image with overlapping text in the first document, and the second image is an image corresponding to the text edited in the first image; by processing the first image and the second image in the plurality of first sample pairs, a second sample pair corresponding to the first sample pair is obtained; based on the first sample pair and the second sample pair, the sample set is determined.

[0056] Among them, the first document can be understood as a pre-acquired sample document with overlapping text. For the sample document, there may only be a situation where there is overlapping text in some areas. Then, the sample document with overlapping text can be used as the first document, and the image corresponding to the area with overlapping text in the first document can be used as the first image. The second image can be understood as the image corresponding to the text edited in the first image. In other words, the second image is an image that corresponds to the first image and has no overlapping text. For example, see Figure 5 and Figure 6 , Figure 5 is an example diagram of the first image, Figure 6 is the second image corresponding to the text edited in the first image and having no overlapping text. It should be noted that there may be text information edited twice or more times in the first image. Then, the text information of the latest edit, or the text information of a preset color, or the text information that conforms to the preset format of the target document is used as the text information included in the second image. For example, Figure 6 uses the black text information as the text information included in the second image.

[0057] Processing multiple first sample pairs can be to perform image enhancement processing on the first sample pairs to increase the sample pairs used to train the image reconstruction model. For example, performing random cropping, horizontal or vertical rotation, or random transposition on the first sample pairs to obtain multiple sample pairs corresponding to the first sample pairs, that is, the second sample pairs. The sample set is a set of sample images containing the first sample pairs and the second sample pairs.

[0058] Specifically, before training the image reconstruction model, a sample set can be obtained first to train the model through the sample set. To improve the accuracy of the model, as many and as rich sample sets as possible can be obtained. That is, obtain multiple first sample pairs, where each first sample pair contains a first image and a second image. The first image is all or part of the image of the first document with overlapping text, and the second image is an image that corresponds to the first image and only contains the text information of the latest edit, or the text information of a preset color, or the text information that conforms to the preset format of the target document in the first image. Further, to enrich the sample set, image enhancement processing can be performed on the first image and the second image in each first sample pair to obtain the second sample pairs corresponding to the first sample pairs. Through the first sample pairs and the second sample pairs, a sample set is obtained to perform training processing on the image reconstruction model through the sample set.

[0059] Exemplarily, the text in the second image being the text of the preset color in the first image is taken as an example for illustration. Obtain a tobacco document with text of two colors appearing in the secondary printing and the text being overlapped, and use this tobacco document as the first document. Obtain a plurality of first documents, and perform image acquisition on the first documents to obtain the first images corresponding to the first documents. Obtain a second image that only contains text of one color, and the text information of this color is consistent with the text of this color in the first image. Obtain the first sample pair according to the above-mentioned first image and second image. For a plurality of first sample pairs, perform at least one of the following processes on the first image and the second image of each first sample pair in a paired manner: randomly cropping, horizontally or vertically rotating, or randomly transposing, etc., to obtain a second sample pair corresponding to each first sample pair. Obtain a sample set according to the first sample pair and the second sample pair.

[0060] Among them, in the process of constructing the sample set, the acquisition method of the second sample pair can be: randomly cropping the first image and the second image in the first sample pair to obtain the second sample pair; and / or, performing horizontal rotation and / or vertical rotation processing on the first image and the second image in the first sample pair and / or the second sample pair after cropping processing to obtain the second sample pair after processing; performing random transposition processing on the first image and the second image in the first sample pair and / or the second sample pair after rotation processing to obtain the second sample pair after processing.

[0061] Among them, randomly cropping the first image and the second image can be understood as randomly selecting a corresponding rectangular area in the first image and the second image, and retaining the pixels within this rectangular area to generate the second sample pair. That is, the first image in the second sample pair is the image of the rectangular area selected from the first image in the first sample pair. Correspondingly, the second image in the second sample pair is the image of the rectangular area selected from the second image in the first sample pair. The first image and the second image in the second sample pair are spatially consistent. And, the image sizes of the first image and the second image in the second sample pair are the same.

[0062] The horizontal and / or vertical rotation processing can be: performing horizontal and / or vertical rotation processing on the first image and the second image of the first sample pair to obtain the first image and the second image of the second sample pair; or performing horizontal and / or vertical rotation processing on the first image and the second image of the second sample pair after random cropping processing to obtain the second sample pair after processing. Correspondingly, it can also be performing horizontal rotation and / or vertical rotation processing on the first sample pair and / or the second sample pair after random cropping processing to obtain the second sample pair after processing.

[0063] The random transposition process may be to exchange the height and width of the image, that is, use the original width value of the image as the height value of the current image, and use the original height value of the image as the width value of the current image. The object of the random transposition process can be the first sample pair, or the second sample pair after the rotation process, or the second sample pair after the random cropping process.

[0064] Specifically, in order to obtain as many and rich second sample pairs as possible, various processes can be performed on the first sample pair. That is, perform a random cropping process on the first image and the second image in the first sample pair in a paired manner to obtain a second sample pair that is spatially consistent. And / or, perform a horizontal rotation and / or vertical rotation process on the first image and the second image of the first sample pair to obtain a second sample pair. And / or, perform a horizontal rotation and / or vertical rotation process on the first image and the second image of the second sample pair after the random cropping process to obtain a processed second sample pair. And / or, perform a random transposition process on the first image and the second image in the first sample pair to obtain a second sample pair. And / or, perform a random transposition process on the first image and the second image of the second sample pair after the rotation process to obtain a processed second sample pair. Optionally, a random transposition process can also be performed on the first image and the second image of the second sample pair after the cropping process to obtain a processed second sample pair. Through the above processes, the number of second sample pairs can be enriched, thereby improving the training accuracy of the image reconstruction model.

[0065] Exemplarily, taking the first image of the first sample pair as Figure 5 , and the second image of the first sample pair as Figure 6 as an example for illustration. Perform a random cropping process on the first image and the second image of the first sample pair in a paired manner to obtain the first image and the second image of the second sample pair. Among them, the first image and the second image of the second sample pair are spatially consistent. Horizontally rotate the first image and the second image of the second sample pair after the cropping process with a probability of 50%, and then vertically rotate the obtained image pair with a probability of 50%. Randomly transpose the rotated image pair with a probability of 50%. That is, exchange the width value and height value of the image, and use the processed image pair as the second sample pair. It should be noted that this embodiment only provides a way to determine the second sample pair. However, for the determination of the second sample pair, it can be obtained by performing any process of random cropping, horizontal rotation and / or vertical rotation process and / or random transposition process on the first sample pair and / or the second sample pair after the cropping process and / or the second sample pair after the rotation process.

[0066] S220. Build an image reconstruction model.

[0067] In an embodiment of the present invention, the image reconstruction model can be constructed as follows: obtain at least four pre-set feature extraction sub-modules, and sequentially sort at least four feature extraction sub-modules to obtain an encoding module, where the vector mapping layer in the first feature extraction sub-module of the encoding module uses a convolution kernel of the first size, and the next three sequentially connected vector mapping layers use convolution kernels of the second size; obtain at least three upsampling layers and at least three feature transformation layers pre-set, and determine a decoding module, where the upsampling layer includes a convolution sub-layer and a PixelShuffle sub-layer, and the output of the encoding module is the input of the decoding module; splice the input of the encoding module and the output of the decoding module together to obtain an image reconstruction model.

[0068] Among them, the image reconstruction model can be a network model for image reconstruction and restoration processing of images with text overlapping. The image reconstruction model includes two parts: an encoding module and a decoding module. The encoding module includes at least four feature extraction sub-modules, and each feature extraction sub-module consists of a vector mapping layer and a feature transformation layer. The vector mapping layer is used to map the image data corresponding to the target image into vector-form data. The vector mapping layer in the first feature extraction sub-module uses a convolution kernel of the first size, and the vector mapping layers in the remaining at least three feature extraction sub-modules use convolution kernels of the second size. Optionally, the first size can be 7×7, and the second size can be 3×3. The feature transformation layer of the encoding module is used to extract the global features and local features corresponding to the input image. The feature transformation layer includes: a dynamic position encoding sub-layer, a normalization sub-layer, a double dynamic token mixer sub-layer, and a multi-scale feed-forward network. Optionally, the parameters of the feature transformation layer can be set, and the parameters of the feature transformation layer can include the depth parameter of the feature transformation layer and the feature dimension parameter of the feature transformation layer. For example, if there are four feature transformation layers, the depth parameters of the four feature transformation layers can be sequentially set to [3, 3, 9, 3], and the feature dimension parameters of the feature transformation layer can be sequentially set to [48, 96, 192, 384]. The double dynamic token mixer sub-layer of the feature transformation layer realizes the extraction and fusion of local features and global features of the input feature map through an overlapping spatial reduction attention mechanism and input-dependent depth convolution. Optionally, the input-dependent depth convolution can be composed of four convolutions with convolution kernel sizes of [1, 3, 5, 7] in parallel. The number of attention heads of the overlapping spatial reduction attention mechanism is set to [1, 2, 4, 8] respectively.

[0069] The decoding module may include at least three groups of upsampling layers and feature transformation layers. The upsampling layer includes a convolutional sub-layer and a PixelShuffle sub-layer. The convolutional sub-layer may be a convolutional layer with a size of 3×3. The PixelShuffle sub-layer can convert the low-resolution feature image into a high-resolution feature image according to the mapping relationship between the high-resolution feature image and the low-resolution feature image. In other words, the PixelShuffle sub-layer is used to rearrange the pixels of the feature image to increase the resolution of the feature image. Optionally, the depth parameter and the feature dimension parameter of the feature transformation layer of the decoding module can be set. For example, if the decoding module includes 3 feature transformation layers, the depth parameters of the three feature transformation layers of the decoding module can be set to [9, 3, 3] in sequence, and the feature dimension parameters can be set to [192, 96, 96] in sequence.

[0070] Specifically, at least four pre-set feature extraction sub-modules are obtained, and at least one feature extraction sub-module is sorted sequentially to obtain an encoding module. Each feature extraction sub-module includes a vector mapping layer and a feature transformation layer. The vector mapping layer of the first feature extraction sub-module in the encoding module uses a convolutional kernel of the first size, and the vector mapping layers of at least three sequentially connected remaining feature extraction sub-modules use convolutional kernels of the second size. At least three upsampling layers and at least three feature transformation layers are obtained in advance to obtain a decoding module. The upsampling layer includes a convolutional sub-layer and a PixelShuffle sub-layer. The output of the encoding module is used as the input of the decoding module, and by splicing the input of the encoding module and the output of the decoding module together, an image reconstruction model is obtained.

[0071] It should be noted that the decoding module also includes a feature splicing layer, convolutional layers of at least two sizes, and a feature fusion layer. The feature splicing layer is used to perform feature splicing processing on the feature map output by the upsampling layer and the feature map output by the corresponding feature extraction sub-module in the encoding module. The convolutional layer is used to perform convolutional processing on the spliced features output by the feature splicing layer and the features output by the feature transformation layer. The feature fusion layer is used to fuse the input of the encoding module and the output features of the decoding module to obtain the output content corresponding to the input image.

[0072] Exemplarily, see Figure 7 , Figure 7 is a structural example diagram of the image reconstruction model. Figure 7 The PE in Figure 7 , that is, the PatchEmbedding layer, corresponds to the vector mapping layer mentioned in the above embodiment. Figure 7The UP in it, that is, the Upsampling layer, corresponds to the upsampling layer mentioned in the above embodiments. Figure 7 The 1×1 in it represents a convolutional layer with a size of 1×1. Figure 7 The 3×3 in it represents a convolutional layer with a size of 3×3.

[0073] For the encoding module of the image reconstruction model, it can be: determining the feature extraction sub-module according to the vector mapping layer and the feature transformation layer. Connecting four feature extraction sub-modules in series to obtain the encoding module. Among them, the vector mapping layer of the first feature extraction sub-module uses a 7×7 convolutional kernel, and the vector mapping layers of the remaining three feature extraction sub-modules use 3×3 convolutional kernels. The feature transformation layer consists of a dynamic positional encoding sub-layer, a normalization (Norm) sub-layer, a dual dynamic token mixer (D-Mixer) sub-layer, and a multi-scale feed-forward network (Multi-Scale Feed-Forward Network, MS-FFN). For example, see Figure 8 , the structure of the feature transformation layer can be as Figure 8 shown. Figure 8 The DPE in it, that is, Dynamic Positional Encoding, corresponds to the dynamic positional encoding sub-layer provided by the embodiments of the present invention. Figure 8 The Norm in it corresponds to the normalization sub-layer provided by the embodiments of the present invention. Figure 8 The D-Mixer in it corresponds to the dual dynamic token mixer sub-layer provided by the embodiments of the present invention. Figure 8 The MS-FFN in it corresponds to the multi-scale feed-forward network provided by the embodiments of the present invention. Optionally, if the encoding module includes four feature transformation layers, the depth parameters of the four feature transformation layers of the encoding module can be set to [3, 3, 9, 3] in sequence, and the feature dimension parameters can be set to [48, 96, 192, 384] in sequence. Among them, the dual dynamic token mixer sub-layer of the feature transformation layer is determined based on the input-dependent depthwise convolution (Input-dependent Depthwise Convolution, IDConv) and the overlapping spatial reduction attention mechanism (Overlapping Spatial Reduction Attention, OSRA) to realize the extraction of global features and local features of the input feature map. Among them, the input-dependent depthwise convolution is composed of four convolutions with convolutional kernel sizes of [1, 3, 5, 7] in parallel. The number of attention heads of the overlapping spatial reduction attention mechanism is [1, 2, 4, 8] respectively. For example, see Figure 9 , Figure 9 is an example diagram of the dual dynamic token mixer sub-layer processing the input features. Figure 9 The OSRA in it corresponds to the overlapping spatial reduction attention mechanism provided by the embodiments of the present invention. Figure 9The IDConv in it corresponds to the input-dependent depth convolution provided by the embodiments of the present invention. Figure 9 The STE in it, namely Squeezed Token Enhancer, is the compressed token enhancer of the dual-dynamic token mixer sublayer. After the features output by the overlapping spatial reduction attention mechanism and the input-dependent depth convolution are spliced, local relationship enhancement, channel compression and expansion, and residual connection are performed on the spliced features to obtain the output features of the dual-dynamic token mixer sublayer.

[0074] For the decoding module, it can be determined by at least three groups of upsampling layers and feature transformation layers. Among them, the preset upsampling layer consists of a 3×3 convolution sublayer and a PixelShuffle sublayer. The structure of the feature transformation layer is the same as that of the encoding module, but the depth parameter and feature dimension parameter of the feature transformation layer of the decoding module can be set. Among them, the depth parameters of the three feature transformation layers can be set as [9, 3, 3] in sequence, and the feature dimension parameters can be set as [192, 96, 96] in sequence.

[0075] It should be noted that the decoding module also includes a feature splicing layer, a 1×1 convolution layer, a fourth feature transformation layer, a 3×3 convolution layer, and a feature fusion layer. Among them, the feature splicing layer corresponds to Figure 7 the SkipConnection in it. The feature splicing layer is used to perform feature splicing processing on the feature map output by the upsampling layer and the feature map output by the feature extraction submodule in the corresponding encoding module. The 1×1 convolution layer is used to perform convolution processing on the spliced features output by the feature splicing layer. The fourth feature transformation layer is used to further optimize the extraction of decoding features. Optionally, the depth parameter of the fourth feature transformation layer can be 4, and the feature dimension parameter can be 96. The 3×3 convolution layer can be used to perform convolution processing on the decoding features output by the fourth feature transformation layer to obtain the output of the decoding module. The feature fusion layer is used to perform element-wise addition processing on the output feature map of the decoding module and the input image of the encoding module to obtain the target content corresponding to the input image.

[0076] Optionally, set hyperparameters and an initial learning rate for the image reconstruction model.

[0077] Among them, the hyperparameters can be understood as the parameter values set for the optimizer before training the image reconstruction model. The initial learning rate is the learning rate of the optimizer preset according to actual needs.

[0078] Optionally, the initial learning rate can be determined based on multiple dimensions. The initial learning rate can be determined based on the number of sample pairs in the sample set. When the number of sample pairs in the sample set is less than a preset number, a larger initial learning rate can be set to speed up the convergence of the model while avoiding overfitting. Correspondingly, when the number of sample pairs in the sample set is greater than a preset number, a smaller initial learning rate can be set to prevent the image reconstruction model from failing to converge due to excessively fast model parameter updates. The initial learning rate can also be determined based on the complexity of the model. The higher the model complexity and the more complex the training process, the smaller the initial learning rate can be set to avoid overfitting problems. The initial learning rate can also be determined based on the type of optimizer. Optionally, if the Adam optimizer is used during training of the image reconstruction model, the initial learning rate of the optimizer can be set to 2×10 -4 .

[0079] For example, the training process of the image reconstruction model is described using the Adam optimizer. The hyperparameters of the Adam optimizer can be set to: β 1 =0.9,β 2 =0.999, the initial learning rate can be set to 2×10 -4 .

[0080] S230: Training the constructed image reconstruction model based on the constructed sample set to obtain a usable image reconstruction model.

[0081] In an embodiment of the present invention, the training process of the image reconstruction model may specifically be: for a sample pair in a sample set, inputting the first image in the sample pair into the image reconstruction model to obtain a predicted image; performing loss processing on the second image and the predicted image based on a loss function in the image reconstruction model to obtain a loss value, and correcting the model parameters in the image reconstruction model based on the loss value; and using the model obtained when the loss function converges as a usable image reconstruction model.

[0082] The model parameters in the image reconstruction model are initial parameters or default parameters. The model parameters in the image reconstruction model are corrected by sample pairs to obtain a usable image reconstruction model. The predicted image is an image containing the edited content corresponding to the first image output after the first image is input into the image reconstruction model. The loss value can be understood as the difference between the second image and the predicted image. The loss value is determined based on a loss function. Optionally, the loss function can be a Charbonnier loss function.

[0083] Specifically, for the sample pairs in the sample set, the first image in the sample pair is input into the encoding module in the image reconstruction model, so as to generate at least four feature maps with different dimensions through at least four feature extraction sub-modules in the encoding module. Using the last feature map of the encoding module as the main input of the decoding module of the image reconstruction model, the last feature map is upsampled through the upsampling layer of the decoding module to obtain an upsampled feature. Through skip connection, the previous feature map of the last feature map is feature concatenated with the upsampled feature to obtain a concatenated feature. The concatenated feature is refined through the feature transformation layer of the decoding module to obtain a refined feature. The above process is repeated until the decoded feature is obtained through the last feature transformation layer of the decoding module. After performing image fusion processing on the decoded feature and the image information of the first image of the sample pair, a predicted image is obtained.

[0084] The second image and the predicted image are processed for loss through the loss function of the image reconstruction model to obtain a loss value, so as to correct the model parameters through the loss value. When using the loss value to correct the model parameters in the image reconstruction model, the convergence of the loss function can be used as the training objective, such as whether the training error is less than a preset error, or whether the error change tends to be stable, or whether the current number of iterations is equal to the preset number. If it is detected that the convergence condition is met, such as the training error of the loss function is less than the preset error, or the error change trend tends to be stable, it indicates that the training of the image reconstruction model is completed, and at this time, the iterative training can be stopped. If it is detected that the current convergence condition is not met, other training samples can be further obtained to continue training the image reconstruction model until the training error of the loss function is within the preset range. When the training error of the loss function reaches convergence, the trained image reconstruction model can be used as the available image reconstruction model, that is, when the target image is input into this image reconstruction model, the target content edited in the target document corresponding to the target image can be accurately obtained.

[0085] The technical solution of this embodiment improves the richness of training samples by constructing a sample set for training an image reconstruction model and obtaining as many and as rich first sample pairs and second sample pairs as possible. When the image reconstruction model is subsequently trained using the sample set, the generalization ability and accuracy of the image reconstruction model are improved. By constructing the image reconstruction model, subsequent processing of the input sample pairs by each layer of the encoding module and the decoding module of the image reconstruction model can improve the accuracy of the image reconstruction model in dealing with the problem of text overlap. Training the constructed image reconstruction model based on the constructed sample set to obtain a usable image reconstruction model facilitates subsequent processing of the target image using the usable image reconstruction model, so as to solve the problem that it is difficult to achieve ideal recognition accuracy by manual means or optical character recognition technology when there are randomly overlapping text characters in tobacco documents, and realize effective separation and reconstruction of the text content in the target document, improving the readability of the target document and the accuracy of subsequent optical character recognition technology in recognizing the target content.

[0086] Embodiment III

[0087] Figure 10 It is a schematic structural diagram of an image processing device provided in Embodiment III of the present invention. As Figure 10 shown, the device includes:

[0088] A target image acquisition module 310, configured to acquire a target image corresponding to the target document when it is detected that the content on the target document meets a preset condition; a target content determination module 320, configured to input the target image into a pre-trained image reconstruction model to obtain the target content edited in the target document; wherein, the image reconstruction model includes an encoding module and a decoding module, at least one feature extraction sub-module in the encoding module, the feature extraction sub-module includes a vector mapping layer and a feature transformation layer, the feature transformation layer includes a dynamic position encoding sub-layer, a normalization sub-layer, a double dynamic token mixer sub-layer, and a multi-scale feed-forward network, the decoding module includes an upsampling layer and a feature transformation layer, and the target content is the content edited by the user in the target document.

[0089] In the technical solution of this embodiment, when it is detected that the content on the target document meets the preset conditions, the target image corresponding to the target document is obtained. The target image is input into the pre-trained image reconstruction model to obtain the target content edited on the target document. Based on this, the effective separation of the overlapping text content in the target document is realized, and the readability of the target document and the recognition accuracy of the subsequent target content are improved. Among them, the image reconstruction model includes an encoding module and a decoding module. The encoding module includes at least one feature extraction sub-module. The feature extraction sub-module includes a vector mapping layer and a feature transformation layer. The feature transformation layer includes a dynamic position encoding sub-layer, a normalization sub-layer, a double dynamic token mixer sub-layer, and a multi-scale feed-forward network. The decoding module includes an upsampling layer and a feature transformation layer. The target content is the content edited by the user in the target document. Through the above-created image reconstruction model, the problem that it is difficult to achieve ideal recognition accuracy by manual means or optical character recognition technology in the case of random overlap of text characters in tobacco documents is solved. The text content in the target document is effectively separated and reconstructed, which not only improves the readability of the target document, but also improves the accuracy of the subsequent optical character recognition technology for recognizing the target content.

[0090] Based on the above embodiment, optionally, the device further includes: a sample set construction module, which includes: a first sample pair acquisition unit, configured to acquire a plurality of first sample pairs, where the first sample pair includes a first image and a second image, the first image is an image with overlapping text in the first document, and the second image is an image corresponding to the text edited in the first image; a second sample pair determination unit, configured to obtain a second sample pair corresponding to the first sample pair by processing the first image and the second image in the plurality of first sample pairs; a sample set determination unit, configured to determine a sample set based on the first sample pair and the second sample pair.

[0091] Optionally, the second sample pair determination unit is configured to perform random cropping on the first image and the second image in the first sample pair to obtain a second sample pair; and / or, perform horizontal rotation and / or vertical rotation on the first image and the second image in the first sample pair and / or the cropped second sample pair to obtain a processed second sample pair; perform random transposition on the first image and the second image in the first sample pair and / or the rotated second sample pair to obtain a processed second sample pair.

[0092] Optionally, the device further includes: a model construction module, which includes: an encoding module determination unit, configured to obtain at least four pre-set feature extraction sub-modules, and sequentially sort the at least four feature extraction sub-modules to obtain an encoding module, where the vector mapping layer in the first feature extraction sub-module of the encoding module uses a convolution kernel of a first size, and the three sequentially connected vector mapping layers thereafter use convolution kernels of a second size; a decoding module determination unit, configured to obtain at least three upsampling layers and at least three feature transformation layers pre-set, and determine a decoding module, where the upsampling layer includes a convolution sub-layer and a PixelShuffle sub-layer, and the output of the encoding module is the input of the decoding module; a model determination unit, configured to splice the input of the encoding module and the output of the decoding module together to obtain an image reconstruction model.

[0093] Optionally, the device further includes: a model training module, which is configured to, for a sample pair in a sample set, input the first image in the sample pair into the image reconstruction model to obtain a predicted image; perform loss processing on the second image and the predicted image based on a loss function in the image reconstruction model to obtain a loss value, so as to correct model parameters in the image reconstruction model based on the loss value; and use the model obtained when the loss function converges as an available image reconstruction model.

[0094] Optionally, set hyperparameters and an initial learning rate for the image reconstruction model, and the target document is a tobacco document.

[0095] The image processing device provided by an embodiment of the present invention can execute the image processing method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0096] Embodiment 4

[0097] Figure 11 It is a schematic structural diagram of an electronic device provided by Embodiment 4 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0098] As Figure 11As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0099] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0100] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as an image processing method.

[0101] In some embodiments, the image processing method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the image processing method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the image processing method in any other appropriate way (for example, by means of firmware).

[0102] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0103] The computer programs for implementing the image processing method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0104] Embodiment Five

[0105] Embodiment Five of the present invention also provides a computer-readable storage medium storing computer instructions for causing a processor to execute an image processing method, the method including:

[0106] When it is detected that the content on the target document meets a preset condition, obtaining a target image corresponding to the target document; inputting the target image into a pre-trained image reconstruction model to obtain the target content edited in the target document; wherein, the image reconstruction model includes an encoding module and a decoding module, the encoding module includes at least one feature extraction sub-module, the feature extraction sub-module includes a vector mapping layer and a feature transformation layer, the feature transformation layer includes a dynamic position encoding sub-layer, a normalization sub-layer, a double dynamic token mixer sub-layer, and a multi-scale feed-forward network, the decoding module includes an upsampling layer and a feature transformation layer, and the target content is the content edited by the user in the target document.

[0107] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0108] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0109] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0110] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs that run on respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0111] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and this is not limited herein.

[0112] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image processing method, characterized in that: include: When it is detected that the content on the target document meets the preset condition, acquiring a target image corresponding to the target document; Inputting the target image into a pre-trained image reconstruction model to obtain the target content edited in the target document; Among them, the image reconstruction model includes an encoding module and a decoding module, the encoding module includes at least one feature extraction submodule, the feature extraction submodule includes a vector mapping layer and a feature conversion layer, the feature conversion layer includes a dynamic position encoding sublayer, a normalization sublayer, a dual dynamic token mixer sublayer and a multi-scale feedforward network, the decoding module includes an upsampling layer and the feature conversion layer, and the target content is the content edited by the user in the target document.

2. The method according to claim 1, characterized in that The method further comprises: Constructing a sample set for training the image reconstruction model; The constructing of a sample set involved in training the image reconstruction model includes: Acquire a plurality of first sample pairs, wherein the first sample pair includes a first image and a second image, the first image is an image with overlapping text in the first document, and the second image is an image corresponding to the text edited in the first image; Obtaining a second sample pair corresponding to the first sample pair by processing the first image and the second image in the plurality of first sample pairs; The sample set is determined based on the first sample pair and the second sample pair.

3. The method according to claim 2, characterized in that The step of processing the first image and the second image in the plurality of first sample pairs to obtain a second sample pair corresponding to the first sample pair includes: Randomly cropping the first image and the second image in the first sample pair to obtain a second sample pair; and / or, Performing horizontal rotation and / or vertical rotation processing on the first sample pair and / or the first image and the second image in the second sample pair after the cropping processing to obtain a processed second sample pair; A random transposition process is performed on the first image and the second image in the first sample pair and / or the second sample pair after the rotation process to obtain a processed second sample pair.

4. The method according to claim 1, characterized in that Also includes: constructing the image reconstruction model, The constructing of the image reconstruction model comprises: Acquire at least four pre-set feature extraction submodules, and sequentially sort the at least four feature extraction submodules to obtain the encoding module, wherein the vector mapping layer in the first feature extraction submodule in the encoding module adopts a convolution kernel of a first size, and the three vector mapping layers sequentially connected thereafter adopt a convolution kernel of a second size; Acquire at least three pre-set upsampling layers and at least three of the feature conversion layers, and determine a decoding module, wherein the upsampling layer includes a convolution sublayer and a PixelShuffle sublayer, and the output of the encoding module is the input of the decoding module; The input of the encoding module and the output of the decoding module are spliced ​​together to obtain the image reconstruction model.

5. The method according to claim 1, characterized in that Also includes: The constructed image reconstruction model is trained based on the constructed sample set to obtain a usable image reconstruction model; The step of training the constructed image reconstruction model based on the constructed sample set to obtain a usable image reconstruction model includes: For a sample pair in the sample set, inputting a first image in the sample pair into the image reconstruction model to obtain a predicted image; Performing loss processing on the second image and the predicted image based on the loss function in the image reconstruction model to obtain a loss value, so as to modify the model parameters in the image reconstruction model based on the loss value; The model obtained when the loss function converges is used as the usable image reconstruction model.

6. The method according to claim 1, characterized in that Hyperparameters and an initial learning rate are set for the image reconstruction model, and the target document is a tobacco document.

7. An image processing device, characterized in that: include: A target image acquisition module, used to acquire a target image corresponding to the target document when it is detected that the content on the target document meets a preset condition; A target content determination module, used for inputting the target image into a pre-trained image reconstruction model to obtain the target content edited in the target document; Among them, the image reconstruction model includes an encoding module and a decoding module, the encoding module includes at least one feature extraction submodule, the feature extraction submodule includes a vector mapping layer and a feature conversion layer, the feature conversion layer includes a dynamic position encoding sublayer, a normalization sublayer, a dual dynamic token mixer sublayer and a multi-scale feedforward network, the decoding module includes an upsampling layer and the feature conversion layer, and the target content is the content edited by the user in the target document.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the image processing method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the image processing method according to any one of claims 1 to 6 when executed.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the image processing method according to any one of claims 1 to 6.