Document reconstruction method and device based on artificial intelligence, equipment and storage medium

Through the document reconstruction method based on artificial intelligence, safe conversion, denoising, detail enhancement and adaptive compression, the problems of large security and capacity of document reconstruction in the existing technology are solved, and safe, convenient and efficient document processing is achieved.

CN120070667APending Publication Date: 2025-05-30BEIJING THUNDERSTONE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510139301.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing document reconstruction methods are difficult to strip away potential threats from the source documents, and the refactored target documents have a large capacity, which is not conducive to subsequent processing.

Method used

Using the document reconstruction method based on artificial intelligence, the initial image is obtained through safe conversion, noise and artifacts are removed, detail enhancement processing is performed, and the compression algorithm is adaptively selected according to different objects.

Benefits of technology

It ensures the security of the document and the convenience of subsequent processing, while achieving high compression efficiency while ensuring image quality, reducing file capacity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070667A_ABST
    Figure CN120070667A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides an artificial intelligence-based document reconstruction method, which comprises the following steps of: performing security conversion on an original document to obtain an initial image; analyzing the initial image to remove noise and / or artifacts in the initial image; performing detail enhancement processing on the image with the noise and / or artifact removed based on an artificial intelligence algorithm to obtain a detail enhanced image; reconstructing the detail enhanced image into a security document; and according to different objects in the security document, adaptively selecting a corresponding compression algorithm to compress the objects in the security document. According to the technical scheme, the security document can be reconstructed, and the subsequent processing process is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, apparatus, device and storage medium for document reconstruction based on artificial intelligence. Background Art

[0002] Modern enterprises and organizations often use different systems and platforms. Especially when it comes to sharing information with external partners, suppliers or customers, the differences in document formats become an obstacle. Therefore, reconstructing internal documents into a cross-platform universal format to ensure that all platforms and devices can properly view and use this content has become an important topic in the industry.

[0003] Existing document reconstruction methods mainly use image extraction tools to directly convert a source document in a certain format (for example, word format) into a target format, such as PDF or an image. However, this simple conversion method is difficult to strip the potential threats in the source document, and the reconstructed target document often has a large capacity, which is not conducive to subsequent processing. Summary of the Invention

[0004] This application provides a method, apparatus, device and storage medium for document reconstruction based on artificial intelligence, which can reconstruct a secure document and is beneficial to the subsequent processing process.

[0005] On the one hand, this application provides a method for document reconstruction based on artificial intelligence, and the method includes:

[0006] Performing a secure conversion on the original document to obtain an initial image;

[0007] By analyzing the initial image, removing the noise and / or artifacts in the initial image;

[0008] Performing detail enhancement processing on the image from which the noise and / or artifacts have been removed based on an artificial intelligence algorithm to obtain a detail-enhanced image;

[0009] Reconstructing the detail-enhanced image into a secure document;

[0010] According to different objects in the secure document, adaptively selecting a corresponding compression algorithm to compress the objects in the secure document.

[0011] On the other hand, this application provides a device for document reconstruction based on artificial intelligence, and the device includes:

[0012] A conversion module, configured to perform a secure conversion on the original document to obtain an initial image;

[0013] A de-noising module, configured to remove the noise and / or artifacts in the initial image by analyzing the initial image;

[0014] An enhancement module for performing detail enhancement processing on an image from which noise and / or artifacts have been removed based on an artificial intelligence algorithm to obtain a detail-enhanced image;

[0015] A reconstruction module for reconstructing the detail-enhanced image into a security document;

[0016] A compression module for adaptively selecting a corresponding compression algorithm to compress the objects in the security document according to different objects in the security document.

[0017] In a third aspect, the present application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the technical solution of the above-mentioned artificial intelligence-based document reconstruction method are implemented.

[0018] In a fourth aspect, the present application provides a storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the technical solution of the above-mentioned artificial intelligence-based document reconstruction method are implemented.

[0019] As can be seen from the technical solutions provided by the present application above, on the one hand, the original document is safely converted to obtain an initial image, which can ensure the safety of the process during subsequent processing. After performing detail enhancement processing on the image from which noise and / or artifacts have been removed based on an artificial intelligence algorithm to obtain a detail-enhanced image, reconstructing the detail-enhanced image into a security document can further ensure the safety of the subsequent processing process; on the other hand, according to different objects in the security document, adaptively selecting a corresponding compression algorithm to compress the objects in the security document not only enables a high compression efficiency while ensuring the image quality, but also reduces the capacity size of the security document file, which is beneficial to subsequent processing processes such as storage and transmission. In summary, the technical solution of the present application reconstructs a security document through image processing means and is beneficial to subsequent processing processes. Description of the Drawings

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a flowchart of the artificial intelligence-based document reconstruction method provided by the embodiment of the present application;

[0022] Figure 2It is a schematic structural diagram of a document reconstruction device based on artificial intelligence provided by an embodiment of the present application;

[0023] Figure 3 It is a schematic structure of an electronic device provided by an embodiment of the present application. Specific embodiments

[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present application.

[0025] In this specification, adjectives such as first and second can only be used to distinguish one element or action from another element or action, and do not necessarily require or imply any actual such relationship or order. Where circumstances permit, reference to an element or component or step (etc.) should not be construed as being limited to only one of the elements, components, or steps, but can be one or more of the elements, components, or steps, etc.

[0026] In this specification, for the convenience of description, the dimensions of the various parts shown in the drawings are not drawn according to the actual proportional relationship. Modern enterprises and organizations often use different systems and platforms. Especially when it is necessary to share information with external partners, suppliers, or customers, document format differences become an obstacle. Therefore, reconstructing internal documents into a cross-platform universal format to ensure that all platforms and devices can view and use this content normally has become an important topic in the industry. Existing document reconstruction methods mainly use image extraction tools to directly convert a source document in a certain format (for example, word format) into a target format, such as PDF or an image. However, this simple conversion method is difficult to strip the potential threats in the source document, and the reconstructed target document often has a large capacity, which is not conducive to subsequent processing.

[0027] In view of the above problems in the prior art, the present application proposes a document reconstruction method based on artificial intelligence, and its flowchart is as shown in the appendix Figure 1 shown, mainly including steps S101 to S105, which are described in detail as follows:

[0028] Step S101: Perform a security conversion on the original document to obtain an initial image.

[0029] If there is dynamic content and malicious code in the document (such as JavaScript, macros, embedded scripts, or other executable content), and these dynamic content and malicious code rely on the specific format and interactive support of the document (such as the dynamic functions supported by Word, PDF, or HTML), they will run under specific triggering conditions, such as when the user opens the document, clicks on a link, or enables certain functions. On the one hand, since image files (such as JPEG, PNG, etc.) only contain pixel information and cannot execute dynamic scripts, embedded macros, or other code, nor can they trigger any interactions. Therefore, if the original document contains malicious code, after converting it into an image, this code only exists in the form of a static picture without a carrier or environment for execution; on the other hand, the imaged document is essentially a "screenshot" of the content, a pure presentation of text and graphics, without the potential threat code in the original document. Even if there are malicious links, executable code, etc. in the document, they are effectively isolated after imaging. Therefore, from a security perspective, based on the assumption that there is dynamic content and malicious code in the original document, the original document can be converted into an image. Safe document conversion tools such as PDF, Word, HTML, etc. can be used to render the original document, and the initial image can be obtained through safe conversion. Specifically, when opening the document, only the content of the page can be rendered without executing any embedded scripts, and at the same time, the image resolution (DPI) can be set to ensure that the generated image is clear enough to meet the requirements of OCR; if the original document is a multi-page document, each page is rendered into an image, and each page image is saved in a static format (such as PNG or JPEG) to ensure that the conversion result is an uneditable picture; if the original document is pure text, grayscale images can be selected to reduce the file size, and if the original document is a document with many graphics and markings, the image can be saved in a color format (such as RGB) to maintain the integrity of the original content, and so on.

[0030] Step S102: Analyze the initial image to remove noise and / or artifacts from the initial image.

[0031] In the field of image processing, noise refers to random interference in the image that has nothing to do with the actual content, such as scattered dots, random black and white pixels, lines, etc. These usually appear during the processes of scanning, compression, and resolution mismatch, and may be caused by factors such as the compression algorithm during image conversion, the quality of the scanner, or resolution mismatch. Artifacts generally refer to false information or distortion introduced during the image processing process, such as JPEG compression artifacts, scanning artifacts, edge blurring, etc. These may cause abnormal display of the document content, affecting the clarity and readability of the content. Whether it is noise or artifacts, they essentially belong to some visual or data "impurities" generated during the process of processing the document into an image. These "impurities" not only affect the visual quality of the image but may also contain potential security risks. Therefore, it is necessary to analyze the initial image to remove noise and / or artifacts from the initial image.

[0032] As an embodiment of the present application, by analyzing the initial image, removing the noise in the initial image can be achieved through steps Sa1021 to Sa1024, which are described in detail as follows:

[0033] Step Sa1021: Calculate the local pixel mean and local standard deviation of the initial image.

[0034] Specifically, a sliding window (e.g., a window of size k*k) can be selected in the initial image to calculate the mean and standard deviation of the local area. For a pixel in a window, assuming that the pixel value of the area is I(x', y'), calculate the local mean μ w and the local standard deviation σ w :

[0035]

[0036] Where I(x',y') is the pixel value in window w, μ w is the mean value of the pixels in the window, σ w is the standard deviation of pixels within the window, k 2 is the number of pixels of the window.

[0037] Step Sa1022: Analyze the value of each pixel in the initial image.

[0038] Step Sa1023: If there is a pixel P in the initial image whose average value differs from that of its local area by more than a preset threshold noise , then pixel P noise Determined to be noise.

[0039] In some noise types (such as Gaussian noise), noise is often a manifestation of the deviation of pixel values ​​in a local area from the local mean. Therefore, when the value of a pixel differs from the mean of its local area by more than a certain threshold, it may be noise. A noise detection threshold can be set. If a pixel I(x,y) meets the following conditions, the pixel is considered to be noise:

[0040] |I(x,y)-μ w |>τ*σ w

[0041] Here, τ is the noise detection threshold, which can be adjusted experimentally and is usually between 2 and 3.

[0042] Step Sa1024: Remove noise in the initial image by neighborhood averaging or interpolation.

[0043] For example, replace a noise pixel with the mean of its neighborhood:

[0044]

[0045] As another embodiment of the present application, by analyzing the initial image, removing the noise in the initial image can be achieved through steps Sb1021 to Sb1024, and the detailed description is as follows:

[0046] Step Sb1021: Input the initial image into the trained neural network, where the trained neural network includes an encoder, a residual connection layer, and a decoder.

[0047] In the embodiment of the present application, the trained neural network is trained based on a deep learning-based algorithm. Specifically, training data sets such as noise-free images and noisy images are input (the noisy image is used as the input for neural network training, and the noise-free image is used as the target label), and functions such as the mean squared error (MSE) function are used as the loss function. The network parameters are optimized through continuous backpropagation to train the neural network until the loss function value is minimized and the training ends. The encoder is used to extract the features of the image, the decoder is mainly used to recover the noise-free image from the features, and the residual connection layer is located between the encoder and the decoder.

[0048] Step Sb1022: Extract the features of the initial image through the encoder to obtain a feature map containing multiple layers of features.

[0049] The encoder consists of multiple convolutional layers, pooling layers, activation functions, etc. These convolutional layers not only extract the local features of the image but also gradually reduce the spatial resolution through pooling operations, mapping the image from a high-dimensional space to a low-dimensional latent space. Through the stacking of multiple convolutional layers and pooling layers, the spatial resolution of the input image gradually decreases, the size of the feature map gradually shrinks, and the features extracted by each layer become more and more abstract, gradually converting the low-level pixel information into high-level semantic features. This means that the output from the encoder can be a feature map containing multiple layers of features, that is, as the convolutional layers and pooling layers of the encoder deepen (meaning getting farther from the input layer and closer to the output layer), the level of the obtained feature map becomes higher, more abstract, or more semantic, that is, a feature map of high-level features is obtained. As the features become more abstract or more semantic, details such as small random dots, spots, or irregular brightness changes in the image become more obvious, and the network is more likely to recognize the noise characterized by these details and thus remove it. This is the mechanism of noise removal by the trained neural network.

[0050] Step Sb1023: Transmit the low-dimensional features of the feature map to the decoder through the residual connection layer.

[0051] Different from traditional neural networks, the initial image is calculated layer by layer through multiple convolutional layers, pooling layers, etc. in the encoder, and finally, after obtaining high-level features, it is input to the decoder for decoding. In the embodiments of the present application, the encoder can transfer the low-dimensional features of the feature map to the decoder through a residual connection layer, and the residual connection layer can directly transfer the shallow features obtained after the encoder processes the initial image to the decoder. This means that in the embodiments of the present application, the decoder input includes not only the relatively abstract and semantically obvious high-level features output from the last layer of the encoder, but also the low-dimensional features containing more image details output from the residual connection layer.

[0052] Step Sb1024: The decoder fuses the high-level features calculated by the encoder with the low-dimensional features transferred from the encoder to obtain an image with noise removed.

[0053] If, like traditional neural networks, the decoder directly performs upsampling on the high-level features output from the encoder to restore the image, the details of the image, such as edges and textures, will be lost. To enhance the detail information, in the embodiments of the present application, the decoder fuses the high-level features calculated by the encoder with the low-dimensional features transferred from the encoder to obtain an image with noise removed. At this time, the image not only removes noise but also retains detail information such as edges and textures.

[0054] As an embodiment of the present application, by analyzing the initial image, removing the artifacts in the initial image can be achieved through steps Sc1021 to Sc1024, and the detailed description is as follows:

[0055] Step Sc1021: By analyzing the structural similarity index of the initial image, determine the artifact regions in the initial image.

[0056] Artifacts usually cause obvious differences in the structural information of some regions of the image from its reference image or from its surrounding regions. The artifacts can be identified by calculating the structural similarity index (SSIM) of each local region. When the SSIM of a certain region is low, it indicates that the structural difference between this region and the surrounding region or the reference image is large, and it may be an artifact. Specifically, a threshold can be set. If the SSIM of a certain region is less than this threshold, then it is considered that this region contains artifacts.

[0057] Step Sc1022: Extract the texture features of the artifact regions and non-artifact regions in the initial image.

[0058] For each local region in the initial image, texture features such as gradient direction, edge information, and texture frequency can be extracted through algorithms such as wavelet transform, including the texture features of the artifact regions and non-artifact regions.

[0059] Step Sc1023: Match non-artifact regions in the initial image whose texture features are similar to those of the artifact regions.

[0060] Based on measurement methods such as mutual information or cosine similarity, by calculating the similarity between two texture blocks, if the similarity between the two exceeds a preset threshold, it can be determined that the two texture blocks have similar textures, thereby matching non-artifact regions in the initial image whose texture features are similar to those of the artifact regions.

[0061] Step Sc1024: According to the matching result of the texture features, fill the texture blocks of the non-artifact regions in the initial image whose texture features are similar to those of the artifact regions into the artifact regions.

[0062] If, through the matching in Step Sc1023, there are texture blocks in the non-artifact regions whose texture features are similar to those of the artifact regions, then these texture blocks can be filled into the artifact regions, so that the image can smoothly transition to the artifact regions, ultimately achieving the removal or repair of the artifacts.

[0063] Step S103: Based on an artificial intelligence algorithm, perform detail enhancement processing on the image from which noise and / or artifacts have been removed to obtain a detail-enhanced image.

[0064] Through the above embodiments of removing noise and / or artifacts in the initial image, the noise removal or artifact removal solution will inevitably or to varying degrees have some impacts on the initial image. For example, the details of the image are lost. To solve this technical problem, in the embodiments of the present application, an artificial intelligence algorithm can be used to perform detail enhancement processing on the image from which noise and / or artifacts have been removed to obtain a detail-enhanced image. Specifically, it can be implemented through Steps S1031 to S1034, and the details are described as follows:

[0065] Step S1031: Extract different-scale features of the image from which noise and / or artifacts have been removed through a multi-scale convolutional neural network.

[0066] In the field of image processing, different scale features of an image characterize different levels of detail in the image. The scale of "scale feature" generally refers to the level of detail in the image or the resolution level of the image, including spatial scale and resolution scale. Among them, the spatial scale refers to the level of detail in the image, and different scales represent different sizes of structures. For example, large-scale structures in the image (such as buildings) usually correspond to lower scales (i.e., large scales), while fine details (such as textures, edges) correspond to higher scales (i.e., small scales). The resolution scale is the different resolution levels of the image. A higher scale of a feature means a finer-grained feature, while a lower scale means a coarser feature. By using multi-scale feature extraction, the overall picture of the image can be understood and captured from different resolution levels of the image. According to the meaning of scale features, in the embodiments of the present application, different scale features of an image from which noise and / or artifacts have been removed are extracted through a multi-scale convolutional neural network. The specific implementation can be achieved by using convolutional kernels of different sizes, upsampling and downsampling, or a multi-scale feature pyramid, etc. Among them, using convolutional kernels of different sizes means capturing features in the image from which noise and / or artifacts have been removed at different scales by using convolutional kernels of different sizes (such as 3×3, 5×5, 7×7, etc.). Smaller convolutional kernels (such as 3×3) usually capture details (such as textures or edges), while larger convolutional kernels (such as 5×5 or 7×7) can capture larger structures (such as the overall shape of an object). Through upsampling and downsampling, pooling operations (such as max pooling, average pooling) are used to reduce the resolution of the image from which noise and / or artifacts have been removed to obtain a low-resolution feature map; then upsampling operations (such as deconvolution or interpolation) are used to increase the resolution of the image from which noise and / or artifacts have been removed to restore more details. The multi-scale feature pyramid is to build a feature pyramid from high resolution to low resolution to process information at different scales, so as to better capture various details of the image from which noise and / or artifacts have been removed.

[0067] Step S1032: Based on different scale features of the image from which noise and / or artifacts have been removed, combine spatial and channel attention mechanisms to obtain a spatial attention map and a channel attention map of the image from which noise and / or artifacts have been removed.

[0068] On the basis of obtaining different scale features of the image from which noise and / or artifacts have been removed, by introducing spatial and channel attention mechanisms, focus on the detail areas and suppress irrelevant information to further enhance the details of the image. Specifically, to obtain the spatial attention map of the image from which noise and / or artifacts have been removed by combining the spatial attention mechanism based on different scale features of the image from which noise and / or artifacts have been removed can be: perform max pooling or average pooling on all feature maps (i.e., different scale features of the image from which noise and / or artifacts have been removed) in the spatial dimension (H×W) to obtain a new feature map, that is, the fused feature map F spatial, representing spatial information at all scales; then, for this fused feature map F spatial Apply a convolution operation and activation function A spatial = sigmoid(W spatial ·F spatial ) to learn the importance at each position, i.e., the spatial attention map A spatial , where, W spatial is the learned convolution kernel for generating the spatial attention map, which is related to each spatial position, and the output is a spatial attention map A spatial of size H×W. Based on the different-scale features of the image with noise and / or artifacts removed, combined with the channel attention mechanism, to obtain the channel attention map of the image with noise and / or artifacts removed, specifically: perform global average pooling and global max pooling on any scale feature of the image with noise and / or artifacts removed (i.e., the image with noise and / or artifacts removed for any scale feature) on any channel c, respectively, to obtain two channel feature vectors F avg and F max on each channel; use a feed-forward neural network to perform the following operations on and:

[0069] M channel = σ(W 2 ·δ(W 1 ·F avg ) + W 2 ·δ(W 1 ·F max ))

[0070] where, and are the parameters of two fully connected layers of the feed-forward neural network, r is the channel dimensionality reduction coefficient (usually taken as r = 16) for reducing the computational amount, δ is the RelU activation function, and σ is the sigmoid activation function for normalizing the attention value to [0, 1]; finally, perform the following calculation on M channel through the sigmoid activation function A channel = σ(M channel ) to obtain the channel attention map A channel .

[0071] Step S1033: Apply the spatial attention map and the channel attention map to the different-scale features of the image with noise and / or artifacts removed to obtain an intermediate enhanced feature map.

[0072] Specifically, the spatial attention map can be multiplied by the different-scale features of the image with noise and / or artifacts removed to obtain a spatially attention-weighted feature map, and then, the spatially attention-weighted feature map is multiplied by the channel attention map, and the product is used as the intermediate enhanced feature map.

[0073] Step S1034: Input the intermediate enhanced feature map into the pre-trained generative adversarial network, and the pre-trained generative adversarial network outputs a detail-enhanced image.

[0074] In the embodiment of the present application, the pre-trained generative adversarial network is obtained by inputting the image features enhanced by multi-scale feature extraction and attention mechanism into the generative adversarial network, performing adversarial training on the generative adversarial network, and optimizing the parameters of its generator and discriminator to the optimal values. After the generative adversarial network training is completed, input the intermediate enhanced feature map into the pre-trained generative adversarial network, and the pre-trained generative adversarial network outputs a detail-enhanced image.

[0075] Step S104: Reconstruct the detail-enhanced image into a security document.

[0076] Image processing techniques can be used, including optical character recognition (OCR) or image segmentation models, to segment the characters in the detail-enhanced image to obtain independent character individuals; then, identify the independent character individuals by extracting the visual features of the independent character individuals to identify the text content in the image; finally, use document generation tools for page layout and typesetting, including layout reconstruction, adaptive typesetting, and style adjustment, etc., to achieve precise format control, and reconstruct the security document according to the parsed layout information such as headings, paragraphs, tables, etc., for example, a document in PDF format. Among them, layout reconstruction aims to recombine the text and image content within each page, maintain the relative positions of the content, and ensure the consistency of typesetting, while adaptive typesetting automatically adjusts the positions and sizes of text blocks and images according to the page width and height to adapt to the standard size of the security document page and ensure that the layout of the content on each page does not shift; as for style adjustment, it mainly unifies the style attributes such as font, color, and spacing to ensure that the visual effect of the generated security document page is neat and standardized. To ensure visual consistency, keep the format parameters such as font, font size, line spacing, and paragraph spacing of the security document similar to those of the original document. During the reconstruction process, completely strip dynamic content such as scripts and macros to ensure that the security document is pure static content. For complex tables or graphics, use rasterization methods or drawing tools to regenerate static versions without dynamic content.

[0077] Step S105: According to different objects in the security document, adaptively select the corresponding compression algorithm to compress the objects in the security document.

[0078] Specifically, by extracting features from different objects in the security document, different objects in the security document can be classified into text, graphics, photos, and background areas; lossless compression, lossy compression, multi-level compression, and lossy compression with a high compression ratio are respectively used to compress the text, graphics, photos, and background areas. It should be noted here that due to the different characteristics of the text, graphics, photos, and background areas, different compression methods need to be adopted. Among them, the text information is crucial for users and must be ensured to be completely restored. Therefore, a lossless compression method (such as ZIP or LZW, etc.) is used for the text to ensure the integrity of the text; the graphics area usually has a high requirement for retaining details and a relatively high tolerance for colors. Therefore, lossy compression (such as JPEG, etc.) can be used to compress the graphics area to reduce the file size; the photo (or image) area usually has rich colors and complex details, and multi-level compression such as JPEG2000 provides higher compression efficiency and more flexible quality adjustment; the background is usually relatively simple and has a uniform color. Using a high compression ratio can effectively reduce the file size without affecting the overall visual effect. Therefore, lossy compression with a high compression ratio can be used to compress the background area.

[0079] From the above-attached Figure 1 As can be seen from the above-described example of the artificial intelligence-based document reconstruction method, on the one hand, the original document is securely converted to obtain an initial image, which can ensure the safety of the process during subsequent processing. After the detail-enhanced image is obtained by performing detail enhancement processing on the image from which noise and / or artifacts have been removed based on an artificial intelligence algorithm, reconstructing the detail-enhanced image into a security document can further ensure the safety of the subsequent processing process; on the other hand, according to different objects in the security document, the corresponding compression algorithm is adaptively selected to compress the objects in the security document. This not only achieves a high compression efficiency while ensuring the image quality, but also reduces the capacity size of the security document file, which is beneficial to subsequent storage, transmission, and other processing processes. In summary, the technical solution of the present application reconstructs the security document through image processing means and is beneficial to the subsequent processing process.

[0080] Please refer to the attached Figure 2 , which is an artificial intelligence-based document reconstruction device provided by an embodiment of the present application. The device may include a conversion module 201, a de-noising module 202, an enhancement module 203, a reconstruction module 204, and a compression module 205, which are described in detail as follows:

[0081] The conversion module 201 is configured to securely convert the original document to obtain an initial image;

[0082] The de-noising module 202 is configured to remove noise and / or artifacts in the initial image by analyzing the initial image;

[0083] An enhancement module 203, configured to perform detail enhancement processing on the image from which noise and / or artifacts have been removed based on an artificial intelligence algorithm to obtain a detail-enhanced image;

[0084] A reconstruction module 204, configured to reconstruct the detail-enhanced image into a security document;

[0085] A compression module 205, configured to adaptively select a corresponding compression algorithm according to different objects in the security document to compress the objects in the security document.

[0086] From the above-described artificial intelligence-based document reconstruction device of the example, on the one hand, performing security conversion on the original document to obtain an initial image can ensure the safety of the process during subsequent processing. After performing detail enhancement processing on the image from which noise and / or artifacts have been removed based on an artificial intelligence algorithm to obtain a detail-enhanced image, reconstructing the detail-enhanced image into a security document can further ensure the safety of the subsequent processing process; on the other hand, adaptively selecting a corresponding compression algorithm according to different objects in the security document to compress the objects in the security document not only enables a high compression efficiency while ensuring the image quality, but also reduces the capacity size of the security document file, which is beneficial for subsequent processing processes such as storage and transmission. In summary, the technical solution of the present application reconstructs a security document through image processing means and is beneficial for subsequent processing processes. Figure 2

[0087] Figure 3 Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 1 shown, the electronic device 3 of this embodiment mainly includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a program for an artificial intelligence-based document reconstruction method. When the processor 30 executes the computer program 32, the steps in the embodiment of the above-described artificial intelligence-based document reconstruction method are implemented, such as Figure 2 shown steps S101 to S105. Alternatively, when the processor 30 executes the computer program 32, the functions of each module / unit in the above-described device embodiments are implemented, such as shown functions of the conversion module 201, the noise removal module 202, the enhancement module 203, the reconstruction module 204, and the compression module 205.

[0088] ​Exemplarily, the computer program 32 for the artificial intelligence-based document reconstruction method mainly includes: securely converting the original document to obtain an initial image; analyzing the initial image to remove noise and / or artifacts in the initial image; performing detail enhancement processing on the image from which noise and / or artifacts have been removed based on an artificial intelligence algorithm to obtain a detail-enhanced image; reconstructing the detail-enhanced image into a secure document; and adaptively selecting a corresponding compression algorithm to compress the objects in the secure document according to different objects in the secure document. The computer program 32 can be divided into one or more modules / units, and one or more modules / units are stored in the memory 31 and executed by the processor 30 to complete the present application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 32 in the electronic device 3. For example, the computer program 32 can be divided into the functions of a conversion module 201, a noise removal module 202, an enhancement module 203, a reconstruction module 204, and a compression module 205 (modules in the virtual device). The specific functions of each module are as follows: The conversion module 201 is used to securely convert the original document to obtain an initial image; the noise removal module 202 is used to analyze the initial image to remove noise and / or artifacts in the initial image; the enhancement module 203 is used to perform detail enhancement processing on the image from which noise and / or artifacts have been removed based on an artificial intelligence algorithm to obtain a detail-enhanced image; the reconstruction module 204 is used to reconstruct the detail-enhanced image into a secure document; and the compression module 205 is used to adaptively select a corresponding compression algorithm to compress the objects in the secure document according to different objects in the secure document.

[0089] The electronic device 3 may include but is not limited to the processor 30 and the memory 31. Those skilled in the art can understand that Figure 3 merely examples of the electronic device 3, which do not constitute a limitation on the electronic device 3, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.

[0090] The so-called processor 30 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0091] The memory 31 may be an internal storage unit of the electronic device 3, such as the hard disk or memory of the electronic device 3. The memory 31 may also be an external storage device of the electronic device 3, such as a plug-in hard disk equipped on the electronic device 3, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 31 may also include both the internal storage unit and the external storage device of the electronic device 3. The memory 31 is used to store computer programs and other programs and data required by the electronic device. The memory 31 may also be used to temporarily store data that has been output or will be output.

[0092] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0093] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0094] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0095] In the embodiments provided in this application, it should be understood that the disclosed apparatus / devices and methods can be implemented in other ways. For example, the apparatus / device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0096] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0097] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0098] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by a computer program instructing relevant hardware. The computer program for the method of document reconstruction based on artificial intelligence can be stored in a storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments, that is, perform a secure conversion on the original document to obtain an initial image; analyze the initial image to remove noise and / or artifacts in the initial image; perform detail enhancement processing on the image from which noise and / or artifacts have been removed based on an artificial intelligence algorithm to obtain a detail-enhanced image; reconstruct the detail-enhanced image into a secure document; adaptively select a corresponding compression algorithm according to different objects in the secure document to compress the objects in the secure document. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The storage medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the storage medium does not include electrical carrier signals and telecommunication signals.

[0099] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application, and should all be included in the protection scope of the present application. The specific implementation manners described above further elaborate on the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above is only the specific implementation manners of the present application, and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application should all be included in the protection scope of the present invention.

Claims

1. A document reconstruction method based on artificial intelligence, characterized in that: The method comprises: The original document is securely converted to obtain an initial image; By analyzing the initial image, removing noise and / or artifacts in the initial image; Performing detail enhancement processing on the image from which noise and / or artifacts have been removed based on an artificial intelligence algorithm to obtain a detail enhanced image; reconstructing the detail-enhanced image into a security document; According to different objects in the security document, a corresponding compression algorithm is adaptively selected to compress the objects in the security document.

2. The document reconstruction method based on artificial intelligence as claimed in claim 1, characterized in that: The removing noise in the initial image by analyzing the initial image comprises: Calculating the local pixel mean and local standard deviation of the initial image; Analyzing the value of each pixel in the initial image; If there is a pixel P in the initial image whose average value differs from that of its local area by more than a preset threshold noise , then the pixel P noise Determined as noise; The noise in the initial image is removed by neighborhood averaging or interpolation.

3. The document reconstruction method based on artificial intelligence as claimed in claim 1, characterized in that: The removing noise in the initial image by analyzing the initial image comprises: Inputting the initial image into a trained neural network, wherein the trained neural network comprises an encoder, a residual connection layer and a decoder; Extracting features of the initial image by the encoder to obtain a feature map containing multiple layers of features; Passing the low-dimensional features of the feature map to the decoder through the residual connection layer; The decoder fuses the high-level features calculated by the encoder with the low-dimensional features to obtain the image from which the noise has been removed.

4. The document reconstruction method based on artificial intelligence as claimed in claim 1, characterized in that: The removing artifacts in the initial image by analyzing the initial image includes: Determining an artifact region in the initial image by analyzing a structural similarity index of the initial image; Extracting texture features of artifact areas and non-artifact areas in the initial image; Matching a non-artifact area in the initial image whose texture features are similar to those of the artifact area; According to the matching result of the texture features, the texture blocks of the non-artifact area in the initial image that are similar to the texture features of the artifact area are filled into the artifact area.

5. The document reconstruction method based on artificial intelligence as claimed in claim 1, characterized in that: The method of performing detail enhancement processing on the image from which noise and / or artifacts have been removed based on an artificial intelligence algorithm to obtain a detail enhanced image includes: Extracting features of different scales of the image from which noise and / or artifacts have been removed by a multi-scale convolutional neural network; Based on different scale features of the image from which noise and / or artifacts have been removed, combining spatial and channel attention mechanisms to obtain a spatial attention map and a channel attention map of the image from which noise and / or artifacts have been removed; Applying the spatial attention map and the channel attention map to different scale features of the image from which noise and / or artifacts have been removed to obtain an intermediate enhanced feature map; The final enhanced feature map is input into a trained generative adversarial network, and the trained generative adversarial network outputs the detail enhanced image.

6. The document reconstruction method based on artificial intelligence as claimed in claim 5, characterized in that: The step of applying the spatial attention map and the channel attention map to different scale features of the image from which noise and / or artifacts have been removed to obtain an intermediate enhanced feature map comprises: Multiplying the spatial attention map with the different scale features of the image from which noise and / or artifacts have been removed to obtain a spatial attention weighted feature map; The spatial attention weighted feature map is multiplied by the channel attention map, and the product is used as the intermediate enhanced feature map.

7. The document reconstruction method based on artificial intelligence as claimed in claim 1, characterized in that: The adaptively selecting a corresponding compression algorithm to compress the objects in the security document according to different objects in the security document includes: By performing feature extraction on different objects in the security document, different objects in the security document are divided into text, graphics, photos and background areas; The text, graphics, photos and background areas are compressed using lossless compression, lossy compression, multi-level compression and lossy compression with a high compression ratio.

8. A document reconstruction device based on artificial intelligence, characterized in that: The device comprises: A conversion module, used for securely converting the original document to obtain an initial image; a de-noising module, configured to remove noise and / or artifacts in the initial image by analyzing the initial image; An enhancement module, used for performing detail enhancement processing on the image from which noise and / or artifacts have been removed based on an artificial intelligence algorithm to obtain a detail enhanced image; a reconstruction module, used for reconstructing the detail-enhanced image into a security document; The compression module is used to adaptively select a corresponding compression algorithm to compress the objects in the security document according to different objects in the security document.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.