High-fidelity cultural relic digital image generation and restoration system and method

By adopting new cross attention modules and cross-window attention methods in the image generation and repair system, the image blur and artifact problems are solved, the image quality is improved, and it is suitable for the fields of cultural relics protection, restoration and cultural relics digitalization, and has performance advantages.

CN120147452AInactive Publication Date: 2025-06-13BEIJING GUANGAN LIGHTING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510216519.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems of image blurring and artifacts in image generation and repair, especially in the fields of cultural relics protection, restoration and cultural relics digitization, which are intolerable.

Method used

A high-fidelity cultural relics digital image generation and repair system is proposed, and a new cross-attention module and cross-window attention method is adopted to enable it to calculate cross-attention on a larger scale and eliminate interference from unrelated information such as image background.

Benefits of technology

Through this system, the problem of image blur and artifact can be effectively solved, image quality can be improved, and it is suitable for use in the fields of cultural relics protection, restoration and cultural relics digitalization, and has performance advantages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147452A_ABST
    Figure CN120147452A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of artificial intelligence, computer vision and computer graphics, and discloses a high-fidelity cultural relic digital image generation and restoration system and method, and the method comprises the steps: extracting the shallow features of an input image; according to the shallow layer features, deep layer features of the input image are extracted; and reconstructing and repairing the image according to the deep features, and outputting an image with quadruple resolution. The invention provides a new cross attention module and a new cross-window attention method, so that the cross attention can be calculated in a larger range, the interference of irrelevant information such as an image background is eliminated, the problems of image blurring and artifacts can be solved, and the method has performance advantages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of artificial intelligence, computer vision, and computer graphics, and particularly relates to an image generation and restoration system and method for high-fidelity cultural relic digitization. Background Art

[0002] In the fields of cultural relic protection, restoration, and cultural relic digitization, the high-fidelity restoration of cultural relics is related to the degree of cultural relic digitization, the restoration degree of cultural relic reconstruction, the effectiveness of cultural relic restoration, etc. Therefore, the generation and restoration of high-fidelity images are crucial.

[0003] In the field of image generation and restoration, the standard Transformer method needs to calculate cross-attention with all pixels of the image, which will bring a huge performance overhead. The current most advanced method is an improved method based on Transformer, which can calculate cross-attention within a smaller pixel range. This not only ensures the effect but also ensures the efficiency. However, precisely because the pixel range it uses is reduced, the probability that it can learn incorrect information is greatly increased, and the window movement mechanism will lead to a reduction in the efficiency of cross-window connections, and information between non-adjacent windows cannot be interconnected. Eventually, when constructing the restoration of high-fidelity images, the situation of image blurring or artifacts will occur.

[0004] In the fields of cultural relic protection, restoration, and cultural relic digitization, image blurring or artifacts are intolerable. Summary of the Invention

[0005] To solve the problems existing in the prior art, the present invention provides an image generation and restoration system and method for high-fidelity cultural relic digitization, proposes a new "cross-attention module", and a new cross-window attention method, enabling it to calculate cross-attention within a larger range and excluding the interference of irrelevant information such as the image background, which can solve the problems of image blurring and artifacts and has performance advantages.

[0006] To achieve the above object, the present invention provides the following solutions:

[0007] An image generation and restoration system for high-fidelity cultural relic digitization, the system includes: a first feature extraction stage, a second feature extraction stage, and an image generation and restoration stage;

[0008] The first feature extraction stage is used to extract the shallow features of the input image;

[0009] The second feature extraction stage is used to extract the deep features of the input image according to the shallow features;

[0010] The image generation and restoration stage is used to reconstruct and restore the image according to the deep features and output an image with four times the resolution.

[0011] Preferably, the first feature extraction stage includes: an input image unit, a residual network, and a first convolutional layer;

[0012] The input image unit is used to input an image;

[0013] The residual network is used to extract shallow features of the input image;

[0014] The first convolutional layer is used to downsample the shallow features and also serves as the output head for residual network feature extraction.

[0015] Preferably, the second feature extraction stage is composed of 4 identical hybrid attention modules. The output feature map size, number of channels, and number of module iterations of each hybrid attention module are different. Each hybrid attention module is respectively composed of a cross-attention module, a cross-window attention module, and two convolutional layers.

[0016] Preferably, the cross-attention module is composed of two modules. The first module iterates 2 times, and the second module uses a third-channel attention module to process cross-window information interaction. Among them, the first module uses a first-channel attention module, a second-channel attention module, and a window-based multi-head self-attention module in parallel, allowing the three modules to extract features respectively and learn different information.

[0017] Preferably, the first-channel attention module has three stages. One is the feature extraction stage, which is composed of a second convolutional layer, a GELU activation function, and a third convolutional layer, and its function is to further extract features; one is the channel attention stage, which is composed of global pooling, a fourth convolutional layer, a GLUE activation function, a fifth convolutional layer, a Sigmoid function, and a skip connection with weight coefficients, and its function is to perform a simple attention calculation on the channels; the last one is the output stage, which is composed of a sixth convolutional layer and is used to adjust the size and number of channels for subsequent processing.

[0018] Preferably, the second-channel attention module is the squeeze-and-excitation attention mechanism SENet, that is, Squeeze-and-Excitation Networks.

[0019] Preferably, the calculation process of the window-based multi-head self-attention module includes:

[0020] Given the input feature map as Split into H cross ×W cross / M 2 windows of size M×M, and then calculate self-attention within each small window. For each window feature map The cross-attention mechanism is expressed as:

[0021]

[0022] Among them, d represents the dimension of the query matrix Q, K represents the key matrix, V represents the value matrix, and B represents the relative position encoding;

[0023] Adopt the method of gradually expanding the window size, divide the window size into three scale levels of M×M, 2M×2M and 4M×4M, then perform cross-attention inside the window respectively, and then for the distance of each window movement

[0024] Preferably, the calculation process of the third-channel attention module includes:

[0025] For the Q, K, and V feature maps of the input feature X, they are respectively

[0026] For X Q Divide it into M 1 ×M 1 non-overlapping windows of size, and for X K and X V Divide it into M 2 ×M 2 overlapping windows of size, where M 2 =(1 + γ)×M 1 , γ is a constant controlling the overlapping size; the cross-attention of these three Q, K, and V still uses the formula of Attention(Q, K, V), and the relative position variation

[0027] The present invention also provides an image generation and restoration method for high-fidelity cultural relic digitization, which is implemented by using the high-fidelity cultural relic digitization image generation and restoration system described in any one of the above, and the method includes:

[0028] Extract the shallow features of the input image;

[0029] According to the shallow features, extract the deep features of the input image;

[0030] According to the deep features, reconstruct and repair the image, and output an image with four times the resolution.

[0031] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method described above is implemented.

[0032] Compared with the prior art, the beneficial effects of the present invention are:

[0033] Through an innovative method, the present invention can solve the problems of image blurring and artifacts, making the image quality higher, and is suitable for the fields of cultural relic protection, restoration, and digitalization of cultural relics.

[0034] Compared with the standard Transformer, the present invention not only has better effects but also better performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions of the present invention, the following briefly introduces the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0036] Figure 1 It is a schematic structural diagram of an image generation and restoration system for high-fidelity digitalization of cultural relics according to an embodiment of the present invention;

[0037] Figure 2 It is an architecture diagram of a cross-attention module according to an embodiment of the present invention;

[0038] Figure 3 It is an architecture diagram of a first-channel attention module according to an embodiment of the present invention;

[0039] Figure 4 It is a schematic structural diagram of an electronic device according to an embodiment of the present invention.

[0040] DESCRIPTION OF REFERENCE NUMERALS:

[0041] 1010, processor; 1020, memory; 1030, input / output interface; 1040, communication interface; 1050, bus. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0043] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0044] Embodiment 1

[0045] As Figure 1As shown in the figure, the present invention discloses an image generation and restoration system for high-fidelity cultural relic digitization. The system includes: a first feature extraction stage, a second feature extraction stage, and an image generation and restoration stage;

[0046] The first feature extraction stage is used to extract the shallow features of the input image;

[0047] The second feature extraction stage is used to extract the deep features of the input image according to the shallow features;

[0048] The image generation and restoration stage is used to reconstruct and restore the image according to the deep features and output an image with four times the resolution.

[0049] In this embodiment, the first feature extraction stage includes: an input image unit, a residual network, and a first convolutional layer;

[0050] (1) Input image That is, the height and width of the input image are H×W, and the number of RGB channels is 3.

[0051] (2) The residual network is used to extract the shallow features of the input image That is, after passing through the residual network, the height and width dimensions of the feature map are still H×W, and the number of channels is C (usually C>3).

[0052] (3) The first function of the "first convolutional layer" is downsampling, and the output feature is The second function is the output head of the residual network feature extraction, preparing for the input of the subsequent network. All convolutional layers in this method are a relatively simple CNN structure, and these convolutional layers are composed of the following structure:

[0053] 1) First, a padding operation will be performed. Since the image size H×W may not be divisible by 4, a padding is done around the image (if it is divisible by 4, no padding operation is required).

[0054] 2) Then it is a CNN with a 4x4 convolutional kernel and a stride of 4.

[0055] 3) Finally, a layer normalization is connected.

[0056] The above three parts are connected in series to form a convolutional layer. The function of the convolutional layer is either upsampling or downsampling. In short, it is to adjust the image size and number of channels to adapt to the input of the subsequent network.

[0057] In this embodiment, (1) the second feature extraction stage is composed of 4 hybrid attention modules with the same structure, namely the "first hybrid attention module", the "second hybrid attention module", the "third hybrid attention module", and the "fourth hybrid attention module".

[0058] (2) Although the structures of these four hybrid attention modules are the same, the sizes and numbers of channels of the output feature maps, as well as the number of module iterations, are different.

[0059] 1) The core module of the "first hybrid attention module" (i.e., the other parts of the "first hybrid attention module" except the "ninth convolutional layer") iterated 2 times ( Figure 1 in the upper right corner of the "first hybrid attention module"), and the first output feature map was F 2 , and then F 2 was input into the "first hybrid attention module" again, so as to output the feature map F 3 .

[0060] 2) The core module of the "second hybrid attention module" iterated 2 times, and the first output feature map was F 4 , and then F 4 was input into the "second hybrid attention module" again, so as to output the feature map F 5 .

[0061] 3) The core module of the "third hybrid attention module" iterated 4 times ( Figure 1 in the upper right corner of the "third hybrid attention module"), and the first output feature map was F 6 , and then F 6 was input into the "third hybrid attention module" to get F 7 , and so on, and the finally output feature map was F 9 .

[0062] 4) The core module of the "fourth hybrid attention module" iterated 2 times, and the first output feature map was F 10 , and then F 10 was input into the "fourth hybrid attention module" again, so as to output the feature map

[0063] (3) We iterated the "hybrid attention module" a total of 10 times, aiming to expand the receptive field, extract deep features, and improve the learning ability of the model. From the "first hybrid attention module" to the "fourth hybrid attention module", the size of the generated feature map was halved each time, and the number of channels doubled, aiming to expand the receptive field of the model, help the model understand the context information in the image, and this can, to a certain extent, alleviate the cross-window connection problem caused by the window shifting mechanism (the window shifting mechanism of SwinTransformer).

[0064] (4) Each "Hybrid Attention Module" is respectively composed of a "Cross Attention Module", a "Cross-Window Attention Module", and two convolutional layers.

[0065] 1) The "Cross Attention Module" and the "Cross-Window Attention Module" will be described in detail below.

[0066] 2) The function of the "Seventh Convolutional Layer" is to adjust the output feature size and channels of the "Cross-Window Attention Module" to so as to perform a concatenation operation with the features connected through the residual connection (i.e., the symbol ) in the overall architecture diagram, and the generated feature map is

[0067] 3) The core module of the "First Hybrid Attention Module" has been iterated 2 times in total. After the iteration ends, it is input into the "Ninth Convolutional Layer", and its function is to adjust the feature size and channels from to that is to prepare for subsequent feature extraction.

[0068] In this embodiment, as Figure 2 shown, (1) the "Cross Attention Module" is composed of two modules. The first module is iterated 2 times ( Figure 2 there is an x2 in the upper right corner, indicating that this module needs to be performed twice, and its first output will be used as the input of the second time). The second module uses the "Third Channel Attention Module" to process the cross-window information interaction. Both of these modules are similar to the Transformer structure.

[0069] (2) In the first module of the "Cross Attention Module", this method uses three modules in parallel: the "First Channel Attention Module", the "Second Channel Attention Module", and the "Window-Based Multi-Head Self-Attention Module". This can allow the three modules to extract features respectively, learn different information, and can speed up the learning process. In addition, channel attention can activate more global pixels because global information is involved in calculating channel attention, thereby enhancing the expression ability of the network.

[0070] The architecture diagram of the "First Channel Attention Module" is as Figure 3 shown:

[0071] 1) It is divided into three stages. One is the feature extraction stage, which consists of the "second convolutional layer" + a GELU activation function + the "third convolutional layer", and its function is to further extract the feature map, that is, the feature map. One is the channel attention stage, which consists of "global pooling" + the "fourth convolutional layer" + the GLUE activation function + the "fifth convolutional layer" + the Sigmoid function, and a skip connection with weight coefficients, and its function is to perform a simple attention calculation on the channels. The last one is the output stage, which consists of the "sixth convolutional layer" and is used to adjust the size and number of channels for subsequent processing.

[0072] Among them, the process of performing a simple attention calculation on the channels is as follows:

[0073]

[0074] Among them represents the input feature map of the "sixth convolutional layer", represents the feature map output after the "fifth convolutional layer" and the Sigmoid, represents the feature map output by the "third convolutional layer", and λ represents the weight coefficient, which is a learnable parameter, represents the pixel product.

[0075] 2) represents the Sigmoid function, represents the pixel product.

[0076] 3) In the channel attention stage, we adopt a skip connection with weight coefficients. This coefficient is a learnable parameter, which can automatically adjust the weight ratio of the global information, so as to automatically adjust the degree of global information in the channel attention according to the task, which is beneficial to the learning of the model.

[0077] 4) From the "second convolutional layer" to the "sixth convolutional layer", its structure and function are the same as those of the "convolutional layer" described above.

[0078] (3) The "second channel attention module" is the squeeze-and-excitation attention mechanism SENet, that is, Squeeze-and-Excitation Networks.

[0079] (4) The calculation method of the "window-based multi-head self-attention module": Given the input feature map as can be split into H cross ×W cross / M 2 windows of size M×M, and then calculate the self-attention within each small window. For each window feature map Its cross-attention mechanism is expressed as:

[0080]

[0081] Among them, d represents the dimension of the query matrix Q, K represents the key matrix, V represents the value matrix, and B represents the relative position encoding. To solve the problem of information connection between windows, the distance of each window movement for us is the floor of 1 / 3 of the window size.

[0082] In addition, this method also adopts the step-by-step expansion of the window size to further eliminate the information connection problem between windows. Specifically, the window size is divided into three scale levels of M×M, 2M×2M, and 4M×4M, and then the above operations are performed on these windows respectively, that is, cross-attention is performed inside the window, and then each movement

[0083] (5) Stack the three modules of "the first channel attention module", "the second channel attention module", and "the window-based multi-head self-attention module" together, that is, the three feature maps have the same size, and the number of channels is added to obtain the fused feature map.

[0084] (6) The "third channel attention module" is also used to solve the information connection problem between windows. Specifically, for the input feature X, its Q, K, and V feature maps are respectively For X Q divide it into non-overlapping windows of size M 1 ×M 1 , while for X K and X V divide it into overlapping windows of size M 2 ×M 2 , where M 2 1 =(1 + γ)×M 1 , and γ is a constant controlling the overlap size. The cross-attention of these three Q, K, and V still uses the above formula of Attention(Q, K, V), and the relative position variation

[0085] (7) Other parts of this "cross-attention module" are not much different from the standard Transformer. We use MLP (i.e., multi-layer perceptron) to replace the original forward feedback network. The "seventh convolutional layer" is used to process the addition and fusion of the skip connection and the output module, that is operation.

[0086] Among them, the processing of the addition and fusion of the skip connection and the output module includes:

[0087]

[0088] wherein represents the input feature map of the "eighth convolutional layer", represents the feature map output by the "first convolutional layer", represents the feature map output by the "seventh convolutional layer", represents pixel addition.

[0089] In this embodiment, the loss function is composed of L1 loss and photometric loss, that is

[0090]

[0091] where λ is a coefficient, and its value range is from 0 to 1, is the L1 loss function, is the photometric loss function.

[0092] In this embodiment, first, pre-training is performed using a general large dataset for visual tasks with a sufficient number of iterations. After pre-training, fine-tuning is performed using small datasets for super-resolution tasks and image inpainting tasks, and at this time, a relatively small learning rate is adopted to avoid overfitting problems for specific datasets. This strategy can effectively improve the learning ability of the model of this method.

[0093] In this embodiment, in the image generation and inpainting stage:

[0094] (1) After downsampling through the "tenth convolutional layer", the "reconstruction layer" reconstructs and inpaints the image, and then the "eleventh convolutional layer" upsamples to output an output image with four times the resolution

[0095] (2) The "reconstruction layer" adopts a mature model, which can be a post-upper sampling network structure or a progressive upper sampling network structure. For example, Pixel-Shuffle realizes the magnification effect by rearranging the pixels in the feature map while maintaining the computational efficiency; for example, Nearest Neighbor Upsampling enlarges the image by replicating adjacent pixels. Although it has a low computational cost and is easy to implement, it usually produces blocky artifacts and the visual quality is inferior to other more complex methods; for example, Bilinear and Bicubic Interpolation, which perform interpolation based on the weighted average of surrounding pixels and can produce relatively smooth results, but may lose some detailed information; for example, Depth-to-Space Transformation, which is an operation in TensorFlow and other deep learning frameworks to convert data in the channel dimension to the spatial dimension for the purpose of upsampling.

[0096] The Pixel-Shuffle reconstruction method is actually used in this method.

[0097] The present invention proposes a new "cross-attention module" and a new cross-window attention method, enabling it to calculate cross-attention within a larger range and excluding the interference of irrelevant information such as the image background, that is, it can solve the problems of image blurring and artifacts and has performance advantages.

[0098] Embodiment 2

[0099] The present invention also provides an image generation and restoration method for high-fidelity cultural relic digitization, which is implemented by applying the image generation and restoration system for high-fidelity cultural relic digitization described in any one of the above. The method includes:

[0100] Extract the shallow features of the input image;

[0101] Extract the deep features of the input image according to the shallow features;

[0102] Reconstruct and restore the image according to the deep features, and output an image with four times the resolution.

[0103] Embodiment 3

[0104] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a three-dimensional structure restoration method for high-quality urban renewal landscape architecture described in any one of the above embodiments.

[0105] Figure 4 FIG. 1 shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0106] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0107] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0108] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0109] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module may implement communication in a wired manner (such as USB (Universal Serial Bus), network cable, etc.) or in a wireless manner (such as a mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).

[0110] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0111] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.

[0112] The system of the above embodiment is used to implement the three-dimensional structure restoration method of a corresponding high-quality urban renewal landscape building in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0113] Embodiment 4

[0114] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute a three-dimensional structure restoration method of a high-quality urban renewal landscape building as described in any of the foregoing embodiments.

[0115] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0116] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute a three-dimensional structure restoration method of a high-quality urban renewal landscape building as described in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0117] Those of ordinary skill in the art should understand that any discussion of the above embodiments is exemplary only and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of brevity.

[0118] In addition, for simplicity of explanation and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the devices may be shown in block diagram form in order to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0119] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0120] Therefore, the units of the examples described in the embodiments of the present application can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0121] The above-described embodiments are only descriptions of the preferred modes of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solution of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A high-fidelity cultural relic digitization image generation and restoration system, characterized in that: The system comprises: a first feature extraction stage, a second feature extraction stage, and an image generation and restoration stage; The first feature extraction stage is used to extract shallow features of the input image; The second feature extraction stage is used to extract deep features of the input image based on the shallow features; The image generation and restoration stage is used to reconstruct and restore the image according to the deep-level features, and output an image with four times the resolution.

2. The system according to claim 1, characterized in that The first feature extraction stage includes: an input image unit, a residual network and a first convolutional layer; The input image unit is used to input an image; The residual network is used to extract shallow features of the input image; The first convolutional layer is used to downsample the shallow features and is also used as the output head of the residual network feature extraction.

3. The system according to claim 1, characterized in that The second feature extraction stage is composed of 4 hybrid attention modules with the same structure. The output feature map size, number of channels, and number of module iterations of each hybrid attention module are different. Each hybrid attention module is composed of a cross attention module, a cross-window attention module and two convolutional layers.

4. The system according to claim 3, characterized in that The cross-attention module consists of two modules. The first module iterates twice, and the second module uses the third-channel attention module to process cross-window information interaction. Among them, the first module uses the first-channel attention module, the second-channel attention module and the window-based multi-head self-attention module in parallel, allowing the three modules to extract features separately and learn different information.

5. The system according to claim 4, characterized in that The first channel attention module is divided into three stages: the first is the feature extraction stage, which consists of the second convolutional layer, a GELU activation function, and the third convolutional layer, which is used to further extract features; One is the channel attention stage, which consists of global pooling, the fourth convolution layer, the GLUE activation function, the fifth convolution layer, the Sigmoid function, and a jump connection with a weight coefficient. It performs a simple attention calculation on the channel. The last one is the output stage, which consists of the sixth convolution layer and is used to adjust the size and number of channels for subsequent processing.

6. The system according to claim 4, characterized in that The "second channel attention module" is the squeeze-and-Excitation attention mechanism SENet, namely Squeeze-and-Excitation Networks.

7. The system according to claim 4, characterized in that The calculation process of the window-based multi-head self-attention module includes: Given the input feature map Split into H cross ×W cross / M 2 A window of size M×M is formed, and then self-attention is calculated in each small window. For each window feature map The cross attention mechanism is expressed as: Where d represents the dimension of the query matrix Q, K represents the key matrix, V represents the value matrix, and B represents the relative position encoding; The window size is divided into three scale levels: M×M, 2M×2M and 4M×4M. Then, cross attention is performed inside the window, and each time the window moves, the window size is divided into three scale levels: M×M, 2M×2M and 4M×4M.

8. The system according to claim 7, characterized in that The calculation process of the third channel attention module includes: For the input feature X, the Q, K, and V feature maps are X Q Divide into non-overlapping windows of size M1×M1, and for X K and X V Divide into The overlapping window of size M2 = (1 + γ) × M1, γ is a constant that controls the overlap size; the cross attention of these three Q, K, V still uses the Attention(Q, K, V) formula, but the relative position becomes worse 9. A method for generating and restoring a digital image of a high-fidelity cultural relic, implemented by using the digital image generation and restoration system of a high-fidelity cultural relic as claimed in any one of claims 1 to 8, characterized in that: The method comprises: Extract shallow features of the input image; Extracting deep features of the input image based on the shallow features; According to the deep-level features, the image is reconstructed and repaired, and an image with four times the resolution is output.

10. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the method according to claim 9 is implemented when the processor executes the program.