Image importance detection method for cultural heritage and historical ancient city protection

Through the image importance detection model based on Transformer and stable diffusion model, the problem of identifying important areas of historical relics under complex background and lighting changes is solved, and stable detection and recognition under various conditions is achieved, with strong robustness and generalization capabilities.

CN120070356AInactive Publication Date: 2025-05-30BEIJING GUANGAN LIGHTING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510129406.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing computer vision methods are difficult to accurately identify important areas in historical sites under complex background and lighting changes, and changes in the natural environment affect image quality, increasing the difficulty of automatic recognition.

Method used

The image importance detection model is constructed based on Transformer and stable diffusion model, including encoding networks and decoding networks. Through the input module, encoding noise addition module, condition module and encoding denoising module, the historical relics image to be identified is obtained and the contour map and importance map are output.

Benefits of technology

It realizes the stable detection of meaningful objects under various conditions, has strong robustness and generalization capabilities, and can effectively utilize global long-range dependencies to detect and identify regional importance and object contour lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070356A_ABST
    Figure CN120070356A_ABST
Patent Text Reader

Abstract

The invention discloses an image importance detection method for cultural heritage and historical ancient city protection. The method comprises the following steps: constructing an image importance detection model based on Transform and a stable diffusion model; wherein the image importance detection model comprises a coding network and a decoding network; the coding network comprises an input module, a coding noise adding module, a condition module and a coding noise removing module; the decoding network comprises a decoding noise adding module, a decoding noise removing module and an output module; and then, obtaining a to-be-identified historical relic image, inputting the to-be-identified historical relic image into the image importance detection model, and outputting a contour graph and an importance graph to realize image importance detection. According to the method, the global long-range dependency relationship can be effectively utilized, more effective region importance and object contour line detection and recognition can be carried out, the robustness and generalization ability are high, and meaningful objects can be stably detected under various conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of artificial intelligence, computer vision, and computer graphics, and particularly relates to a method for detecting the importance of images for the protection of cultural heritage and historical ancient cities. Background Art

[0002] In the field of cultural heritage and historical ancient city protection, accurately identifying and separating key architectural structures or decorative elements in images is the basis for detailed research and restoration work. To ensure the authenticity and integrity of historical relics are maintained, experts need to rely on high-precision image processing technologies to assist their work. These tasks require algorithms not only to accurately identify regions with cultural value and importance but also to adapt to the influence of various factors such as background complexity, illumination changes, and material aging, which poses challenges to existing image analysis methods.

[0003] Traditional computer vision methods are difficult to address the above challenges because they cannot fully capture global information, especially when faced with complex background textures or uneven illumination conditions. In addition, since many historical sites may be located in outdoor environments, changes in the natural environment such as seasonal changes and weather conditions may also affect the quality of images, increasing the difficulty of automatic recognition. Therefore, in this specific application scenario, an ideal solution should possess strong robustness and generalization ability and be able to stably detect meaningful objects under various conditions. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a method for detecting the importance of images for the protection of cultural heritage and historical ancient cities to solve the problems existing in the above prior art.

[0005] To achieve the above object, the present invention provides a method for detecting the importance of images for the protection of cultural heritage and historical ancient cities, including:

[0006] Constructing an image importance detection model based on the Transformer and Stable Diffusion models; wherein, the image importance detection model includes an encoding network and a decoding network; the encoding network includes an input module, an encoding noise addition module, a conditional module, and an encoding denoising module; the decoding network includes a decoding noise addition module, a decoding denoising module, and an output module;

[0007] Obtaining a historical relic image to be recognized, inputting the historical relic image to be recognized into the image importance detection model, and outputting a contour map and an importance map to achieve image importance detection.

[0008] Optionally, input the historical relic image to be recognized into the image importance detection model. The condition module processes the image to be recognized through a pre-trained depth detection network to obtain a depth map, divides the depth map into blocks and inputs them into a Transformer to obtain a first feature map, and processes the first feature map through a forward sorting network to obtain a recombined feature map.

[0009] Optionally, the forward sorting network includes an input conversion layer, a segmentation layer, and an output conversion layer; the input conversion layer converts the first feature map into a 2D image based on multi-head self-attention and a multi-layer perceptron; divides the 2D image according to a preset ratio, recombines the segmented feature vectors, and obtains a recombined feature map.

[0010] Optionally, the encoding denoising module projects the noisy feature map output by the encoding noise addition module and the recombined feature map into the cross-attention mechanism of the Transformer through trilinear projection respectively to obtain corresponding key matrices, query matrices, and value matrices; performs cross-attention processing on the key matrices, query matrices, and value matrices and then fuses them to obtain a fused feature map; wherein, the encoding noise addition module performs noise addition processing on the feature map output by the input module through a stable diffusion model to obtain a noisy feature map.

[0011] Optionally, the cross-attention is as follows:

[0012]

[0013] Among them, the key matrix obtained by projecting the recombined feature map into the cross-attention mechanism of the Transformer through trilinear projection is represented as The query matrix is represented as The value matrix is represented as Where e is the size of the feature map and c is the number of channels of the feature map; the key matrix obtained by projecting the noisy feature map into the cross-attention mechanism of the Transformer through trilinear projection is represented as The query matrix is represented as The value matrix is represented as

[0014] Optionally, the decoding noise addition module performs noise addition processing on the fused feature map through a stable diffusion model to obtain a noisy image; splices the contour line block element and the importance block element with the noisy image in the same size and increasing depth manner to obtain a spliced image.

[0015] Optionally, the decoding denoising module decodes and denoises the spliced image based on the self-attention mechanism of the decoder itself and the design of the forward sorting network to obtain the processed contour line block element, importance block element, and noisy image.

[0016] Optionally, project the processed contour line block elements onto a binary map of 0 and 1 scalars through the contour line head to obtain a contour line map; project the processed importance block elements onto a binary map of 0 and 1 scalars through the importance head to obtain an importance map.

[0017] Optionally, the loss function of the image importance detection model is the sum of the loss function of the Stable Diffusion model and the binary cross-entropy loss function of the Transformer.

[0018] The present invention also provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the steps of the above-mentioned method for detecting the importance of images for the protection of cultural heritage and historical ancient cities.

[0019] Compared with the prior art, the present invention has the following advantages and technical effects:

[0020] The present invention first constructs an image importance detection model based on the Transformer and the Stable Diffusion model; wherein, the image importance detection model includes an encoding network and a decoding network; the encoding network includes an input module, an encoding noise addition module, a conditional module, and an encoding denoising module; the decoding network includes a decoding noise addition module, a decoding denoising module, and an output module; then, obtain the historical relic image to be recognized, input the historical relic image to be recognized into the image importance detection model, and output a contour line map and an importance map to achieve image importance detection. Based on the encoding and decoding architecture of the Transformer and Stable Diffusion, the present invention can effectively utilize the global long-range dependence relationship, thereby performing more effective detection and recognition of regional importance and object contour lines, and has strong robustness and generalization ability, and can stably detect meaningful objects under various conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0022] Figure 1 is a schematic diagram of the system architecture of an embodiment of the present invention;

[0023] Figure 2 is a forward sorting architecture diagram of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0025] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0026] Embodiment 1

[0027] As Figure 1 shown, in this embodiment, a method for detecting the importance of images for the protection of cultural heritage and historical ancient cities is provided, including:

[0028] Construct an image importance detection model based on the Transformer and the Stable Diffusion model; wherein, the image importance detection model includes an encoding network and a decoding network; the encoding network includes an input module, an encoding noise addition module, a conditional module, and an encoding denoising module; the decoding network includes a decoding noise addition module, a decoding denoising module, and an output module;

[0029] Obtain the historical relic image to be recognized, input the historical relic image to be recognized into the image importance detection model, and output a contour map and an importance map to realize image importance detection.

[0030] The input module is to pass the input image through a backbone network (the specific backbone network is not limited, it can be ResNet, VGG, etc.), and the purpose is to initially extract the feature map.

[0031] The encoding noise addition module is the same as the noise addition process of the Stable Diffusion model, and gradually adds noise to the feature map (or called feature vector, or called latent space) Z to Z T .

[0032] Furthermore, input the historical relic image to be recognized into the image importance detection model. The conditional module processes the image to be recognized through a pre-trained depth detection network to obtain a depth map, cuts the depth map into pieces and inputs it into the Transformer to obtain a first feature map, and processes the first feature map through a forward sorting network to obtain a recombined feature map.

[0033] The forward sorting network includes an input conversion layer, a segmentation layer, and an output conversion layer; the input conversion layer converts the first feature map into a 2D image through multi-head self-attention and a multi-layer perceptron; the 2D image is segmented according to a preset ratio, and the segmented feature vectors are recombined to obtain a recombined feature map.

[0034] Specifically, the conditional module first inputs the input image into a pre-trained depth detection module (not limited to a specific depth detection model, which can be Monodepth2 based on CNN, self-supervised GeoNet, or Detection Transformer based on Transformer, etc.), and outputs a depth map. In Figure 1 In the system architecture diagram shown, a lock is drawn in the upper right corner of the depth detection module, indicating that it is pre-trained and directly used.

[0035] Then, the depth map is sliced and input into the Transformer to obtain the first feature map. The specific operation is similar to the ViT model (i.e., Vision Transformer).

[0036] Finally, the first feature map is input into the "positive sorting". The architecture diagram of the "positive sorting" is as Figure 2 shown:

[0037] "Positive sorting" is forward sorting, which includes three layers: an "input conversion layer" (a network composed of a multi-layer perceptron, i.e., MLP, and a multi-head attention mechanism), a "segmentation layer", and an "output conversion layer". The purpose of "positive sorting" is to establish the relationship between local feature vectors (similar to Transformer), which is a self-attention mechanism. And there is also a Transformer layer before "positive sorting", so it also has long-range global relationships, making this method have both global and local dual relationships.

[0038] The "input conversion layer" is to convert the input multi-dimensional feature vector into a form similar to a 2D image; as Figure 2 shown, each number is a feature vector, and the purpose is to piece them together into a feature map. Given a block element T' with a length of e output by the previous Transformer, the "input conversion layer" is used to transform T' to obtain a new block element Mathematically expressed as: T = MLP(MSA(T')), where MSA and MLP respectively represent the multi-head self-attention and multi-layer perceptron in the original Transformer. In addition, layer normalization is performed before each block. Then, T is deformed into a 2D image I, and where l = h × w to restore the spatial structure.

[0039] The "segmentation layer" divides the image I into k × k, similar to the way of image convolution.

[0040] The "output conversion layer" is to recombine the segmented feature vectors to form a new sequence.

[0041] Furthermore, the encoding denoising module embeds the noisy feature map and the recombined feature map output by the encoding noise addition module into the cross-attention mechanism of the Transformer through trilinear projection respectively to obtain the corresponding key matrix, query matrix and value matrix; after performing cross-attention processing on the key matrix, query matrix and value matrix, they are fused to obtain a fused feature map; among them, the encoding noise addition module performs noise addition processing on the feature map output by the input module through the Stable Diffusion model to obtain a noisy feature map.

[0042] The cross-attention is as follows:

[0043]

[0044] Among them, the key matrix obtained by embedding the recombined feature map into the cross-attention mechanism of the Transformer through trilinear projection is denoted as The query matrix is denoted as The value matrix is denoted as where l is the size of the feature map and c is the number of channels of the feature map; the key matrix obtained by embedding the noisy feature map into the cross-attention mechanism of the Transformer through trilinear projection is denoted as The query matrix is denoted as The value matrix is denoted as

[0045] Specifically, the "encoding denoising module" is similar to the denoising process of the Stable Diffusion model. There are two differences. One is the use of the cross-attention mechanism (i.e., QKV) of the Transformer, and the other is the addition of a "forward indexing" operation after each output.

[0046] Embed the recombined feature map output by the conditional module into QKV through trilinear projection to obtain three matrices, namely the key matrix The query matrix And the value matrix where l is the size of the feature map, that is, width * height, and c is the number of channels of the feature map.

[0047] Similarly, embed the noisy feature map output after passing the "input image" through the "input module" and the "encoding noise addition module" into KQV through trilinear projection to obtain three matrices, namely the key matrix The query matrix And the value matrix

[0048] Then perform cross-attention, and the formula is as follows (this formula is the same as the cross-attention of the Transformer, but the KQV in this embodiment is a combination of two sets):

[0049]

[0050] The other parts are the same as those of the Transformer, including the multi-head attention mechanism, the position feed-forward network, the residual connection layer, and the normalization layer, etc.

[0051] Finally, add a fusion layer (which can be a Transformer or an MLP, and we recommend using a Transformer to enhance the ability of cross-embedding) to integrate the two cross-attention mechanisms together.

[0052] In Figure 1 In the system architecture shown, only one denoising process is drawn for the "encoding denoising module". There are T + 1 feature vectors in the whole denoising process, so T iterations are required. In addition, Figure 1 It is also shown that there are skip connections inside, which is the same as the original Stable Diffusion model.

[0053] Furthermore, the decoding and adding noise module adds noise to the fused feature map through the Stable Diffusion model to obtain a noisy image; the contour block element and the importance block element are concatenated with the noisy image in a way of having the same size and increasing the depth to obtain a concatenated image.

[0054] Specifically, as Figure 1 shown, the Z' in the "encoding denoising module" and the "decoding and adding noise module" is the same one. The "decoding and adding noise module" is similar to the "encoding and adding noise module", both are processes of adding noise step by step. The "decoding and adding noise module" adds noise from Z' to Z' T .

[0055] For Z' T two learnable "contour block elements" and "importance block elements" are added, which are two feature vectors (the initial values are randomly assigned and then the required features are learned through the network), and their functions are to output the "contour map" and the "importance map", which are two different tasks.

[0056] Concatenate the "contour block element" P Con and the "importance block element" P Imp to Z' T Generally, concatenation is performed in a way of having the same size and increasing the depth. For example, Then

[0057] Furthermore, the decoding denoising module decodes and denoises the concatenated image based on the self-attention mechanism of the decoder itself and the design of the forward sorting network to obtain the processed contour block element, importance block element, and noisy image.

[0058] Specifically, the "decoding denoising module" is similar to the "encoding denoising module", except that the QKV here is the self-attention of itself, without introducing external cross again. In addition, the "forward arrangement" here is changed to "reverse arrangement", and the processes of the two are the same, except that the order of the two processes is exactly opposite (because this is the decoding process, which is reverse to the encoding process).

[0059] Further, the processed contour line block elements are projected onto a binary image of 0 and 1 scalars through the contour line head to obtain a contour line image; the processed importance block elements are projected onto a binary image of 0 and 1 scalars through the importance head to obtain an importance image.

[0060] Specifically, the "output module" has two linear heads, one is the "contour line head" and the other is the "importance head". Each head uses a Sigmoid activation layer to achieve projection, and respectively projects the processed P Con and P Imp onto a binary image of 0 and 1 scalars, and reshapes them into an importance image and a contour line image respectively.

[0061] The loss function of the image importance detection model is the sum of the loss function of the stable diffusion model and the binary cross-entropy loss function of the Transformer.

[0062] This embodiment provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the steps of the above-mentioned image importance detection method for the protection of cultural heritage and historical ancient cities.

[0063] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above-mentioned image importance detection method for the protection of cultural heritage and historical ancient cities.

[0064] This embodiment provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the above-mentioned image importance detection method for the protection of cultural heritage and historical ancient cities.

[0065] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for detecting the importance of images for the protection of cultural heritage and historical ancient cities, characterized in that: The following steps are involved: An image importance detection model is constructed based on Transformer and stable diffusion models; wherein the image importance detection model includes an encoding network and a decoding network; the encoding network includes an input module, an encoding denoising module, a conditional module and an encoding denoising module; the decoding network includes a decoding denoising module, a decoding denoising module and an output module; The image of the historical relics to be identified is obtained, the image of the historical relics to be identified is input into the image importance detection model, and a contour line map and an importance map are output to realize image importance detection.

2. The image importance detection method for cultural heritage and historical city protection according to claim 1 is characterized in that: The image of the historical relics to be identified is input into the image importance detection model. The conditional module processes the image to be identified through a pre-trained depth detection network to obtain a depth map. The depth map is cut into blocks and input into the Transformer to obtain a first feature map. The first feature map is processed through a forward sorting network to obtain a recombined feature map.

3. The image importance detection method for cultural heritage and historical city protection according to claim 2 is characterized in that: The forward sorting network includes an input conversion layer, a segmentation layer and an output conversion layer; the input conversion layer converts the first feature map into a 2D image based on multi-head self-attention and a multi-layer perceptron; the 2D image is segmented according to a preset ratio, and the feature vectors obtained by the segmentation are recombined to obtain a recombined feature map.

4. The image importance detection method for cultural heritage and historical city protection according to claim 2 is characterized in that: The coding denoising module embeds the noisy feature map output by the coding denoising module and the reorganized feature map into the cross attention mechanism of the Transformer through trilinear projection respectively to obtain the corresponding key matrix, query matrix and value matrix; cross-attention processing is performed on the key matrix, query matrix and value matrix and then fused to obtain a fused feature map; wherein, the coding denoising module performs noise processing on the feature map output by the input module through a stable diffusion model to obtain a noisy feature map.

5. The image importance detection method for cultural heritage and historical city protection according to claim 4 is characterized in that: The cross attention is as follows: Among them, the reorganized feature map is embedded into the key matrix obtained by the cross attention mechanism of Transformer through trilinear projection and is expressed as The query matrix is ​​represented as The value matrix is ​​represented as in is the size of the feature map, c is the number of channels of the feature map; the key matrix obtained by embedding the noisy feature map into the cross attention mechanism of Transformer through trilinear projection is expressed as The query matrix is ​​represented as The value matrix is ​​represented as 6. The image importance detection method for cultural heritage and historical city protection according to claim 4 is characterized in that: The decoding and denoising module performs denoising processing on the fused feature map through a stable diffusion model to obtain a noise image; and splices the contour block element, the importance block element and the noise image in a manner of having the same size and increasing depth to obtain a spliced ​​image.

7. The image importance detection method for cultural heritage and historical city protection according to claim 6 is characterized in that: The decoding denoising module decodes and denoises the spliced ​​image based on the self-attention mechanism of the decoder itself and the design of the forward sorting network to obtain processed contour blocks, importance blocks and noise images.

8. The image importance detection method for cultural heritage and historical city protection according to claim 7 is characterized in that: The processed contour block element is projected to a binary map of 0 and 1 scalars through the contour head to obtain a contour map; the processed importance block element is projected to a binary map of 0 and 1 scalars through the importance head to obtain an importance map.

9. The image importance detection method for cultural heritage and historical city protection according to claim 1 is characterized in that: The loss function of the image importance detection model is the sum of the loss function of the stable diffusion model and the binary cross entropy loss function of the Transformer.

10. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the image importance detection method for the protection of cultural heritage and historical ancient cities as described in any one of claims 1-9.