Multi-scale reconstruction method and system for building landscape and historical relic images
By analyzing and cropping the importance of architectural landscape and historical cultural relics images, the problem of inaccurate identification of important areas of images in the existing technology is solved, and the accurate multi-scale reconstruction of the image and the improvement of visual effects are achieved.
Patent Information
- Application Number
- CN202510105861.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is not accurate enough when automatically identifying important areas of the image, especially when processing large single-body building landscapes and historical relics images, resulting in deformation or loss of important objects, or introducing local artifacts.
By analyzing the importance of architectural landscape and historical relics images, a binary map with important information is generated, and multi-scale reconstruction is carried out through cropping, alignment and content completion methods to ensure the accurate retention of important areas and objects.
Accurate multi-scale reconstruction of architectural landscapes and historical relics images is achieved, avoiding the deformation or loss of important objects, and improving the adaptability and visual effect of the images.
Smart Images

Figure CN120047823A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image reconstruction, and particularly to a multi-scale reconstruction method and system for architectural landscape and historical relic images. Background Art
[0002] Image multi-scale reconstruction is the task of adjusting the aspect ratio of an image to adapt to different display devices or display environments. It is necessary to reconstruct the "input image", and then output an image of any size, while retaining the important regions, important objects, and important parts of the objects in the image, and at the same time maintaining the image quality (avoiding deformation and distortion). This plays an important role in many fields. For example, in scenic spots or museums, when making promotional pictures, it is necessary to customize according to the specific size of the venue, or print them into books, advertising pages, gift boxes of specific sizes.
[0003] However, the current methods are not accurate enough in automatically identifying important regions of images, especially different parts of the same object (such as a human face), resulting in deformation or loss of important objects, or introducing local artifacts. And the current methods mostly use the method of reducing the object spacing to reduce the size for multi-scale reconstruction. However, in the application scenario of large single building landscapes, usually there is only one object in the picture, and only a part of it. In this case, the current methods cannot handle it. Moreover, when there is only one object in the picture, such as a historical relic, and its picture contains the complete object contour but only one object, the current methods also cannot handle it. Summary of the Invention
[0004] To solve the above technical problems in the background, for the application scenario of large single building landscapes, the present invention conducts an importance analysis on this large single object and reduces it by means of cropping. For pictures of single objects such as "historical relics", we can identify the gap between this single object and the image edge, delete the gap, and focus on displaying (such as zooming) the important regions of this single object. In addition, this method additionally uses text information as a guiding condition for guiding the generation of image background and gap content, which has a great effect and can make the generated effect more excellent.
[0005] To achieve the above object, the present invention provides the following solutions:
[0006] A multi-scale reconstruction method for architectural landscape and historical relic images, characterized in that the steps include:
[0007] Collect the architectural landscape and historical relic images to be reconstructed;
[0008] Conduct an importance analysis on the architectural landscape and historical relic images to be reconstructed to obtain a binary image with importance information;
[0009] Crop and align the obtained binary image to obtain a cropped image;
[0010] Complete the content of the cropped image and output it to complete the reconstruction.
[0011] Preferably, the steps of performing the importance analysis include:
[0012]
[0013] where P(x, y) represents the pixel at the image coordinate (x, y); I E (x, y) represents the energy value at the image coordinate (x, y); I E represents the energy distribution map of the image; where α, β, and γ represent coefficients; ⊙ represents the per-pixel dot product; N represents the number of objects detected by the target in the input image.
[0014] Preferably, calculate α, β, and γ according to the input image and the input ratio. When the width of the input image is W, the height is H, and the input ratio is a:b, the three coefficients are respectively:
[0015]
[0016] β = 1 - α
[0017] γ = 1.
[0018] Preferably, the steps of performing the crop alignment include:
[0019]
[0020] where W target represents the target width, W target = H * r; H represents the height of the input image; r represents the input ratio; η represents the tolerable importance loss rate ; The formula for calculating the image height is: By comparing H M and H;
[0021]
[0022] where H target represents the target height, H target = W * r; W represents the width of the input image. The formula for calculating the image width is:
[0023] Preferably, select W M , H M or H′ M , W′ M as the size W Final, H Final :
[0024] If the input ratio r >> 1, it indicates width priority, i.e., W Final = W M , and H Final = H M ;
[0025] If the input ratio r < 1, it indicates height priority, i.e., W Final = W' M , and H Final = H' M .
[0026] Preferably, the steps for content completion include:
[0027]
[0028] where t represents a random variable of the time step from 1 to T; T represents the time step of the diffusion process; x 0 represents the original data sample; ∈ represents the noise of the standard normal distribution; x t represents the data with added noise at time step t; ∈ θ (x t , t) represents the noise predicted by the model, parameterized as θ; ||·|| 2 represents the L2 norm, i.e., the Euclidean distance.
[0029] The present invention also provides a multi-scale reconstruction system for architectural landscape and historical relic images, which is used to implement the above method and includes: a collection module, an analysis module, a cropping module, and a completion module;
[0030] The collection module is used to collect architectural landscape and historical relic images to be reconstructed;
[0031] The analysis module is used to perform importance analysis on the architectural landscape and historical relic images to be reconstructed to obtain a binary map with importance information;
[0032] The cropping module is used to crop and align the obtained binary map to obtain a cropped image;
[0033] The completion module is used to complete the content of the cropped image and output it to complete the reconstruction.
[0034] Preferably, the working process of the analysis module includes:
[0035]
[0036] where P(x, y) represents the pixel at the image coordinate (x, y); I E(x, y) represents the energy value at the image coordinates (x, y); I E represents the energy distribution map of the image; where α, β, and γ represent coefficients; ⊙ represents the per-pixel dot product; and N represents the number of objects detected by the target in the input image.
[0037] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above method is implemented.
[0038] The present invention also provides a computer-readable storage medium storing a computer program, which implements the above method when the computer program is executed.
[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0040] Through the multi-scale reconstruction method, the present invention can accurately identify and retain important regions and objects in architectural landscape and historical relic images, and at the same time perform intelligent cropping and content completion according to the input ratio and text information, effectively avoiding deformation or loss of important objects, and improving the adaptability and visual effect of the images. It is applicable to a variety of display devices and application scenarios, and has broad application prospects and practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 is a schematic flowchart of the method of the present invention;
[0043] Figure 2 is a detailed flowchart of the present invention;
[0044] Figure 3 is a schematic structural diagram of the electronic device according to an embodiment of the present invention.
[0046] 1010. Processor; 1020 Memory; 1030 Input / Output Interface; 1040 Communication Interface; 1050 Bus. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0049] As Figure 1 shown, it is a schematic flowchart of the method in this embodiment, and the steps include:
[0050] S1. Collect images of the building landscape and historical relics to be reconstructed.
[0051] The images collected in this embodiment are mainly for the application scenarios of large single building landscapes, and the sources can be open-source platforms or promotional websites, etc.
[0052] S2. Conduct importance analysis on the images of the building landscape and historical relics to be reconstructed to obtain a binary image with importance information.
[0053] In this embodiment, an importance analysis module is constructed to perform importance analysis. The images of the building landscape and historical relics to be reconstructed are used as the "input" to obtain a binary image with importance information.
[0054] The "input" includes three contents: image, ratio, and text.
[0055] The "input image" is the main body, which needs to be reconstructed at multiple scales and processed into the "output image".
[0056] The "input ratio" is the limiting condition for the length-width ratio of the "output image" (for example, restricting the "output image" to a size of 16:9, etc. The specific size requirements are determined according to the needs).
[0057] The "input text" has two functions. One is to determine the importance of a single object or a part of an object in the "input image" (for example, if there are several cultural relics in the "input image", it is necessary to determine which cultural relic is the most important through the "input text" for highlighting; another example is that if there is only one cultural relic in the "input image", it is necessary to determine which part of this cultural relic is the most important through the "input text" for highlighting), and manage the occlusion relationship of multiple objects according to this importance (the more important one is in the front, and the less important one is in the back), as well as the cropped parts of the object (including the cases of a single object and multiple objects, and the cropped parts should be those with lower importance). Another function is to serve as a condition for the "diffusion module" (that is, encoding the text into Tokens and inputting them into the UNet or Transformer of the "diffusion module", specifically depending on the technical solution of the "diffusion module") to generate the background and blank content.
[0058] The role of the importance analysis module is to accurately calculate which parts in the "input image" are more important according to algorithms, pre-trained models, etc., and generate a binary map. Its main workflow includes: importance detection, object detection, and determining the importance of size.
[0059] Importance detection is processed using a pre-trained importance detection model. This embodiment does not limit the specific detection method or model, including but not limited to PiCANet, BASNet, CPD-R, PoolNet, TSPOANet, A2dele, JL-DCF, SSF-RGBD, UC-Net, etc. Use I I to represent the output importance map, which is a grayscale image with the same size as the input image (that is, there is only one channel, and the value ranges from 0 to 255).
[0060] Object detection is processed using a pre-trained object detection model. This method does not limit the specific detection method or model, including but not limited to R-CNN, FastR-CNN, FasterR-CNN, YOLO, SSD, RetinaNet, MaskR-CNN, EfficientDet, CascadeR-CNN, CenterNet, DETR, HRNet, etc. This embodiment uses I T to represent the output object detection map.
[0061] After that, I I and I TPerform dot product, perform segmentation, and obtain several grayscale images, denoted as (the so-called "several" is mainly based on the number of objects detected by "object detection" in the input image, that is, segmented by object. If there is only one object in the input image, then there is only one binary image. If there are multiple objects in the input image, then multiple objects will be recognized, and there will be multiple binary images. In addition, the subscript j of I ITj represents the index, with values ranging from 0 to M, and M is the number of all objects in the input image and also the number of binary images, that is, the number of objects, and the number of binary images recognized and segmented will be the same).
[0062] Then perform binarization on I ITj . In this embodiment, by setting a threshold (such as 127), for all pixels on I ITj , those greater than the threshold have their pixel values set to 1; those less than the threshold have their pixel values set to 0. In this way, these grayscale images are converted into binary images, denoted as
[0063] The role of determining size importance is to find important regions in the horizontal and vertical scales of the input image. In this embodiment, through the horizontal and vertical (where W and H are the width and height of the input image; x and y are the horizontal and vertical coordinates of the current pixel; x m represents the abscissa of the importance center point of the m-th segmented binary image; y m represents the ordinate of the importance center point of the m-th segmented binary image), further calculations are performed on the important regions of the input image, and the importance gradually decreases from the importance center point to the edge region.
[0064] The importance analysis module of this embodiment adopts an image energy algorithm to calculate the energy distribution of the image through the formula:
[0065]
[0066] where P(x, y) represents the pixel at the image coordinate (x, y); I E (x, y) represents the energy value at the image coordinate (x, y); I E represents the energy distribution map of the image, which is a grayscale image.
[0067] Combining the above steps, the final formula for importance analysis is:
[0068]
[0069] where α, β, and γ represent coefficients; ⊙ represents per-pixel dot product; N represents the number of objects detected by object detection in the input image.
[0070] α, β, and γ can be calculated based on the input image and the input ratio. When the width of the input image is W, the height is H, and the input ratio is a:b, the three coefficients are as follows:
[0071]
[0072] β = 1 - α
[0073] γ = 1
[0074] In addition, in the above represents Figure 2 the "multiply - add" operation in Figure 2 the "add" operation in is the addition of the above three parts.
[0075] After determining the important regions, the unimportant regions need to be deleted. By calculating the energy of each pixel through I e (x, y), a distribution map is formed. By setting a threshold, all regions smaller than the threshold become unimportant regions and are deleted. Figure 2 The comparison operation in
[0076] is based on the above formula to ensure that the width of all high - energy regions is not less than the preset ratio ξ of the input ratio, aiming to prevent objects from being too crowded together, resulting in an imbalance in the overall image ratio. key The final output of the importance analysis module is a binary map with importance information, denoted as I
[0077] At this time, the image size is the same as the input image size, and no cropping operation has been performed.
[0078] The obtained binary map is cropped and aligned through the construction of a cropping and alignment module to obtain the cropped image.
[0079] The "first convolution" in the cropping and alignment module is a one - dimensional convolution operation with a convolution kernel of length l (in this embodiment, l = 10). All values of this convolution kernel are the number 1. The purpose of the "first convolution" is to further retain the important regions and delete the trivial unimportant edge regions. After the "first convolution" of I key denoted as I' key is a binary map, and at this time, the image size is the same as the input image size.
[0080] Figure 2 The "multiplication" operation in keyPerforming a pixel-by-pixel multiplication with the input image is equivalent to extracting the important region from the input image, and the resulting image is denoted as I″ key , which is an RGB image. At this time, the image size is the same as that of the input image.
[0081] "Size cropping" in the system architecture diagram is to crop I″ key to obtain an image size that conforms to the "input ratio". The steps include:
[0082] Calculating the width of the important region:
[0083]
[0084] Among them, W imp represents the width of the important region; represents the pixel with energy higher than the preset threshold at the (x, y) position, represents the sum of all low-energy pixels in the width direction, and its width is the width of the important region.
[0085] Calculating the height of the important region:
[0086]
[0087] Among them, H imp represents the height of the important region; represents the sum of all low-energy pixels in the height direction, and its width is the height of the important region.
[0088] From this, the final image size can be obtained:
[0089]
[0090] Among them, W target represents the target width, W target = H * r; H represents the height of the input image; r represents the input ratio; η represents the tolerable importance loss rate (In this embodiment, λ = 0.1) , and the formula for calculating the image height is: By comparing H M and H, it is determined whether the image needs to be extended in the height direction. If extension is required, the extended height H M - H will be evenly distributed on the upper and lower sides of the image.
[0091] Similarly, it can be obtained:
[0092]
[0093] Among them, H target represents the target height, H target= W * r; where W represents the width of the input image; r represents the input ratio; η is the tolerable importance loss rate (in this embodiment λ = 0.1 ), the formula for calculating the image width is: By comparing W′ M and W, determine whether the image needs to be expanded in the width direction. If expansion is required, the expanded width W′ M - W will be evenly distributed on both the left and right sides of the image.
[0094] Select W M , H M or H′ M , W′ M as the size W Final , H Final of the final image:
[0095] If the input ratio r >> 1, it indicates that width is prioritized, i.e., W Final = W M , and H Final = H M ;
[0096] If the input ratio r < 1, it indicates that height is prioritized, i.e., W Final = W′ M , and H Final = H′ M .
[0097] After determining the final size, delete the unnecessary regions. When the input image contains multiple objects, arrange the objects according to the spacing ratio in the input image; if there is only one object, there is no need to consider the spacing issue.
[0098] S4. Complement the content of the cropped image and output it to complete the reconstruction.
[0099] Complement the content of the cropped image by constructing a content complement module. The diffusion model in the content complement module does not specify a specific model structure. It can use a diffusion model based on StableDiffusion or a diffusion model based on DiffusionTransformer.
[0100] Encode the input image and input it into U - Net (a diffusion model based on StableDiffusion) or Transformer (a diffusion model based on DiffusionTransformer) as a conditional input to the diffusion model to guide the "diffusion model" to generate the blank gaps and the image background content.
[0101] The input ratio is input into the second convolution to generate a feature map, which is then input into a U-Net or a Transformer and used as a conditional input to the diffusion model to guide the "diffusion model" in generating the blank gaps and the image background content.
[0102] The input text is encoded and then input into a U-Net or a Transformer, which is used as a conditional input to the diffusion model to guide the "diffusion model" in generating the blank gaps and the image background content. The guiding effect of this "input text" is very significant. It can perform customized output for the blank gaps and the image background, change the occlusion relationship of the objects in the image, specify the importance of the objects or the important parts of the objects themselves, and can also perform scaling according to the importance.
[0103] The content completion module uses the following loss function:
[0104]
[0105] where t represents a random variable of the time step from 1 to T; T represents the time step of the diffusion process; x 0 represents the original data sample; ∈ represents the noise of the standard normal distribution; x t represents the noise-added data at time step t, which is calculated from x 0 and ∈ according to a specific noise schedule; ∈ θ (x t , t) represents the noise predicted by the model, parameterized as θ, which is used to try to estimate the noise ∈ added to the data at time t; ||·|| 2 represents the L2 norm, that is, the Euclidean distance.
[0106] The objective of this loss function is to make the noise ∈ θ (x t , t) predicted by the model as close as possible to the true noise ∈ actually added at time step t. During the training process, by minimizing this loss, the model learns how to reverse the diffusion process, that is, how to denoise and finally generate new samples similar to the original data.
[0107] Finally, an image that conforms to the input ratio and the description of the input text is output. The detailed process of this embodiment is as Figure 2 shown.
[0108] Embodiment 2
[0109] This embodiment also provides a multi-scale reconstruction system for architectural landscape and historical relic images, including: a collection module, an analysis module, a cropping module, and a completion module; the collection module is used to collect the architectural landscape and historical relic images to be reconstructed; the analysis module is used to perform importance analysis on the architectural landscape and historical relic images to be reconstructed to obtain a binary image with importance information; the cropping module is used to crop and align the obtained binary image to obtain a cropped image; the completion module is used to complete the content of the cropped image and output it to complete the reconstruction.
[0110] Next, in combination with this embodiment, it will be detailed how the present invention solves the technical problems in actual work.
[0111] First, use the collection module to collect the architectural landscape and historical relic images to be reconstructed.
[0112] The images collected in this embodiment are mainly for the application scenarios of large single architectural landscapes, and the sources can be open-source platforms or promotional websites, etc.
[0113] After that, the analysis module performs importance analysis on the architectural landscape and historical relic images to be reconstructed to obtain a binary image with importance information.
[0114] In this embodiment, an importance analysis module is constructed to perform importance analysis. The architectural landscape and historical relic images to be reconstructed are used as the "input" to obtain a binary image with importance information.
[0115] The "input" includes three contents: image, ratio, and text.
[0116] The "input image" is the main body, which needs to be multi-scale reconstructed and processed into the "output image".
[0117] The "input ratio" is the limitation condition for the length-width ratio of the "output image" (for example, restricting the "output image" to a size of 16:9, etc. The specific size requirements are determined according to the needs).
[0118] The "input text" has two functions. One is to determine the importance of a single object or a part of an object in the "input image" (for example, if there are several cultural relics in the "input image", it is necessary to determine which cultural relic is the most important through the "input text" for highlighting; another example is that if there is only one cultural relic in the "input image", it is necessary to determine which part of this cultural relic is the most important through the "input text" for highlighting), and manage the occlusion relationship of multiple objects according to this importance (the more important one is in the front, and the less important one is in the back), as well as the cropped parts of the object (including the cases of a single object and multiple objects, and the cropped parts should be those with lower importance). Another function is to serve as a condition for the "diffusion module" (that is, encoding the text into Tokens and inputting them into the UNet or Transformer of the "diffusion module", depending on the technical solution of the "diffusion module") to generate the background and blank content.
[0119] The role of the importance analysis module is to accurately calculate which parts in the "input image" are more important according to algorithms, pre-trained models, etc., and generate a binary map. Its main workflow includes: importance detection, object detection, and determining the importance of size.
[0120] Importance detection is processed using a pre-trained importance detection model. This embodiment does not limit the specific detection method or model, including but not limited to PiCANet, BASNet, CPD-R, PoolNet, TSPOANet, A2dele, JL-DCF, SSF-RGBD, UC-Net, etc. Use I I to represent the output importance map, which is a grayscale image with the same size as the input image (that is, there is only one channel, and the value ranges from 0 to 255).
[0121] Object detection is processed using a pre-trained object detection model. This method does not limit the specific detection method or model, including but not limited to R-CNN, FastR-CNN, FasterR-CNN, YOLO, SSD, RetinaNet, MaskR-CNN, EfficientDet, CascadeR-CNN, CenterNet, DETR, HRNet, etc. This embodiment uses I T to represent the output object detection map.
[0122] After that, I I and I TPerform dot multiplication, perform segmentation, and obtain several grayscale images, denoted as (the so-called "several" is mainly based on the number of objects "detected by target detection" in the input image, that is, segmented by object. If there is only one object in the input image, then there is only one binary image. If there are multiple objects in the input image, then multiple objects will be recognized, and there will be multiple binary images. In addition, the subscript j of I ITj represents the index, with a value range of 0 to M, and M is the number of all objects in the input image and also the number of binary images, that is, the number of objects, and the same number of binary images will be recognized and segmented).
[0123] Then perform binarization on I ITj . In this embodiment, by setting a threshold (such as 127), for all pixels on I ITj , if the pixel value is greater than the threshold, its pixel value is set to 1; if it is less than the threshold, its pixel value is set to 0. In this way, these grayscale images are converted into binary images, denoted as
[0124] The role of determining size importance is to find important regions in the horizontal and vertical scales of the input image. In this embodiment, through horizontal and vertical (where W and H are the width and height of the input image; x and y are the horizontal and vertical coordinates of the current pixel; x m represents the abscissa of the importance center point of the m-th segmented binary image; y m represents the ordinate of the importance center point of the m-th segmented binary image), further calculate the important regions of the input image, and the importance gradually decreases from the importance center point to the edge region.
[0125] The importance analysis module of this embodiment adopts an image energy algorithm and calculates the energy distribution of the image through the formula:
[0126]
[0127] where P(x, y) represents the pixel at the image coordinate (x, y); I E (x, y) represents the energy value at the image coordinate (x, y); I E represents the energy distribution diagram of the image, which is a grayscale image.
[0128] Combining the above steps, the final formula for importance analysis is:
[0129]
[0130] where α, β, and γ represent coefficients; ⊙ represents per-pixel dot product; N represents the number of objects detected by target detection in the input image.
[0131] α, β, and γ can be calculated based on the input image and the input ratio. When the width of the input image is W, the height is H, and the input ratio is a:b, the three coefficients are as follows:
[0132]
[0133] β = 1 - α
[0134] γ = 1
[0135] In addition, in the above, represents Figure 2 the "multiply-add" operation in; Figure 2 the "add" operation in is the addition of the above three parts, represents the importance distribution. Since there are often more than one importance center, such as a person's face and body, which may both be importance centers, an importance distribution is needed to balance the distribution of the entire energy.
[0136] After determining the important regions, the unimportant regions need to be deleted. By I e (x, y), the energy of each pixel is obtained to form a distribution map. By setting a threshold, all regions smaller than the threshold become unimportant regions and are deleted. Figure 2 The comparison operation in is based on the above formula, ensuring that the width of all high-energy regions is not less than the preset ratio ξ of the input ratio, aiming to prevent objects from being too crowded together and causing the proportion of the entire image to be unbalanced.
[0137] The final output of the importance analysis module is a binary map with importance information, denoted as I key , and at this time, the image size is the same as the input image size, and no cropping operation has been performed.
[0138] The obtained binary map is cropped and aligned using the cropping module to obtain the cropped image.
[0139] The obtained binary map is cropped and aligned by constructing a cropping alignment module to obtain the cropped image.
[0140] The "first convolution" in the cropping alignment module is a convolution operation performed by a one-dimensional convolution kernel of length l (in this embodiment, l = 10), and all values of this convolution kernel are the number 1. The purpose of the "first convolution" is to further retain the important regions and delete the trivial unimportant edge regions. After the "first convolution" of I key , denoted as I' key , which is a binary map, and at this time, the image size is the same as the input image size.
[0141] Figure 2 The "multiply" operation in is to multiply I' keyPerforming a pixel-by-pixel multiplication with the input image is equivalent to extracting the important region from the input image, and the resulting image is denoted as I″ key , which is an RGB image. At this time, the image size is the same as the input image size.
[0142] In the "size cropping" in the system architecture diagram, it is to crop I″ key to obtain an image size that conforms to the "input ratio". The steps include:
[0143] Calculating the width of the important region:
[0144]
[0145] where W imp represents the width of the important region; represents the pixels with energy higher than the preset threshold at the (x, y) position, represents the sum of all low-energy pixels in the width direction, and its width is the width of the important region.
[0146] Calculating the height of the important region:
[0147]
[0148] where H imp represents the height of the important region; represents the sum of all low-energy pixels in the height direction, and its width is the height of the important region.
[0149] Thus, the final image size can be obtained:
[0150]
[0151] where W target represents the target width, W target = H * r; H represents the height of the input image; r represents the input ratio; η represents the tolerable importance loss rate (λ = 0.1 in this embodiment), and the formula for calculating the image height is: By comparing H M and H, it is determined whether the image needs to be expanded in the height direction. If expansion is required, the expanded height H M - H will be evenly distributed on the upper and lower sides of the image.
[0152] Similarly, it can be obtained:
[0153]
[0154] where H target represents the target height, H target= W * r; where W represents the width of the input image; r represents the input ratio; η is the tolerable importance loss rate (λ = 0.1 in this embodiment), and the formula for calculating the image width is: By comparing W′ M and W, it is determined whether the image needs to be expanded in the width direction. If expansion is required, the expanded width W′ M - W will be evenly distributed on both the left and right sides of the image.
[0155] Select W M , H M or H′ M , W′ M as the size W Final , H Final :
[0156] If the input ratio r >> 1, it indicates that width is prioritized, i.e., W Final = W M , and H Final = H M ;
[0157] If the input ratio r < 1, it indicates that height is prioritized, i.e., W Final = W′ M , and H Final = H′ M .
[0158] After determining the final size, delete the unnecessary regions. When the input image contains multiple objects, arrange the objects according to the spacing ratio in the input image; if there is only one object, there is no need to consider the spacing issue.
[0159] Finally, use the content completion module to complete the content of the cropped image and output it to complete the reconstruction.
[0160] The cropped image is completed in content by constructing a content completion module. The diffusion model in the content completion module does not specify a specific model structure. A diffusion model based on StableDiffusion can be used, or a diffusion model based on DiffusionTransformer can also be used.
[0161] The input image is encoded and then input into U-Net (a diffusion model based on StableDiffusion) or Transformer (a diffusion model based on DiffusionTransformer) as a conditional input into the diffusion model to guide the "diffusion model" to generate the blank gaps and the image background content.
[0162] The input ratio is input into the second convolution to generate a feature map, which is then input into a U-Net or a Transformer and used as a conditional input to the diffusion model to guide the "diffusion model" in generating the blank gaps and the image background content.
[0163] The input text is encoded and then input into a U-Net or a Transformer, which is used as a conditional input to the diffusion model to guide the "diffusion model" in generating the blank gaps and the image background content. The guiding role of this "input text" is very significant. It can perform customized output for the blank gaps and the image background, change the occlusion relationship of the objects in the image, specify the importance of the objects or the important parts of the objects themselves, and can also perform scaling according to the importance.
[0164] The content completion module adopts the following loss function:
[0165]
[0166] where t represents a random variable of the time step from 1 to T; T represents the time step of the diffusion process; x 0 represents the original data sample; ∈ represents the noise of the standard normal distribution; x t represents the noise-added data at time step t, which is calculated from x 0 and ∈ according to a specific noise schedule; e θ (x t , t) represents the noise predicted by the model, parameterized as θ, which is used to try to estimate the noise ∈ added to the data at time t; ||·|| 2 represents the L2 norm, that is, the Euclidean distance.
[0167] The objective of this loss function is to make the noise ∈ θ (x t , t) predicted by the model as close as possible to the actual true noise ∈ added at time step t. During the training process, by minimizing this loss, the model learns how to reverse the diffusion process, that is, how to denoise and finally generate new samples similar to the original data.
[0168] Finally, an image that conforms to the input ratio and the input text description is output. The detailed process of this embodiment is as Figure 2 shown.
[0169] Embodiment III
[0170] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the multi-scale reconstruction method of a building landscape and historical relics image described in any one of the above embodiments.
[0171] Figure 3 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0172] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0173] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0174] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0175] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to enable communication and interaction between this device and other devices. The communication module can achieve communication through wired means (such as USB (Universal Serial Bus), network cable, etc.) or wireless means (such as mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).
[0176] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0177] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0178] Embodiment 4
[0179] Based on the same inventive concept, corresponding to any of the above embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute a multi-scale reconstruction method for building landscape and historical relic images as described in any of the above embodiments.
[0180] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0181] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A multi-scale reconstruction method for architectural landscape and historical relics images, characterized in that the steps include: Collect images of architectural landscapes and historical relics to be reconstructed; The importance of the architectural landscape and historical relics images to be reconstructed is analyzed to obtain a binary image with importance information; Performing cropping and alignment on the binary image to obtain a cropped image; The cropped image is supplemented with content and outputted to complete the reconstruction.
2. The multi-scale reconstruction method of architectural landscape and historical relics images according to claim 1 is characterized in that: It is characterized in that The steps of performing the importance analysis include: Where P(x, y) represents the pixel at the image coordinate (x, y); I E (x, y) represents the energy value at the image coordinate (x, y); I E Represents the energy distribution diagram of the image; where α, β and γ represent coefficients; ⊙ represents the pixel-by-pixel dot product; and N represents the number of objects detected by the target in the input image.
3. The multi-scale reconstruction method of architectural landscape and historical relics images according to claim 2 is characterized in that: Calculate α, β and γ according to the input image and input ratio. When the input image width is W, height is H, and the input ratio is a:b, the three coefficients are: β=1-α γ=1.
4. The multi-scale reconstruction method of architectural landscape and historical relics images according to claim 1, characterized in that: It is characterized in that The steps to perform cropping alignment include: Among them, W targt Indicates the target width, W target =H*r; H represents the height of the input image; r represents the input ratio; η represents the tolerable importance loss rate ; The formula for finding the image height is: By comparing H M and H. Among them, H target Indicates the target height, H target =W*r; W represents the width of the input image. The formula for calculating the image width is:
5. The multi-scale reconstruction method of architectural landscape and historical relics images according to claim 4 is characterized in that: Choose W according to the following rules M , H M or H′ M , H′ M As the final image size W Final , H Final : If the input ratio r>>1, it means width priority, that is, W Final =W M , and H Final =H M ; If the input ratio r < 1, it means that the height is prioritized, that is, W Final =W′ M , and H Final =H′ M .
6. The multi-scale reconstruction method of architectural landscape and historical relics images according to claim 1, characterized in that: It is characterized in that The steps to complete the content include: Where t represents a random variable with a time step from 1 to T; T represents the time step of the diffusion process; x0 represents the original data sample; ∈ represents the noise of the standard normal distribution; x t represents the noisy data at time step t; ∈ θ (x t , t) represents the noise predicted by the model, parameterized by θ; ||·|| 2 represents the L2 norm, that is, the Euclidean distance.
7. A multi-scale reconstruction system for architectural landscape and historical relics images, the system being used to implement the method according to any one of claims 1 to 6, characterized in that: include: Acquisition module, analysis module, cropping module and completion module; The acquisition module is used to acquire images of architectural landscapes and historical relics to be reconstructed; The analysis module is used to perform importance analysis on the architectural landscape and historical relics images to be reconstructed, and obtain a binary image with importance information; The cropping module is used to crop and align the binary image to obtain a cropped image; The completion module is used to complete the content of the cropped image and output it to complete the reconstruction.
8. The multi-scale reconstruction system of architectural landscape and historical relics images according to claim 7, characterized in that: The workflow of the analysis module includes: Where P(x, y) represents the pixel at the image coordinate (x, y); I E (x, y) represents the energy value at the image coordinate (x, y); I E Represents the energy distribution diagram of the image; where α, β and γ represent coefficients; ⊙ represents the pixel-by-pixel dot product; and N represents the number of objects detected by the target in the input image.
9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 6 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 6 is implemented.