An ancient building image self-adaptive style transfer method based on a visual encoder

By proposing an adaptive style transfer method for ancient building images based on a visual encoder, a multi-stage feature extraction and fusion method is used with a visual transformer encoder and decoder. This method solves the problem of low automation in the digital depiction of ancient buildings and achieves a more efficient style transfer effect.

CN116309022BActive Publication Date: 2026-03-24HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-08
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Current technologies for digitally depicting ancient buildings rely on human experience and have a low degree of automation, resulting in cluttered creation scenes, irregular shapes, inconsistent textures and colors, and are time-consuming and labor-intensive.

Method used

An adaptive style transfer method for ancient building images based on a visual encoder is adopted. By constructing a dataset, combining a visual transformer encoder and decoder, and utilizing adaptive color and structure adjustment modules, multi-stage feature extraction and fusion are performed to achieve style transfer.

Benefits of technology

It improves the ability to represent color and structural features during style transfer of ancient building images, solves the problems of irregular local structural forms and inconsistent texture colors, and achieves more efficient automated processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116309022B_ABST
    Figure CN116309022B_ABST
Patent Text Reader

Abstract

The application discloses an ancient building image self-adaptive style migration method based on a visual encoder, and relates to a visual coding-decoding style migration neural network. In the feature encoder stage, multi-stage relative position coding features of a building image are obtained, and an adaptive color and structure feature migration fusion module is designed to perform multi-stage and multi-scale fusion on extracted building content features and building style features. The adaptive color and structure feature migration fusion module adopts multiple style feature migration methods to perform feature migration fusion. Then, in the feature decoder stage, the building content features and the building style features are further migrated and fused. The model can predict an ancient building style migration test set, greatly improve the style migration estimation accuracy, solve the problems of building scene description refinement, form standardization, texture and color standardization and the like, and realize building heritage space narrative reproduction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, specifically to an adaptive style transfer method for ancient building images based on a visual encoder. Background Technology

[0002] Ancient architecture is an important part of cultural heritage, possessing significant historical, scientific, and artistic value.

[0003] Digital depiction of ancient architecture, picture book storytelling, picture book re-creation, and digital packaging are important means of conveying ancient architecture and culture. However, current large-scale and dense digital depictions and stylistic narratives of architectural heritage still heavily rely on manual experience, have low automation levels, and suffer from problems such as cluttered creative scenes, non-standard forms, inconsistent textures and colors, and are time-consuming and labor-intensive. These issues place high demands on the technical skills and experience of practitioners. Summary of the Invention

[0004] This invention provides an adaptive style transfer method for ancient building images based on a visual encoder, in order to solve the problem that the digital depiction of ancient buildings in the prior art relies on human experience and has a low degree of automation.

[0005] To achieve the above objectives, the technical solution of the present invention is as follows:

[0006] An adaptive style transfer method for ancient building images based on a visual encoder, characterized by the following steps:

[0007] Step S1: Collect multiple images of ancient buildings, style images, color reference images, structural reference images, and real images for style transfer learning, and construct a dataset using the above five types of images; divide the dataset to obtain a training set and a test set;

[0008] Step S2: Construct a visual encoder-decoder style transfer neural network; the visual encoder-decoder style transfer neural network includes an encoder and a decoder;

[0009] Step S3: Randomly select a set of samples from the training set, use the encoder to encode the ancient building images and style images in the samples, extract the features of the ancient building images and style images at different layers, output their architectural content features and architectural style features, fuse the architectural content features and architectural style features to obtain the fused features.

[0010] Step S4: During the encoding feature extraction process, the fused features and the style image from step S3 are input into the decoder for reconstruction to obtain the image after predicted style transfer.

[0011] Step S5: Based on the style-transferred images and the corresponding style-transfer learning real images in the dataset, iteratively train the visual encoding-decoding style transfer neural network to obtain the trained visual encoding-decoding style transfer neural network.

[0012] Step S6: Test the trained visual encoder-decoder style transfer neural network using a test set.

[0013] Preferably, step S1 specifically includes the following steps:

[0014] Step S11: Collect multiple images of ancient buildings, style images, color reference images, structural reference images, and real images for style transfer learning. The number of these five types of images is equal and they correspond to each other. Group them one by one to construct a dataset. Each group contains five different types of images. The dataset can be represented as: ;

[0015] Where D represents the dataset, Images representing ancient architecture Image representing style, Represents a color reference image. Represents the structural reference image. This represents style transfer learning from real images;

[0016] Step S12: Randomly select images from the dataset in groups and construct training and test sets in an 8:2 ratio.

[0017] Preferably, the encoder in step S2 is a visual transformer encoder, which includes two weighted encoding feature extraction sub-networks: an adaptive color adjustment network and an adaptive structure adjustment network.

[0018] The adaptive color adjustment network contains i sequentially connected adaptive color adjustment modules based on a color reference image. The i adaptive color adjustment modules adjust the colors of features in different dimensions from low dimension to high dimension.

[0019] The adaptive structure adjustment network consists of i sequentially connected adaptive structure adjustment modules based on the structural baseline image. The i adaptive structure adjustment modules perform structural adjustments on features of different dimensions from low dimension to high dimension.

[0020] Preferably, the decoder in step S2 is a visual transformer decoder, which includes i sequentially connected feature decoding modules. The i feature decoding modules decode features of different dimensions sequentially from high dimension to low dimension.

[0021] Preferably, step S3 specifically includes the following steps:

[0022] Step S31: Randomly select a set of samples from the training set and input the selected samples into the adaptive color adjustment network and the adaptive structure adjustment network respectively.

[0023] Step S32: In the adaptive color adjustment network, after... After the adaptive color adjustment module, among which Extracting features from ancient building images With style image features The feature dimension is , For batches of feature extraction, The number of channels for feature extraction. and The width and height of the input image;

[0024] Encoding is embedded at specified locations in a visual encoder-decoder style transfer neural network to incorporate features from ancient building images. and style image features Enter the number In the adaptive color adjustment module, the query features of the ancient building image are output using embedding encoding. Query features of ancient building images Content characteristics of ancient architectural images and query features of style images Query features of style images Content features of style image features The AdaIN function is used for feature transfer fusion, and the fusion process is defined as follows:

[0025]

[0026]

[0027]

[0028]

[0029] in, and These are the mean and variance functions, respectively; and These are the query features and the queryed features of the fused ancient building images, respectively. and These are the query features of the fused style image and the query features of the fused style image, respectively.

[0030] Next, the content features of the ancient building images are fused using the MixAdaIN function through feature transfer. Content features of style image features Feature fusion is performed, and the fused feature is defined as follows:

[0031]

[0032]

[0033] in, and These are the content features of the fused ancient building image and the content features of the fused style image, respectively. , and , These are all parameters of the affine transformation of the adaptive color adjustment network, defined as follows:

[0034]

[0035]

[0036]

[0037]

[0038] in, and These are the mean and variance functions, respectively; It is a constant, and its range of values ​​is 1. ; A color reference image; This indicates cross-attention feature extraction;

[0039] A first self-attention module is created. The fused features are passed through the first self-attention module to output reconstructed features, defined as follows:

[0040]

[0041]

[0042] Among them, the function For the characteristic normalization function, For feature dimension, This represents the transpose of the characteristic matrix. and The first The features of the ancient building image reconstruction output by the first adaptive color adjustment and the first Each adaptive color adjustment output style image reconstruction feature has a dimension of 1. ;

[0043] Step S33: In the adaptive structure adjustment network, after... After an adaptive structural adjustment module, the image features of the ancient building are extracted. With style image features The feature dimension is , ;

[0044] Features of ancient building images With style image features Enter them into the first one respectively In each adaptive structure adjustment module, the query features of the ancient building images are output using embedding encoding. Query features of style images Query features of ancient building images The query features of the style image and the content characteristics of ancient architectural images Content characteristics of style images The feature transfer fusion is performed using the AdaIN function, and the fusion process is defined as follows:

[0045]

[0046]

[0047]

[0048]

[0049] in, , , and These are the query features of the merged ancient building images, the query features of the merged style images, the query features of the merged ancient building images, and the query features of the merged style images, respectively. and These are the variance and mean functions of the feature, respectively;

[0050] Next, the content features of the ancient building images are fused using the Adaptive Structural Feature Transfer Fusion MixAdaIN function. Content features of style images Feature fusion is performed, defined as follows:

[0051]

[0052]

[0053] in, and These are the content features of the fused ancient building image and the content features of the fused style image, respectively. , and , These are all parameters of the feature cross-attention transfer fusion affine transformation, defined as follows:

[0054]

[0055]

[0056]

[0057]

[0058] here, For structural reference images; and It is a function of mean and variance; It is a constant, and its range of values ​​is 1. ; This indicates cross-attention feature extraction;

[0059] A second self-attention module is created. The fused features are processed by the second self-attention module to output reconstructed features, defined as follows:

[0060]

[0061]

[0062] in, and The first The ancient building image reconstruction and fusion features reconstructed from the first adaptive structure adjustment module, and the first... The style image reconstruction fusion features reconstructed from the output of each adaptive structure adjustment module, with each reconstruction feature dimension being [missing information]. ;

[0063] Step S34: Use the visual transformer encoder to... and The mixture is fused and then processed using a visual transformer encoder. The fusion process is defined as follows:

[0064]

[0065]

[0066] in, For convolutional layers; function Indicates feature concatenation operation; and The ancient architecture fusion feature and style fusion feature are respectively proposed by the visual transformer encoder.

[0067] Preferably, step S4 specifically includes the following steps:

[0068] Step S41, after the first Each feature decoding module extracts decoding features from the ancient building image. With style image decoding features The data is then input into the i-th feature decoding module, which uses embedding encoding to output the query features of the ancient building images. Query features of style images Query features of ancient building images Query features of style images Content characteristics of ancient architectural images Content characteristics of style images ;

[0069] A third self-attention module is created, and the reconstructed features are output through the third self-attention module, defined as follows:

[0070]

[0071]

[0072] in, and These are the reconstructed features from the decoding features of ancient building images and the reconstructed features from the decoding features of style images, respectively. The dimensions of the reconstructed features are both... ;

[0073] Step S42: For and Feature transfer fusion is performed, and the feature transfer fusion process is defined as follows:

[0074]

[0075]

[0076] in, This represents the migration fusion features of the ancient building image output by the i-th feature decoding module. This represents the transfer fusion feature of the style image output by the i-th feature decoding module;

[0077] Step S43: Use the visual transformer decoder to... The grid transfer estimation output module is defined as follows: The module processes and outputs the image after predicted style transfer.

[0078]

[0079] in, It is a convolutional layer; To predict the style-transferred image, the dimensions are the same as the input image.

[0080] Preferably, step S5 specifically includes the following steps:

[0081] Step S51: Calculate the content feature loss function based on the generated predicted style-transferred images. Style feature loss function Semantic feature loss function and style reconstruction loss function ;

[0082] Step S52: Determine the content feature loss function respectively. Style feature loss function Semantic feature loss function and style reconstruction loss function Supervised training weights 1. 2. 3 and 4;

[0083] Step S53: Establish the total loss function Repeat steps S3 to S5, iteratively training to minimize the total loss function until 50 epochs of training are completed or the loss value is less than 1. .

[0084] Preferably, the content feature loss function in step S51 Style feature loss function Semantic feature loss function and style reconstruction loss function Specifically:

[0085] Among them, the content feature loss function The definition is as follows:

[0086]

[0087] in, and For the first The layer encoder feature extraction module extracts architectural content image features and style features; ; Reconstructing images for style transfer; function for Norm; function This is a channel normalization operation for the average variance. This indicates the number of layers in the feature encoder module, which is set to 4 in this example;

[0088] Style feature loss function The definition is as follows:

[0089]

[0090] in, and For the first The layer encoder extracts architectural content image features and style features. and It is a function of mean and variance. This indicates the number of feature decoder module layers, which is set to 4 in this example;

[0091] Semantic feature loss function The definition is as follows:

[0092]

[0093] Style Reconstruction Loss Function The definition is as follows:

[0094]

[0095] Preferably, in step S52 , and Satisfy the following formula:

[0096] ,and , , .

[0097] Preferably, the total loss function in step S53 is:

[0098] .

[0099] Beneficial effects:

[0100] Compared with the prior art, the advantages of this invention are mainly reflected in the following aspects:

[0101] (1) This invention proposes a visual encoding-decoding style transfer neural network, which can fully learn the relative position texture and structural feature information of architectural images, greatly improve the feature representation ability of color and structure during style transfer, and solve problems such as irregular local structural morphology, inconsistent texture and color during architectural image style transfer.

[0102] (2) The visual encoding-decoding style transfer neural network uses a multi-stage extraction of shallow and deep features of the image, progressive global context fusion of architectural and style image features, and progressive encoding and reconstruction of the fused features.

[0103] (3) The present invention designs an adaptive color adjustment module, which is integrated into the visual encoding-decoding style transfer neural network and can use an adaptive global supervised transfer method to fuse style features and color reference features.

[0104] (4) The present invention designs an adaptive structure adjustment module, which is integrated into the visual encoding-decoding style transfer neural network and can use an adaptive global supervision method to fuse building structure features and structural baseline features. Attached Figure Description

[0105] Figure 1 This is a flowchart of the method described in this invention;

[0106] Figure 2 The above is a schematic diagram of a portion of the ancient architectural style transfer and imitation learning dataset according to an embodiment of the present invention. Figure (a) is an architectural image, and Figure (b) is a style image.

[0107] Figure 3 This is a structural diagram of the visual encoder-decoder style transfer neural network used in this invention;

[0108] Figure 4 Figures (a) to (c) represent the modules of the visual encoder-decoder style transfer neural network, respectively. The modules are the adaptive color adjustment network, the adaptive structure adjustment network, and the visual transformer decoder. Detailed Implementation

[0109] The present invention will now be further described in conjunction with the accompanying drawings and embodiments.

[0110] like Figure 1 The diagram shows a flowchart of the method described in this invention. An adaptive style transfer method for ancient building images based on a visual encoder includes the following steps:

[0111] Step S1: Collect multiple images of ancient buildings, style images, color reference images, structural reference images, and real images for style transfer learning, and construct a dataset using the above five types of images; divide the dataset to obtain a training set and a test set;

[0112] Step S2: Construct a visual encoder-decoder style transfer neural network; the visual encoder-decoder style transfer neural network includes an encoder and a decoder;

[0113] Step S3: Randomly select a set of samples from the training set, use the encoder to encode the ancient building images and style images in the samples, extract the features of the ancient building images and style images at different layers, output their architectural content features and architectural style features, fuse the architectural content features and architectural style features to obtain the fused features.

[0114] Step S4: During the encoding feature extraction process, the fused features and the style image from step S3 are input into the decoder for reconstruction to obtain the image after predicted style transfer.

[0115] Step S5: Based on the predicted style-transfer images and the corresponding real images of style transfer learning in the dataset, iteratively train the visual encoding-decoding style transfer neural network to obtain the trained visual encoding-decoding style transfer neural network.

[0116] Step S6: Test the trained visual encoder-decoder style transfer neural network using a test set.

[0117] In this embodiment, step S1 specifically includes the following steps:

[0118] Step S11: Collect multiple images of ancient buildings, style images, color reference images, structural reference images, and real images for style transfer learning. The number of these five types of images is equal and they correspond to each other. Group them one by one to construct a dataset. Each group contains five different types of images. The dataset can be represented as: ;

[0119] Where D represents the dataset, Images representing ancient architecture Image representing style, Represents a color reference image. Represents the structural reference image. This represents style transfer learning from real images;

[0120] Specifically, images of ancient buildings were taken from different angles in multiple ancient building complexes. It should be noted that this example uses ancient buildings such as Taiping Street and Tongguan Kiln as shooting scenes, but the ancient building complexes used in the dataset constructed in this example are not limited to these. Figure 2 To capture scene images, images of the architectural style to be learned, color baseline images, and structural baseline images were collected. Based on these images, professionals were invited to depict real images for style transfer learning.

[0121] Step S12: Randomly select images from the dataset in groups and construct training and test sets in an 8:2 ratio.

[0122] In this embodiment, the encoder in step S2 is specifically a visual transformer encoder, which includes two weighted encoding feature extraction sub-networks: an adaptive color adjustment network and an adaptive structure adjustment network.

[0123] The adaptive color adjustment network contains i sequentially connected adaptive color adjustment modules based on a color reference image. The i adaptive color adjustment modules adjust the colors of features in different dimensions from low dimension to high dimension.

[0124] The adaptive structure adjustment network consists of i sequentially connected adaptive structure adjustment modules based on the structural baseline image. The i adaptive structure adjustment modules perform structural adjustments on features of different dimensions from low dimension to high dimension.

[0125] Specifically, adaptive color adjustment networks and adaptive structure adjustment networks, such as Figure 4 As shown in (a) and (b), the feature transfer fusion of the adaptive color adjustment module is as follows: the output features of the (i-1)th adaptive color adjustment module are combined with the ancient building image features. With style image features Enter them into the first one respectively In each adaptive color adjustment module, the query features of the ancient building images are output using a feature local location embedding encoding layer. Query features of style images Query features of ancient building images The query features of the style image and the content characteristics of ancient architectural images Content characteristics of style images The features of ancient building images are obtained by using attention transfer fusion of the AdaIN function and the MixAdaIN function through feature transfer fusion. With style image features Similarly, the feature transfer fusion of the adaptive structure adjustment module is as follows: the output features of the (i-1)th adaptive structure adjustment module are combined with the ancient building image features. With style image features Enter them into the first one respectively In each adaptive structure adjustment module, the query features of the ancient building images are output by embedding the encoding layer at the local feature location. Query features of style images Query features of ancient building images The query features of the style image and the content characteristics of ancient architectural images Content characteristics of style images The features of ancient building images are obtained by using attention transfer fusion of the AdaIN function and the MixAdaIN function through feature transfer fusion. With style image features .

[0126] In this embodiment, the decoder in step S2 is specifically a visual transformer decoder. The visual transformer decoder includes i sequentially connected feature decoding modules, and the i feature decoding modules decode features of different dimensions sequentially from high dimension to low dimension.

[0127] In this embodiment, step S3 specifically includes the following steps:

[0128] Step S31: Randomly select a set of samples from the training set and input the selected samples into the adaptive color adjustment network and the adaptive structure adjustment network respectively.

[0129] In the training set, one image of an ancient building to be style transferred, one style image as the style source, one color reference image, and one structural reference image are selected as inputs to the adaptive color adjustment network and the adaptive structure adjustment network. All input images are of the same size. , The width of the image. The height of the image;

[0130] Step S32: In the adaptive color adjustment network, after... After the adaptive color adjustment module, among which Extracting features from ancient building images With style image features The feature dimension is , For batches of feature extraction, The number of channels for feature extraction. and The width and height of the input image;

[0131] Encoding is embedded at specified locations in a visual encoder-decoder style transfer neural network to incorporate features from ancient building images. and style image features Enter the number In the adaptive color adjustment module, the query features of the ancient building image are output using embedding encoding. Query features of ancient building images Content characteristics of ancient architectural images and query features of style images Query features of style images Content features of style image features The AdaIN function is used for feature transfer fusion, and the fusion process is defined as follows:

[0132]

[0133]

[0134]

[0135]

[0136] in, and These are the mean and variance functions, respectively; and These are the query features and the queryed features of the fused ancient building images, respectively. and These are the query features of the fused style image and the query features of the fused style image, respectively.

[0137] Next, the content features of the ancient building images are fused using the MixAdaIN function through feature transfer. Content features of style image features Feature fusion is performed, and the fused feature is defined as follows:

[0138]

[0139]

[0140] in, and These are the content features of the fused ancient building image and the content features of the fused style image, respectively. , and , These are all parameters of the affine transformation of the adaptive color adjustment network, defined as follows:

[0141]

[0142]

[0143]

[0144]

[0145] in, and These are the mean and variance functions, respectively; It is a constant, and its range of values ​​is 1. ; A color reference image; This indicates cross-attention feature extraction;

[0146] A first self-attention module is created. The fused features are passed through the first self-attention module to output reconstructed features, defined as follows:

[0147]

[0148]

[0149] Among them, the function For the characteristic normalization function, For feature dimension, This represents the transpose of the characteristic matrix. and The first The features of the ancient building image reconstruction output by the adaptive color adjustment and the first The style image reconstruction features are generated from adaptive color adjustment outputs, and the reconstruction feature dimensions are all... ;

[0150] Step S33: In the adaptive structure adjustment network, after... After the adaptive structure adjustment module, the image features of the ancient building are extracted. With style image features The feature dimension is , ;

[0151] Features of ancient building images With style image features Enter them into the first one respectively In each adaptive structure adjustment module, the query features of the ancient building images are output using embedding encoding. Query features of style images Query features of ancient building images The query features of the style image and the content characteristics of ancient architectural images Content characteristics of style images The feature transfer fusion is performed using the AdaIN function, and the fusion process is defined as follows:

[0152]

[0153]

[0154]

[0155]

[0156] in, , , and These are the query features of the merged ancient building images, the query features of the merged style images, the query features of the merged ancient building images, and the query features of the merged style images, respectively. and These are the variance and mean functions of the feature, respectively;

[0157] Next, the content features of the ancient building images are fused using the Adaptive Structural Feature Transfer Fusion MixAdaIN function. Content features of style images Feature fusion is performed, defined as follows:

[0158]

[0159]

[0160] in, and These are the content features of the fused ancient building image and the content features of the fused style image, respectively. , and , These are all parameters of the feature cross-attention transfer fusion affine transformation, defined as follows:

[0161]

[0162]

[0163]

[0164]

[0165] here, For structural reference images; and It is a function of mean and variance; It is a constant, and its range of values ​​is 1. ; This indicates cross-attention feature extraction;

[0166] A second self-attention module is created. The fused features are processed by the second self-attention module to output reconstructed features, defined as follows:

[0167]

[0168]

[0169] in, and The first The ancient building image reconstruction and fusion features reconstructed from the first adaptive structure adjustment module, and the first... The style image reconstruction fusion features reconstructed from the output of each adaptive structure adjustment module, with each reconstruction feature dimension being [missing information]. ;

[0170] Step S34: Use the visual transformer encoder to... and The mixture is fused and then processed using a visual transformer encoder. The fusion process is defined as follows:

[0171]

[0172]

[0173] in, For convolutional layers; function Indicates feature concatenation operation; and The ancient architecture fusion feature and style fusion feature are respectively proposed by the visual transformer encoder.

[0174] In this embodiment, step S4 specifically includes the following steps:

[0175] Step S41, after the first Each feature decoding module extracts decoding features from the ancient building image. With style image decoding features The data is then input into the i-th feature decoding module, which uses embedding encoding to output the query features of the ancient building images. Query features of style images Query features of ancient building images Query features of style images Content characteristics of ancient architectural images Content characteristics of style images ;

[0176] A third self-attention module is created, and the reconstructed features are output through the third self-attention module, defined as follows:

[0177]

[0178]

[0179] in, and These are the reconstructed features from the decoding features of ancient building images and the reconstructed features from the decoding features of style images, respectively. The dimensions of the reconstructed features are both... ;

[0180] Step S42: For and Feature transfer fusion is performed, and the feature transfer fusion process is defined as follows:

[0181]

[0182]

[0183] in, This represents the migration fusion features of the ancient building image output by the i-th feature decoding module. This represents the transfer fusion feature of the style image output by the i-th feature decoding module;

[0184] Step S43: Use the visual transformer decoder to... The grid transfer estimation output module is defined as follows: The module processes and outputs the image after predicted style transfer.

[0185]

[0186] in, It is a convolutional layer; To predict the style-transferred image, the dimensions are the same as the input image. See the visual transformer decoder. Figure 4 (c);

[0187] Preferably, step S5 specifically includes the following steps:

[0188] Step S51: Calculate the content feature loss function based on the generated predicted style-transferred images. Style feature loss function Semantic feature loss function and style reconstruction loss function ;

[0189] Step S52: Determine the content feature loss function respectively. Style feature loss function Semantic feature loss function and style reconstruction loss function Supervised training weights 1. 2. 3 and 4;

[0190] Step S53: Establish the total loss function Repeat steps S3 to S5, iteratively training to minimize the total loss function until 50 epochs of training are completed or the loss value is less than 1. .

[0191] Preferably, the content feature loss function in step S51 Style feature loss function Semantic feature loss function and style reconstruction loss function Specifically:

[0192] Among them, the content feature loss function The definition is as follows:

[0193]

[0194] in, and For the first The layer encoder feature extraction module extracts architectural content image features and style features; ; Reconstructing images for style transfer; function for Norm; function This is a channel normalization operation for the average variance. This indicates the number of layers in the feature encoder module, which is set to 4 in this example;

[0195] Style feature loss function The definition is as follows:

[0196]

[0197] in, and For the first The layer encoder extracts architectural content image features and style features. and It is a function of mean and variance. This indicates the number of feature decoder module layers, which is set to 4 in this example;

[0198] Semantic feature loss function The definition is as follows:

[0199]

[0200] Style Reconstruction Loss Function The definition is as follows:

[0201]

[0202] Preferably, in step S52 , and Satisfy the following formula:

[0203] ,and , , .

[0204] Preferably, the total loss function in step S53 is:

[0205] .

[0206] Specifically, network training uses the backpropagation algorithm, and the training optimizer uses the Adam algorithm, where the optimizer parameters... , The network training epoch is 50, and the initial learning rate is set to... And after training to 20, At 40 epochs, the learning rate is reduced to its original value. The network model training software platform is PyTorch.

Claims

1. A method for adaptive style transfer of ancient architectural images based on a visual encoder, characterized in that, include: Step S1: Collect multiple images of ancient buildings, style images, color reference images, structural reference images, and real images for style transfer learning, and construct a dataset using the above five types of images; divide the dataset to obtain a training set and a test set; Step S2: Construct a visual encoding-decoding style transfer neural network; Visual encoder-decoder style transfer neural networks consist of an encoder and a decoder; Step S3: Randomly select a set of samples from the training set, use the encoder to encode the ancient building images and style images in the samples, extract the features of the ancient building images and style images at different layers, output their architectural content features and architectural style features, fuse the architectural content features and architectural style features to obtain the fused features. Step S4: During the process of the encoder extracting features from the ancient building image and style image at different layers, the fused features and the style image from step S3 are input into the decoder for reconstruction to obtain the image after predicted style transfer. Step S5: Based on the predicted style-transfer images and the corresponding real images of style transfer learning in the dataset, iteratively train the visual encoding-decoding style transfer neural network to obtain the trained visual encoding-decoding style transfer neural network. Step S6: Test the trained visual encoder-decoder style transfer neural network using the test set; The encoder in step S2 is specifically a visual transformer encoder, which includes two weighted encoding feature extraction sub-networks: an adaptive color adjustment network and an adaptive structure adjustment network. The adaptive color adjustment network contains i sequentially connected adaptive color adjustment modules based on a color reference image. The i adaptive color adjustment modules adjust the colors of features in different dimensions from low dimension to high dimension. The adaptive structure adjustment network consists of i sequentially connected adaptive structure adjustment modules based on the structural baseline image. The i adaptive structure adjustment modules perform structural adjustments on features of different dimensions from low dimension to high dimension.

2. The adaptive style transfer method for ancient building images according to claim 1, characterized in that, Step S1 specifically includes the following steps: Step S11: Collect multiple images of ancient buildings, style images, color reference images, structural reference images, and real images for style transfer learning. The number of these five types of images is equal and they correspond to each other. Group them one by one to construct a dataset. Each group contains five different types of images. The dataset can be represented as: ; Where D represents the dataset, Images representing ancient architecture Image representing style, Represents a color reference image. Represents the structural reference image. This represents style transfer learning from real images; Step S12: Randomly select images from the dataset in groups and construct training and test sets in an 8:2 ratio.

3. The adaptive style transfer method for ancient building images according to claim 2, characterized in that, The visual transformer decoder in step S2 includes i sequentially connected feature decoding modules, which decode features of different dimensions from high dimension to low dimension.

4. The adaptive style transfer method for ancient building images according to claim 3, characterized in that, Step S3 specifically includes the following steps: Step S31: Randomly select a set of samples from the training set and input the selected samples into the adaptive color adjustment network and the adaptive structure adjustment network respectively. Step S32: In the adaptive color adjustment network, after... After the adaptive color adjustment module, among which Extracting features from ancient building images With style image features The feature dimension is , For batches of feature extraction, The number of channels for feature extraction. and The width and height of the input image; Specifically, the image features of ancient buildings and style image features Enter the number In the adaptive color adjustment module, the query features of the ancient building image are embedded in the encoding layer using local feature locations. Query features of ancient building images Content characteristics of ancient architectural images and query features of style images Query features of style images Content features of style image features The AdaIN function is used for feature transfer fusion, and the fusion process is defined as follows: in, and These are the mean and variance functions, respectively; and These are the query features and the queryed features of the fused ancient building images, respectively. and These are the query features of the fused style image and the query features of the fused style image, respectively. Next, the content features of the ancient building images are fused using the MixAdaIN function through feature transfer. Content features of style image features Feature fusion is performed, and the fused feature is defined as follows: in, and These are the content features of the fused ancient building image and the content features of the fused style image, respectively. , and , These are all parameters of the affine transformation of the adaptive color adjustment network, defined as follows: in, and These are the mean and variance functions, respectively; It is a constant, and its range of values ​​is 1. ; A color reference image; This indicates cross-attention feature extraction; A first self-attention module is created. The fused features are passed through the first self-attention module to output reconstructed features, defined as follows: Among them, the function For the characteristic normalization function, For feature dimension, This represents the transpose of the characteristic matrix. and The first The features of the ancient building image reconstruction output by the first adaptive color adjustment and the first Each adaptive color adjustment output style image reconstruction feature has a dimension of 1. ; Step S33: In the adaptive structure adjustment network, after... After an adaptive structural adjustment module, the image features of the ancient building are extracted. With style image features The feature dimension is ; Features of ancient building images With style image features Enter them into the first one respectively In each adaptive structure adjustment module, the query features of the ancient building images are output by embedding the encoding layer at the local feature location. Query features of style images Query features of ancient building images The query features of style images and the content characteristics of ancient architectural images Content characteristics of style images The feature transfer fusion is performed using the AdaIN function, and the fusion process is defined as follows: in, , , and These are the query features of the merged ancient building images, the query features of the merged style images, the query features of the merged ancient building images, and the query features of the merged style images, respectively. and These are the variance and mean functions of the feature, respectively; Next, the content features of the ancient building images are fused using the Adaptive Structural Feature Transfer Fusion MixAdaIN function. Content features of style images Feature fusion is performed, defined as follows: in, and These are the content features of the fused ancient building image and the content features of the fused style image, respectively. , and , These are all parameters of the feature cross-attention transfer fusion affine transformation, defined as follows: here, For structural reference images; and It is a function of mean and variance; It is a constant, and its range of values ​​is 1. ; This indicates cross-attention feature extraction; A second self-attention module is created. The fused features are processed by the second self-attention module to output reconstructed features, defined as follows: in, and The first The ancient building image reconstruction and fusion features reconstructed from the first adaptive structure adjustment module, and the first... The style image reconstruction fusion features reconstructed from the output of each adaptive structure adjustment module, with each reconstruction feature dimension being [missing information]. ; Step S34: Use the visual transformer encoder to... and The mixture is fused and then processed using a visual transformer encoder. The fusion process is defined as follows: in, For convolutional layers; function Indicates feature concatenation operation; and The ancient architecture fusion feature and style fusion feature are respectively proposed by the visual transformer encoder.

5. The adaptive style transfer method for ancient building images according to claim 4, characterized in that, Step S4 specifically includes the following steps: Step S41, after the first Each feature decoding module extracts decoding features from the ancient building image. With style image decoding features The data is then input into the i-th feature decoding module, which uses embedding encoding to output the query features of the ancient building images. Query features of style images Query features of ancient building images Query features of style images Content characteristics of ancient architectural images Content characteristics of style images ; A third self-attention module is created, and the reconstructed features are output through the third self-attention module, defined as follows: in, and These are the reconstructed features from the decoding features of ancient building images and the reconstructed features from the decoding features of style images, respectively. The dimensions of the reconstructed features are both... ; Step S42: For and Feature transfer fusion is performed, and the feature transfer fusion process is defined as follows: in, This represents the migration fusion features of the ancient building image output by the i-th feature decoding module. This represents the transfer fusion feature of the style image output by the i-th feature decoding module; Step S43: Use the visual transformer decoder to... The process is performed and the predicted style-transferred image is output. The style transfer estimation output module is defined as follows: in, It is a convolutional layer; To predict the style-transferred image, the dimensions are the same as the input image.

6. The adaptive style transfer method for ancient building images according to claim 5, characterized in that, Step S5 specifically includes the following steps: Step S51: Calculate the content feature loss function based on the generated predicted style-transferred images. Style feature loss function Semantic feature loss function and style reconstruction loss function ; Step S52: Determine the content feature loss function respectively. Style feature loss function Semantic feature loss function and style reconstruction loss function Supervised training weights 1.

2. 3 and 4; Step S53: Establish the total loss function Repeat steps S3 to S5, iteratively training to minimize the total loss function until 50 epochs of training are completed or the loss value is less than 1. .

7. The adaptive style transfer method for ancient building images according to claim 6, characterized in that, The content feature loss function in step S51 Style feature loss function Semantic feature loss function and style reconstruction loss function Specifically: Among them, the content feature loss function The definition is as follows: in, and For the first The layer encoder feature extraction module extracts architectural content image features and style features; ; Reconstructing images for style transfer; function for Norm; function This is a channel normalization operation for the average variance. This indicates the number of layers in the feature encoder module, set to 4. Style feature loss function The definition is as follows: in, and For the first The layer encoder extracts architectural content image features and style features. and It is a function of mean and variance. This indicates the number of layers in the feature decoder module, set to 4. Semantic feature loss function The definition is as follows: Style Reconstruction Loss Function The definition is as follows: 。 8. The adaptive style transfer method for ancient building images according to claim 7, characterized in that, In step S52 , and Satisfy the following formula: ,and , , .

9. The adaptive style transfer method for ancient building images according to claim 7 or 8, characterized in that, The total loss function in step S53 is: .

Citation Information

Patent Citations

  • Image style migration method combining meta-learning mechanism and feature fusion

    CN111325681A

  • Arbitrary style migration method and system, storage medium, computer equipment and terminal

    CN114240735A