Text-guided image style migration method and device based on global attention and fourth-generation deformable convolution, and readable storage medium

By constructing a U-shaped network using global attention and fourth-generation deformable convolution, and combining it with a multi-objective collaborative loss function to optimize the model, the problem of poor geometric adaptability of traditional convolution kernels is solved, achieving high-precision text-guided image style transfer. The generated images retain both content structure and accurately match text style.

CN121236769APending Publication Date: 2025-12-30GUANGDONG UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511457671.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

In existing text-guided image style transfer techniques, traditional convolutional kernels have poor geometric adaptability and insufficient alignment between global and local features, resulting in image distortion after style transfer. They also lack global semantic guidance and are prone to semantic misalignment between style features and text descriptions.

Method used

A U-shaped network is constructed using global attention and fourth-generation deformable convolution. The convolution kernel is dynamically adjusted by generating a channel-space dual-domain weight matrix. The model is trained by combining a multi-objective collaborative composite loss function to optimize the style transfer process and achieve multi-granular fusion of style features and content structure.

Benefits of technology

It improves the accuracy and multi-view consistency of image style transfer, avoids image distortion after style transfer, and generates images that retain content structure and accurately match text style.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121236769A_ABST
    Figure CN121236769A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the field of computer vision, and provides a text-guided image style migration method and device based on global attention and fourth-generation deformable convolution and a readable storage medium, and the method comprises the steps: obtaining a style description text, and extracting semantic features in the style description text; obtaining an initial image to be subjected to style conversion, and extracting multi-level visual features in the initial image; constructing a text-guided image style migration model based on a U-shaped network based on global attention and fourth-generation deformable convolution, and training the text-guided image style migration model according to a pre-designed multi-target collaborative composite loss function; and inputting the semantic features and the multi-level visual features into a trained text-guided image style migration model to output a style-migrated target image associated with the initial image. By adopting the method, the effect that the text guides the image to perform style migration can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a text-guided image style transfer method, apparatus and readable storage medium based on global attention and fourth-generation deformable convolution. Background Technology

[0002] In the field of computer vision technology, with the rapid development of deep learning technology, image style transfer, as an important research direction in digital art creation and computer vision, has gradually evolved from the traditional transfer mode based on reference images to a text-guided cross-modal generation mode. Users can directly control the image style through natural language descriptions without relying on preset style images, which greatly expands the application scenarios.

[0003] However, in existing text-guided style transfer techniques, the sampling position of traditional convolutional kernels is fixed and cannot be dynamically adjusted to match the irregular geometric shape of the content image (such as curved edges and deformed objects), resulting in local structural distortion after style transfer. Moreover, existing methods focus more on local texture transfer and lack global semantic guidance, which easily leads to the problem of semantic misalignment between style features and text descriptions, such as "bold outlines" not being accurately transferred to the target area.

[0004] In summary, due to the poor geometric adaptability of traditional convolutional kernels and insufficient alignment of global and local features, current text-guided style transfer techniques suffer from image distortion after style transfer, resulting in poor performance of text-guided image style transfer. Summary of the Invention

[0005] This application provides a method, apparatus, and readable storage medium for style transfer of text-guided images based on global attention and fourth-generation deformable convolution, which can improve the style transfer effect of text-guided images.

[0006] In a first aspect, embodiments of this application provide a text-guided image style transfer method based on global attention and fourth-generation deformable convolution, including:

[0007] The style description text is obtained, and semantic features are extracted from the style description text based on a text encoder and a pre-trained converter model for contrastive language images.

[0008] An initial image to be style-transformed is obtained, and multi-level visual features are extracted from the initial image based on an encoder of a 19-layer network of visual geometry group.

[0009] A text-guided image style transfer model based on a U-shaped network is constructed based on global attention and fourth-generation deformable convolution. The text-guided image style transfer model is trained according to a pre-designed multi-objective collaborative composite loss function to obtain the trained text-guided image style transfer model.

[0010] The semantic features and the multi-level visual features are input into the trained text-guided image style transfer model to output the style-transferred target image associated with the initial image.

[0011] Secondly, embodiments of this application provide a text-guided image style transfer apparatus based on global attention and fourth-generation deformable convolution, which has the function of implementing the text-guided image style transfer method based on global attention and fourth-generation deformable convolution provided in the first aspect above. The function can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above function, and the modules can be software and / or hardware.

[0012] In one possible design, the device includes:

[0013] The text acquisition module is used to acquire style description text and extract semantic features from the style description text based on a text encoder and a pre-trained converter model for contrastive language images.

[0014] The initial image acquisition module is used to acquire the initial image to be style-transformed, and extracts multi-level visual features from the initial image based on the encoder of the 19-layer visual geometry network.

[0015] The model building and training module is used to build a text-guided image style transfer model based on a U-shaped network based on global attention and fourth-generation deformable convolution. The text-guided image style transfer model is trained according to a pre-designed multi-objective collaborative composite loss function to obtain the trained text-guided image style transfer model.

[0016] The target image output module is used to input the semantic features and the multi-level visual features into the trained text-guided image style transfer model to output the style-transferred target image associated with the initial image.

[0017] Another aspect of this application provides a text-guided image style transfer apparatus based on global attention and fourth-generation deformable convolution, which includes at least one connected processor and memory, wherein the memory is used to store program code, and the processor is used to call the program code in the memory to execute the methods described in the above aspects.

[0018] In another aspect, this application provides a computer storage medium including instructions that, when executed on a computer, cause the computer to perform the methods described in the above aspects.

[0019] In another aspect, this application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods described in the above aspects.

[0020] Compared to traditional style transfer methods in existing technologies, this application's embodiments first acquire style description text and extract semantic features, acquire the initial image to be style transferred, and extract multi-level visual features. Then, a text-guided image style transfer model based on a U-shaped network is constructed based on global attention and fourth-generation deformable convolution. The model is trained according to a pre-designed multi-objective collaborative composite loss function. Finally, the semantic features and multi-level visual features are input into the trained model to output the target image after style transfer. A channel-space dual-domain weight matrix is ​​generated through global attention to guide the fourth-generation deformable convolution to dynamically adjust the sampling position of the convolution kernel, realizing the priority injection of semantically relevant regions of style features. The model is optimized by combining the multi-objective collaborative composite loss function, ultimately generating an image that retains the content structure and accurately matches the text style, achieving multi-granular fusion of style features and content structure, improving transfer accuracy and multi-view style consistency. It overcomes the problems of poor geometric adaptability of traditional convolution kernels and insufficient alignment of global and local features, avoids the problem of image distortion after style transfer, and improves the effect of text-guided image style transfer. Attached Figure Description

[0021] Figure 1 This is an application environment diagram of a text-guided image style transfer method based on global attention and fourth-generation deformable convolution in one embodiment;

[0022] Figure 2 This is a flowchart illustrating a text-guided image style transfer method based on global attention and fourth-generation deformable convolution in one embodiment.

[0023] Figure 3 This is a diagram of a style transfer network framework in one embodiment.

[0024] Figure 4 This is a schematic diagram of the improved U-net network structure in one embodiment;

[0025] Figure 5 This is a schematic diagram of a deformable convolutional framework that incorporates global attention in one embodiment;

[0026] Figure 6This is a flowchart illustrating a text-guided image style transfer method based on global attention and fourth-generation deformable convolution, as described in another embodiment.

[0027] Figure 7 A comparison chart of iterative optimization of the loss function in one embodiment;

[0028] Figure 8 This is a structural block diagram of a text-guided image style transfer device based on global attention and fourth-generation deformable convolution in one embodiment.

[0029] Figure 9 This is an internal structure diagram of a text-guided image style transfer device based on global attention and fourth-generation deformable convolution in one embodiment.

[0030] Figure 10 This is an internal structure diagram of a text-guided image style transfer device based on global attention and fourth-generation deformable convolution in another embodiment. Detailed Implementation

[0031] The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or devices. The module divisions appearing in the embodiments of this application are merely logical divisions; in actual applications, there may be other division methods. For example, multiple modules may be combined into or integrated into another system, or some features may be ignored or not performed. Furthermore, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces, and the indirect couplings or communication connections between modules may be electrical or other similar forms. These are not limited in the embodiments of this application. Moreover, modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed across multiple circuit modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the embodiments of this application.

[0032] This application provides a text-guided image style transfer method based on global attention and fourth-generation deformable convolution, which can be applied to, for example... Figure 1 In the application scenario shown, terminal 102 communicates with server 104 via a network.

[0033] The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle systems. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0034] It should be specifically noted that the terminal 102 involved in this application embodiment can be a wired terminal or a wireless terminal. It can be a device that provides voice and / or data connectivity to a user, a handheld device with wireless connectivity, or other processing devices connected to a wireless modem. The wireless terminal can communicate with one or more core networks via a Radio Access Network (RAN). The wireless terminal can be a mobile terminal, such as a mobile phone (or "cellular" phone) and a computer with a mobile terminal. For example, it can be a portable, pocket-sized, handheld, computer-embedded, or vehicle-mounted mobile device that exchanges voice and / or data with the radio access network. Examples include Personal Communication Service (PCS) phones, cordless phones, Session Initiation Protocol (SIP) phones, Wireless Local Loop (WLL) stations, and Personal Digital Assistants (PDAs). Wireless terminals can also be referred to as systems, subscriber units, subscriber stations, mobile stations, mobile stations, remote stations, access points, remote terminals, access terminals, user terminals, terminal equipment, user agents, user devices, or user equipment.

[0035] Please refer to Figure 2 The following describes a text-guided image style transfer method based on global attention and fourth-generation deformable convolution, provided by embodiments of this application. The embodiments of this application mainly include:

[0036] S201, Obtain style description text. Based on the contrastive language image pre-trained text encoder and pre-trained converter model, extract semantic features from the style description text.

[0037] The style description text is the text content used to guide style transfer. The text encoder of Contrastive Language-Image Pretraining (CLIP) is used to encode the text data. The pre-trained transformer model can be a publicly available pre-trained large model, such as the ViT-B / 32 model. The semantic features are features extracted from the style description text.

[0038] S202: Obtain the initial image to be style-transformed, and extract multi-level visual features from the initial image using an encoder based on a 19-layer network of visual geometry.

[0039] The initial image is the image to be style-transformed. The encoder of the Visual Geometry Group 19-layer network (VGG19) is used to encode the image data, and the multi-level visual features are features extracted from the initial image.

[0040] S203. A text-guided image style transfer model based on a U-shaped network is constructed based on global attention and fourth-generation deformable convolution. The text-guided image style transfer model is trained according to a pre-designed multi-objective collaborative composite loss function to obtain a trained text-guided image style transfer model.

[0041] Global Attention Module (GAM) is a dual attention mechanism that combines channel attention and spatial attention, enhancing the representation of key features through global information modeling.

[0042] Among them, the fourth-generation deformable convolutional network (DCNv4) is an improved version of deformable convolution. Its core lies in the ability of the convolutional kernel to adaptively adjust the sampling position through a dynamic offset-aware mechanism, so as to better capture the local deformation features in the image.

[0043] Among them, the U-Net is an encoder-decoder architecture that fuses high- and low-level features through skip connections.

[0044] Among them, the multi-objective collaborative composite loss function is a comprehensive loss function that combines multiple loss functions designed in this embodiment, and is used to train the model.

[0045] Among them, the text-guided image style transfer model is an improved U-Net network structure, which can be denoted as Style_Unet. It is a dual-path architecture with multi-scale deformation perception and global semantic enhancement, which is constructed by the collaborative design of DCNv4 deformable convolutional module and GAM global attention module.

[0046] S204 inputs semantic features and multi-level visual features into the trained text-guided image style transfer model to output the style-transferred target image associated with the initial image.

[0047] The target image is the image obtained after style transfer.

[0048] Compared to traditional style transfer methods in existing technologies, this application's embodiments first acquire style description text and extract semantic features, acquire the initial image to be style transferred, and extract multi-level visual features. Then, a text-guided image style transfer model based on a U-shaped network is constructed based on global attention and fourth-generation deformable convolution. The model is trained according to a pre-designed multi-objective collaborative composite loss function. Finally, the semantic features and multi-level visual features are input into the trained model to output the target image after style transfer. A channel-space dual-domain weight matrix is ​​generated through global attention to guide the fourth-generation deformable convolution to dynamically adjust the sampling position of the convolution kernel, realizing the priority injection of semantically relevant regions of style features. The model is optimized by combining the multi-objective collaborative composite loss function, ultimately generating an image that retains the content structure and accurately matches the text style, achieving multi-granular fusion of style features and content structure, improving transfer accuracy and multi-view style consistency. It overcomes the problems of poor geometric adaptability of traditional convolution kernels and insufficient alignment of global and local features, avoids the problem of image distortion after style transfer, and improves the effect of text-guided image style transfer.

[0049] Optionally, in some embodiments of this application, inputting semantic features and multi-level visual features into a trained text-guided image style transfer model to output a style-transferred target image associated with the initial image includes: inputting semantic features and multi-level visual features into a trained text-guided image style transfer model to obtain an output image associated with the initial image; and processing the output image based on the decoder corresponding to the encoder of the 19-layer network of the visual geometry group to obtain a style-transferred target image associated with the initial image.

[0050] Figure 3 As a style transfer network framework diagram, the network structure or network framework designed in this application is as follows: Figure 3 As shown, the network consists of an encoder, a style transfer module, and a decoder.

[0051] In this embodiment, the style description text is first processed by the CLIP text encoder and the pre-trained large model ViT-B / 32 to generate semantic features. At the same time, the content map is processed by the VGG19 encoder to extract multi-level visual features. In the core architecture, the GAM attention module realizes cross-level feature fusion, and the DCNv4 deformable convolution dynamically adapts to texture deformation. The two work together to inject text semantics into visual features. By using content loss to maintain the consistency of the main structure, using text alignment loss to drive the style transfer direction, and using offset regularization loss to constrain the deformation amplitude, the three are jointly optimized under the overall loss framework. Finally, the VGG19 decoder generates a high-quality image that retains the content structure and accurately matches the text style, thus fully realizing end-to-end style transfer from text semantic understanding to pixel-level style transfer.

[0052] Optionally, in some embodiments of this application, before training the text-guided image style transfer model according to a pre-designed multi-objective collaborative composite loss function, the method further includes: obtaining a local semantic alignment loss function, a global orientation loss function, and a total variation regularization loss function; constructing a text loss function based on the local semantic alignment loss function, the global orientation loss function, and the total variation regularization loss function; obtaining a content preservation loss function and a deformable convolution offset regularization loss function; and constructing a multi-objective collaborative composite loss function based on the text loss function, the content preservation loss function, and the deformable convolution offset regularization loss function.

[0053] The text loss function can include three parts: local semantic alignment loss function, global orientation loss function, and total variation regularization loss function. The multi-objective collaborative composite loss function can include three parts: text loss function, content preservation loss function, and deformable convolution offset regularization loss function.

[0054] For example, the process of designing a composite loss function for multi-objective collaboration is as follows.

[0055] The overall loss function is:

[0056]

[0057] in, For content preservation loss, For local semantic alignment loss, For global directional loss, For total variation regularization, This is for deformable convolution offset regularization.

[0058] Specifically, and These are the core methods for implementing Text_loss (responsible for local details and global direction respectively). These are optimization tools (improving the visual quality of the stylization results, indirectly enhancing the effectiveness of Text_loss). Together, these three elements constitute the loss function system for text-guided style transfer, serving the goal of Text_loss. The following sections will detail these various loss parameters.

[0059] Among them, memory retention loss To maintain the consistency of the generated image and the content image in the deep semantic structure and avoid the stylization process from destroying the key features of the original content. Specifically, the mean square error of the content feature map is calculated in the conv4_2 and conv5_2 layers of the VGG-19 network, and its expression is shown in (2).

[0060]

[0061] in, This represents the feature map of the i-th layer of VGG, where T is the generated image and C is the content image. Control the intensity of the content.

[0062] It should be noted that you only need to select one of the middle layers (stages 3 and 4) of the VGG-19 network to extract local texture information and abstract semantic information, capturing detailed features such as the edges and textures of objects in the image; then select another layer (stage 5) of the deep VGG-19 network to extract high-level semantic information (the overall outline and structure of the object). The conv4_2 and conv5_2 layers are just one stable selection method, and can be adjusted in practice. This embodiment does not limit their specific settings.

[0063] Among them, local semantic alignment loss It utilizes fine-grained alignment between local image patches and text descriptions to achieve precise control of style details and enhance the semantic relevance of texture transfer. It mainly calculates the CLIP space cosine similarity by randomly cropping and enhancing the generated image, as shown in (3).

[0064]

[0065] in, S represents the style text, where S is the randomly cropped block. The intensity coefficient is used when the similarity is below a threshold. Optimization stops when the time is right. The function is for extracting visual features of images based on the VGG19 encoder, specifically using the feature map output by the conv4_2 layer; For: a function that extracts semantic features of text using the CLIP text encoder and the ViT-B / 32 pre-trained model; For: a set of local image patches randomly cropped from the generated image; Norm operations are performed on vectors.

[0066] Among them, global directional loss By calculating the feature difference direction between the generated image and the content image, and aligning it with the feature difference direction of the text description, the global motion direction of the generated image and the text in the CLIP feature space is kept consistent, thereby improving the overall style coordination. Its expression is shown in (4).

[0067]

[0068] in, For source text parameters, Used to control alignment strength.

[0069] Among them, the total variation regularization loss This method suppresses high-frequency noise and irregular textures in the generated image and improves visual smoothness by calculating the sum of squared differences between adjacent pixels in the horizontal, vertical, and diagonal directions. Its expression is shown in (5). The smoothing intensity control coefficient is set to 0.002.

[0070]

[0071] in, To generate the pixel value of the pixel located in the i-th row and j-th column of the image; , , And so on; Norm operations are performed on vectors.

[0072] Among them, deformable convolution offset regularization loss The offset magnitude of deformable convolutional layers in the style transfer network is constrained by applying L1 regularization to the offset parameters of the deformable convolutional kernels, preventing excessive deformation from damaging the content structure. The expression is shown in (6). For the convolution kernel neighborhood, The deformation strength control coefficient is set to 0.07.

[0073]

[0074] in, Let be the offset of the k-th sampling point in the deformable convolution kernel at position (x, y) in the feature map; k is the index of the sampling point in the convolution kernel, x is the row coordinate of the feature map, and y is the column coordinate of the feature map. This represents the absolute value of the offset of the k-th sampling point of the deformable convolution kernel at the position (x, y) in the feature map.

[0075] Optionally, in some embodiments of this application, a text-guided image style transfer model based on a U-shaped network is constructed based on global attention and fourth-generation deformable convolution, including: constructing a global attention module and a fourth-generation deformable convolution module; based on the network structure of the U-shaped network, the global attention module and the fourth-generation deformable convolution module are introduced into the upsampling stage and downsampling stage of the U-shaped network to construct the text-guided image style transfer model.

[0076] In this embodiment, by embedding the global attention module and the fourth-generation deformable convolution module into the network structure of the U-Net, and changing the processing procedures of its upsampling and downsampling stages, a text-guided image style transfer model is constructed, thereby realizing the improvement of the U-Net network structure. That is, through the collaborative design of the DCNv4 deformable convolution module and the GAM global attention module, a dual-path architecture of multi-scale deformation perception and global semantic enhancement is constructed.

[0077] Optionally, in some embodiments of this application, based on the network structure of the U-shaped network, a global attention module and a fourth-generation deformable convolutional module are introduced into the upsampling and downsampling stages of the U-shaped network to construct a text-guided image style transfer model. This includes: in the downsampling stage of the U-shaped network, adjusting the channel dimensions to convert the feature map into a preset format before inputting it into the fourth-generation deformable convolutional module; in each downsampling layer of the U-shaped network, restoring the original channel dimensions of the features processed by the fourth-generation deformable convolutional module and cascading them with the global attention module; in the upsampling stage of the U-shaped network, fusing the enhanced features of the global attention modules at each stage of the encoder through skip connections; and after each downsampling layer of the U-shaped network, introducing the global attention module again to recalibrate the spatial channel weights of the fused features to construct the text-guided image style transfer model.

[0078] In this embodiment, Figure 4 A schematic diagram of the improved U-net network structure, as shown below. Figure 4 As shown, the improvement process is as follows.

[0079] First, during downsampling, the channel dimensions are adjusted using conv_to_dcnv4 convolution (e.g., expanding the input channels from ngf=16 to 64), converting the feature map to NHWC format before inputting it into the DCNv4 module. Since DCNv4 achieves adaptive adjustment of the sampling grid based on dynamic offsets, enhances deformation representation by removing the softmax constraint of spatial aggregation, and optimizes vectorization loading and NHWC arrangement, it can improve forward speed during downsampling.

[0080] In each downsampling layer, the features processed by DCNv4 are restored to the original channel dimension through conv_back and cascaded with the GAM module. GAM adopts a channel-space dual attention mechanism, performing SE-style compression excitation in the channel dimension and capturing the spatial mask through spatial convolution in the spatial dimension, effectively suppressing background noise.

[0081] Secondly, in upsampling, the GAM enhancement features of each stage of the encoder, such as d1_f, d2_f, and d3_f, are fused by skip connections to form a multi-level feature pyramid.

[0082] Then, through bilinear upsampling and channel concatenation in UpBlock, cross-level detail-semantic information fusion is achieved.

[0083] It should also be noted that a GAM module is introduced again after each upsampling layer to recalibrate the spatial-channel weights of the fused features. For example, after upsampling to the original... Figure 1 At 4x4 resolution, GAM uses spatial attention to focus on the edge region of the target to enhance boundary sharpness.

[0084] Additionally, it should be noted that DCNv4's offset prediction may introduce "over-deformation" of local features (such as excessive stretching of brush strokes leading to texture breakage), and the upsampling process amplifies such defects; the GAM module, through global context awareness (such as using spatial masks to identify over-deformed regions), can suppress artifacts caused by offset anomalies and ensure global consistency of style transfer; therefore, the introduction of GAM in upsampling is to further enhance the synergistic effect between DCNv4 and GAM.

[0085] Optionally, in some embodiments of this application, the global attention module includes a channel attention module and a spatial attention module; the channel attention module is used to generate channel attention weights based on the input feature map, and fuse the original features of the input feature map with the channel attention weights channel by channel to obtain the channel attention output result; the spatial attention module is used to capture spatial dependencies of the input channel attention output result through convolution operation, output spatial attention weights, and fuse the spatial attention weights with the channel attention output result to obtain the global attention output result.

[0086] Among them, fusion can be multiplication, the channel attention output is the data result obtained after processing by the channel attention module, and the global attention output is the data processing result of the entire global attention module after processing by the spatial attention module.

[0087] For example, in the process of improving the DCNv4 module, a global attention mechanism is introduced. The global attention module realizes global interaction across channels and space through a two-dimensional attention module. Its core structure includes a channel attention module and a spatial attention module. The specific implementation process of the global attention module is as follows.

[0088] First, input the feature map at the input terminal. The channel attention module generates channel attention weights. Specifically, F is first reshaped into a three-dimensional tensor. And through a multilayer perceptron, i.e., MLP, a weight matrix for cross-channel interaction is generated:

[0089]

[0090] in, It contains two fully connected layers. Using the Sigmoid activation function, the output channel attention map is generated. .

[0091] Next, the original features are multiplied by the attention weights channel by channel, i.e. This is to achieve feature enhancement while preserving global information in the channel dimension, thereby avoiding information loss due to pooling.

[0092] Then, the channel attention features are input into the spatial attention module, and spatial dependencies are captured through convolution operations. Specifically, two grouped convolutions with 7×7 kernels are used to reduce computational cost, as shown in the following expression:

[0093]

[0094] This convolution method reduces the number of channels to [amount missing]. Where r is the compression ratio, the second convolution restores it to the original number of channels, and outputs a spatial attention feature map: Simultaneously, the spatial weights are multiplied by the channel-enhanced features: This operation is primarily used to enhance important spatial areas and suppress irrelevant background.

[0095] Through these sequences, GAM performs global feature interaction between the channel attention and spatial attention modules, specifically expressed as follows:

[0096]

[0097] It is worth noting that GAM utilizes MLP and unique convolutional operations to model the global relationship between channel and spatial attention respectively, solving the dimensionality separation problem of traditional attention mechanisms. By comparing the individual effects of channel attention and spatial attention, GAM replaces the pooling operation of channel attention with MLP, enhancing the cross-dimensional interaction of text-guided style transfer. At the same time, GAM effectively reduces information reduction by replacing pooling convolution with double convolution, and reduces feature loss by avoiding pooling operations, thus preserving image information to the greatest extent. This approach in style transfer technology can pay more attention to the edge structure and depth information of the input image, and enhance its attention to image details in view synthesis, thereby achieving a three-dimensional style transfer that conforms to real-world features.

[0098] Optionally, in some embodiments of this application, the fourth-generation deformable convolution module is used to: calculate the dynamic offset and adjust the sampling point position according to the dynamic offset; calculate the modulation amplitude and adjust the importance of the feature points according to the modulation amplitude; and determine the output feature value at the center position according to the dynamic offset and the modulation amplitude.

[0099] The DCNv4 module incorporates global attention, which means that offset calculations are performed only in the deformable convolutional layer (DCNv4).

[0100] Unlike traditional convolution, deformable convolution in DCNv4 introduces dynamic offsets. And calculate the modulation amplitude. The calculation formula is as follows:

[0101]

[0102] in, The values ​​are predicted by DCNv4 and used to adjust the sampling point positions so that the convolution window can adapt to different deformation patterns. As a modulation amplitude weight, it is used to adjust the importance of feature points and enhance the focus on key regions; The coordinates of the center sampling position of the convolution kernel; central position Output feature value at; The nth fixed sampling point of the convolution kernel relative to the center position Offset coordinates (such as in a 3×3 convolution kernel) Possible values ​​include ((-1, -1), (-1, 0), ..., (1, 1)), etc. Input feature map; The convolution kernel at the nth sampling point The weight value at that location.

[0103] The DCNv4 module proposed in this application integrates global attention and consists of a fourth-generation deformable convolution and global attention to form a feature extraction network. It achieves synergistic optimization of local deformation modeling and global semantic perception through a dual-path architecture. The DCNv4 module adopts a dynamic offset perception mechanism. By removing the softmax normalization of spatial aggregation and optimizing the memory access mode, its 3×3 deformable convolution kernel can adaptively adjust the sampling grid and quickly output the offset while maintaining parameter efficiency.

[0104] Specifically, the core purpose of removing softmax normalization from spatial aggregation is to improve feature representation capabilities, and the core purpose of optimizing memory access patterns is to improve computational efficiency. Specifically, this refers to adjusting the data arrangement and access order by converting the feature map from NCHW format (channel priority) to NHWC format (height-width-channel priority) and using vectorized loading (such as 32x32x4 tensor access in Tensor Core) to reduce the randomness of memory access.

[0105] Additionally, in some embodiments... Figure 5 A schematic diagram of a deformable convolutional framework that incorporates global attention, refer to Figure 5 Before being fed into DCNv4, feature maps undergo channel dimension adjustment, using conv2d for channel expansion and restoration to ensure tensor dimension compatibility. The Global Attention Module (GAM) employs a dual spatial-channel attention mechanism, generating attention weights through channel compression and spatial masks through spatial convolution, complementing and enhancing the offset predictions of DCNv4. Figure 5 In this context, X_Feature (input feature map) refers to the original feature map input to the DCNv4 module or attention module; P_conv (offset prediction convolution) refers to the convolution in the DCNv4 module used to predict the offset. A lightweight convolutional network; M_conv (modulation amplitude convolution) refers to the module in DCNv4 used to predict modulation amplitude ( The convolutional network is used to adjust the feature weights of the sampled points after the offset; Input Features are the input features; Channel Compression is the channel compression; Spatial Mask Generation is the spatial mask generation; Output Features are the output features; DCNv4offset prediction is the offset prediction of the fourth generation deformable convolution.

[0106] Additionally, it should be noted that CLIP in this application can be replaced with other multimodal models (such as ALIGN) to achieve text-image semantic alignment, only requiring adjustment of the weight matrix design for cross-modal feature fusion; other deformable convolutional versions (such as DCNv3) can also be used to replace DCNv4, only requiring re-optimization of the parameter efficiency of the offset prediction network, with the specific design scheme based on the ability to improve the transfer effect.

[0107] In another embodiment, such as Figure 6 As shown, a text-guided image style transfer method based on global attention and fourth-generation deformable convolution is provided, including the following steps:

[0108] S601, extract semantic features from style description text, and extract multi-level visual features from initial image;

[0109] S602, obtain the local semantic alignment loss function, the global orientation loss function, and the total variation regularization loss function, and construct the text loss function based on the local semantic alignment loss function, the global orientation loss function, and the total variation regularization loss function;

[0110] S603, obtain the content preservation loss function and the deformable convolutional offset regularization loss function, and construct a multi-objective collaborative composite loss function based on the text loss function, the content preservation loss function and the deformable convolutional offset regularization loss function;

[0111] S604, constructs a global attention module and a fourth-generation deformable convolution module;

[0112] S605, based on the network structure of the U-shaped network, introduces the global attention module and the fourth-generation deformable convolution module into the upsampling and downsampling stages of the U-shaped network to construct a text-guided image style transfer model.

[0113] S606, The text-guided image style transfer model is trained according to the pre-designed multi-objective collaborative composite loss function to obtain the trained text-guided image style transfer model;

[0114] S607: Input semantic features and multi-level visual features into the trained text-guided image style transfer model to obtain the output image associated with the initial image;

[0115] S608 processes the output image based on the decoder corresponding to the encoder of the 19-layer network of the visual geometry group, and obtains the style-transferred target image associated with the initial image.

[0116] It should be noted that the specific limitations of the above steps can be found in the above description of the specific limitations of a text-guided image style transfer method based on global attention and fourth-generation deformable convolution, which will not be repeated here.

[0117] The following specific embodiment illustrates the research process and other technical details of the text-guided image style transfer method based on global attention and fourth-generation deformable convolution provided in this application.

[0118] In existing technologies, text-guided image style transfer technology has gradually become a research hotspot in the fields of digital art creation and computer vision. However, the development of this technology has encountered the following obstacles:

[0119] The challenge of balancing style and content: Most existing text-guided image style transfer methods struggle to balance the extraction of global style and local texture details. Overemphasizing style features may damage the structure of the original image, while excessive preservation of content may result in insignificant style transfer effects. It is difficult to finely present details such as brushstrokes and textures while performing large-scale style transfer, leading to generated images that appear monotonous or blurry at the level of detail.

[0120] The problem of semantic alignment between style and content: Traditional methods are mostly based on local feature transformation and lack global semantic awareness, which makes the generated results prone to "style dominance imbalance" (such as excessive transfer of textures that destroys the content structure) or "insufficient text response" (such as the inability to accurately match abstract style descriptions such as "Van Gogh's Starry Night").

[0121] Insufficient cross-modal feature alignment: Existing text-guided style transfer techniques have weak semantic response capabilities when fusing text descriptions with image style features, resulting in poor correlation between stylization results and text semantics. This is especially true in complex scenes where texture misalignment or inconsistent regional styles are likely to occur.

[0122] Lack of global perception capability: Existing methods are mostly based on pixel or local feature transformation, lacking the capture of the global structure and semantics of the image, resulting in generated images that lack realism and spatial sense, and are difficult to maintain style consistency in multi-view scenes, easily resulting in deformation, perspective distortion or artifacts.

[0123] Limited adaptability of network architecture: Traditional convolutional networks have a fixed receptive field, making it difficult to dynamically adapt to complex geometric deformations and texture details under text guidance, resulting in insufficient style transfer accuracy, especially in the presentation of edge areas and detailed textures (such as brushstrokes and materials).

[0124] Adaptability limitations of stereoscopic scenes: When applied to multi-view or stereoscopic images, existing 2D style transfer techniques struggle to maintain consistency between depth information and structure, often resulting in issues such as perspective distortion and parallax inconsistency, failing to meet the demands for stereoscopic visual effects in fields such as AR / VR and 3D content creation.

[0125] Although researchers have proposed some improvements, the following problems still exist.

[0126] Text-visual feature fusion optimization: such as enhancing the alignment ability of text and image features through cross-modal attention mechanisms (such as StyleStudio's Cross-modal AdaIN), or using Transformer architecture (such as TRTST) to improve long-range dependency modeling. However, such methods still have problems such as fuzzy priority of style feature injection and insufficient synergy between local texture and global structure.

[0127] Enhanced Dynamic Geometric Modeling: Deformable Convolution (DCN) improves its ability to capture features of irregular targets by adaptively adjusting the sampling position of the convolution kernel. The fourth-generation Deformable Convolution (DCNv4) further introduces dynamic kernel weight prediction and multi-scale offset fusion, which optimizes parameter efficiency and deformation adaptability. However, it is difficult to combine with text semantics for targeted style transfer when applied alone.

[0128] Exploration of 3D style transfer: Combining view synthesis techniques such as Neural Radiation Field (NeRF), researchers have attempted to extend style transfer to 3D scenes (such as StyleRF and NeRF-Art) and maintain style consistency across multiple perspectives through implicit radiation field modeling. However, such methods have high computational complexity and are prone to edge artifacts and color breaks during dynamic viewpoint switching.

[0129] To address the aforementioned issues, this application provides a text-guided image style transfer method based on global attention and fourth-generation deformable convolution. This method can be applied to scenarios such as image style transfer software, digital art creation tools, and film and television post-processing systems. By dynamically adjusting the sampling position of the convolution kernel (DCNv4) and global semantic guidance (GAM), it achieves multi-granular fusion of style features and content structure, thereby improving transfer accuracy and multi-view style consistency.

[0130] This application proposes a text-guided style transfer framework that integrates improved DCNv4 and Global Attention (GAM), comprising an encoder, a style transfer module (containing DCNv4 and GAM), and a decoder. Text semantic features are extracted using CLIP, and multi-level visual features of the image are extracted using VGG19. In the style transfer module, GAM generates a channel-space dual-domain weight matrix to guide DCNv4 in dynamically adjusting the sampling position of the convolution kernel, prioritizing the injection of semantically relevant regions of style features. Combined with optimization using multiple loss functions (content loss, local alignment loss, global orientation loss, etc.), the final image generated retains both content structure and accurately matches the text style.

[0131] In the DCNv4 module that integrates global attention, dynamic offsets are used ( ) and modulation amplitude ( To achieve adaptive adjustment of the convolution kernel, and combined with the channel-space dual-domain weight matrix of GAM, the problem of insufficient adaptation of fixed kernel convolution to geometric deformation is solved. In terms of cross-modal multi-granularity feature fusion mechanism, multi-scale alignment of style and content is achieved by capturing micro-brush details (such as low-level features of VGG19) at the shallow level and modeling macro-construction logic (such as high-level features of VGG19) at the deep level. In terms of collaborative optimization of multiple loss functions, the style transfer accuracy and content structure preservation are balanced by joint constraints such as content loss, local alignment loss and global orientation loss.

[0132] To more objectively demonstrate the superiority of the method of this invention over other attention mechanisms, 10 images were extracted and subjected to text style transfer using different methods, and their corresponding saliency images were generated. The SSIM (Structural Similarity Index Measure), PSNR (Peak Signal-to-Noise Ratio), and LPIPS (Learned Perceptual Image Patch Similarity) values ​​between the saliency map of the text style transfer result and the saliency map of the original content map for each method were calculated, and their average values ​​were obtained. The results are shown in Table 1.

[0133] Table 1: Comparison of saliency map metrics for style transfer results using the CLIP method and different attention combinations.

[0134]

[0135] As can be seen from the information in Table 1, the mean SSIM reached 0.438, which is 17.4% higher than the 0.373 of the existing best attention method AA attention. This indicates that the image generated by the method of the present invention is closer to the original content image in terms of key regional structures such as object outlines and texture details, avoiding the edge blurring problem caused by local overfitting in traditional attention mechanisms.

[0136] In terms of PSNR, the method provided in this application achieves an average value of 14.69 dB, which is 8.7% higher than the existing best method CLIP's 13.51, demonstrating the advantages of this application in color consistency and noise suppression, especially in reducing pixel-level distortion in complex lighting scenes.

[0137] On LPIPS, the method provided in this application achieves an average score as low as 0.362, which is 23.1% lower than the best existing method CLIP's 0.471. This indicates that the generated image is highly consistent with the original content image in human visual perception, thus solving the "incongruity" problem caused by semantic misalignment in traditional methods.

[0138] To verify the superiority of the proposed method in iterative optimization of the single-graph training process, training convergence curves of the loss function were generated based on the loss values ​​generated by the proposed method and other methods during the training process. The results are as follows. Figure 7 As shown, Figure 7 A comparison chart showing iterative optimization of the loss function.

[0139] The results demonstrate that the proposed method offers significant advantages in improving model training efficiency and stability. The rapid decrease in loss value in the early stages of training indicates that the proposed method can quickly capture key features from content images, thereby accelerating the learning process. Furthermore, the stable convergence observed throughout the training process suggests that the proposed method is more robust to fluctuations in training data, minimizing potential instability factors that could hinder training progress.

[0140] To quantitatively demonstrate the advantages of the proposed solution, two sets of stylistic text statements in the style of Van Gogh's Starry Night and Fields were used to generate effect diagrams, saliency maps, depth estimation maps, and edge detection maps, respectively. The proposed method was then compared and analyzed with existing mainstream style transfer methods. This application also calculated the SSIM, PSNR, and LPIPS values ​​of the saliency maps, edge structure maps, and depth estimation maps of the aforementioned style transfer methods compared to the original content maps, and obtained the mean values, as shown in Table 2.

[0141] Table 2: Comparison of metrics for existing style transfer semantic feature maps

[0142]

[0143] As can be directly seen from Table 2, on the salient region map, the proposed solution significantly outperforms other methods such as AAMS (SSIM 0.449, LPIPS 0.556) with SSIM 0.765 and LPIPS 0.146, demonstrating its ability to accurately locate style feature injection regions. On the edge detection map, the proposed solution is only 4.5% and 6.6% lower than the optimal Stytr2 method in SSIM and PSNR, respectively, and 27.3% lower than the second-best method in LPIPS, indicating a micro-level restoration of content map strokes. On the depth estimation map, the proposed solution is close to the theoretical limit of 1.0 with SSIM 0.898 and LPIPS 0.146, which is 43.5% lower than the comparison method, verifying the advantage of the proposed solution in modeling the physical rationality of spatial perspective.

[0144] Figures 1 to 7 Any technical feature in the embodiments corresponding to any of the above items is also applicable to the embodiments of this application. Figures 8 to 10The corresponding implementation examples will not be repeated hereafter.

[0145] The above describes a text-guided image style transfer method based on global attention and fourth-generation deformable convolution in the embodiments of this application. The following describes the apparatus for performing the above-mentioned text-guided image style transfer method based on global attention and fourth-generation deformable convolution.

[0146] Reference Figure 8 This paper describes a text-guided image style transfer device based on global attention and fourth-generation deformable convolution. The device includes:

[0147] The text acquisition module 801 is used to acquire style description text. Based on the contrastive language image pre-trained text encoder and pre-trained converter model, semantic features in the style description text are extracted.

[0148] The initial image acquisition module 802 is used to acquire the initial image to be style-transformed, and extracts multi-level visual features from the initial image based on the encoder of the 19-layer network of the visual geometry group.

[0149] The model building and training module 803 is used to build a text-guided image style transfer model based on a U-shaped network based on global attention and fourth-generation deformable convolution. The text-guided image style transfer model is trained according to a pre-designed multi-objective collaborative composite loss function to obtain a trained text-guided image style transfer model.

[0150] The target image output module 804 is used to input semantic features and multi-level visual features into the trained text-guided image style transfer model to output the style-transferred target image associated with the initial image.

[0151] In this embodiment of the application, the cooperation between the above modules can improve the effect of style transfer of text-guided images.

[0152] In another embodiment, a text-guided image style transfer apparatus based on global attention and fourth-generation deformable convolution is provided. This apparatus can be a computer device, such as a server, and its internal structure diagram can be as follows: Figure 9As shown, the device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores relevant data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements the text-guided image style transfer method based on global attention and fourth-generation deformable convolution described in the above embodiment.

[0153] In yet another embodiment, a text-guided image style transfer apparatus based on global attention and fourth-generation deformable convolution is provided. This apparatus can be a computer device, such as a terminal, and its internal structure diagram can be as follows: Figure 10 As shown, the device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements the text-guided image style transfer method based on global attention and fourth-generation deformable convolution as described in the above embodiment. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the computer equipment casing, or external keyboards, touchpads, or mice, etc.

[0154] Those skilled in the art will understand that Figure 9 and Figure 10The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the text-guided image style transfer device based on global attention and fourth-generation deformable convolution applied thereto. The specific text-guided image style transfer device based on global attention and fourth-generation deformable convolution may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements, in order to achieve the functions of computer devices such as terminals or servers.

[0155] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0156] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0157] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, apparatuses, or modules, and may be electrical, mechanical, or other forms.

[0158] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0159] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0160] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0161] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)).

[0162] The technical solutions provided in the embodiments of this application have been described in detail above. Specific examples have been used in the embodiments of this application to illustrate the principles and implementation methods of the embodiments of this application. The description of the above embodiments is only for the purpose of helping to understand the methods and core ideas of the embodiments of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the embodiments of this application. Therefore, the content of this specification should not be construed as a limitation on the embodiments of this application.

Claims

1. A text-guided image style transfer method based on global attention and fourth generation deformable convolution, characterized in that, The method comprises: obtaining a style description text, extracting semantic features in the style description text based on a text encoder pre-trained by comparing language images and a pre-trained converter model; obtaining an initial image to be subjected to style conversion, and extracting multi-level visual features in the initial image based on an encoder of a visual geometry group 19 layer network; constructing a text-guided image style transfer model based on a U-shaped network based on global attention and a fourth generation deformable convolution, training the text-guided image style transfer model according to a pre-designed multi-target cooperative compound loss function, and obtaining a trained text-guided image style transfer model; inputting the semantic features and the multi-level visual features into the trained text-guided image style transfer model to output a target image after style transfer associated with the initial image.

2. The method of claim 1, wherein, Before the step of training the text-guided image style transfer model according to the pre-designed multi-target cooperative compound loss function, the method further comprises: obtaining a local semantic alignment loss function, a global direction loss function and a total variation regularization loss function, and constructing a text loss function based on the local semantic alignment loss function, the global direction loss function and the total variation regularization loss function; obtaining a content preservation loss function and a deformable convolution offset regularization loss function, and constructing the compound loss function based on the multi-target cooperation according to the text loss function, the content preservation loss function and the deformable convolution offset regularization loss function.

3. The method of claim 1, wherein, The text-guided image style transfer model based on the U-shaped network based on global attention and the fourth generation deformable convolution comprises: a global attention module is constructed, and a fourth generation deformable convolution module is constructed; the global attention module and the fourth generation deformable convolution module are introduced into an up-sampling stage and a down-sampling stage of the U-shaped network based on a network structure of the U-shaped network, and the text-guided image style transfer model is constructed.

4. The method of claim 3, wherein, The text-guided image style transfer model based on the U-shaped network based on global attention and the fourth generation deformable convolution comprises: in the down-sampling stage of the U-shaped network, adjusting the channel dimension to convert the feature map into a preset format and then inputting the fourth generation deformable convolution module; in each down-sampling layer of the U-shaped network, restoring the original channel dimension of the feature processed by the fourth generation deformable convolution module, and cascading the global attention module; in the up-sampling stage of the U-shaped network, the enhanced features of the global attention module in each stage of the encoder are fused through a skip connection; after each down-sampling layer of the U-shaped network, the global attention module is introduced again, the spatial channel weight of the fused features is recalibrated, and the text-guided image style transfer model is constructed.

5. The method of claim 3, wherein, The global attention module comprises a channel attention module and a spatial attention module; The channel attention module is configured to generate channel attention weights according to an input feature map, and fuse original features of the input feature map with the channel attention weights channel by channel to obtain a channel attention output result. The spatial attention module is configured to capture spatial dependencies through convolution operation on the input channel attention output result to output spatial attention weights, and fuse the spatial attention weights with the channel attention output result to obtain a global attention output result.

6. The method of claim 3, wherein, The fourth generation deformable convolution module is configured to: calculate a dynamic offset, and adjust a sampling point position according to the dynamic offset; calculate a modulation amplitude, and adjust importance of a feature point according to the modulation amplitude; determine an output feature value at a center position according to the dynamic offset and the modulation amplitude.

7. The method of claim 1, wherein, The input of the semantic features and the multi-level visual features into the trained text-guided image style transfer model to output the target image after style transfer associated with the initial image includes: inputting the semantic features and the multi-level visual features into the trained text-guided image style transfer model to obtain an output image associated with the initial image; processing the output image based on a decoder corresponding to an encoder of the visual geometry group 19-layer network to obtain the target image after style transfer associated with the initial image.

8. A device for text-guided image style transfer based on global attention and fourth generation deformable convolution, characterized in that, The device includes: a text acquisition module configured to acquire a style description text, extract semantic features in the style description text based on a text encoder pre-trained based on a contrastive language-image pre-training and a pre-trained converter model; an initial image acquisition module configured to acquire an initial image to be subjected to style conversion, and extract multi-level visual features in the initial image based on an encoder of a visual geometry group 19-layer network; a model construction and training module configured to construct a text-guided image style transfer model based on a U-shaped network based on global attention and a fourth generation deformable convolution, train the text-guided image style transfer model according to a pre-designed multi-objective collaborative compound loss function, and obtain a trained text-guided image style transfer model; a target image output module configured to input the semantic features and the multi-level visual features into the trained text-guided image style transfer model to output the target image after style transfer associated with the initial image.

9. A device for text-guided image style transfer based on global attention and fourth generation deformable convolution, characterized in that, The device includes: at least one processor, a memory; wherein the memory is configured to store program code, and the processor is configured to invoke the program code stored in the memory to execute the method of any one of claims 1 to 7.

10. A computer storage medium, characterized in that, It includes instructions that, when executed on a computer, cause the computer to perform the method of any one of claims 1 to 7.

Citation Information

Cited By

  • Style migration method, system and equipment based on regional decomposition and HJB equation

    CN121481830A