Semantic texture double-branch image restoration method and device and electronic equipment

Through the joint training and feature extraction method of semantic texture dual-branch image repair network, the problems of semantic inconsistency and insufficient global consistency in image repair are solved, and high-quality image repair effect is achieved.

CN120219237AActive Publication Date: 2025-06-27NANCHANG UNIV
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510167292.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-15
Publication Date
2025-06-27
Estimated Expiration
2045-02-15

AI Technical Summary

Technical Problem

The prior art has problems of semantic inconsistency and insufficient global consistency in image repair, especially when dealing with the objectives of multiple different semantic categories.

Method used

A semantic texture dual-branch image repair method is proposed. Through the semantic texture dual-branch image repair network, the image repair task and semantic segmentation task are jointly trained, and the coding-decoding structure of shared weights is extracted, and the optimization and expression ability of features are improved through mutual constraint processing and hierarchical enhancement processing.

Benefits of technology

By making full use of texture information and semantic information, the semantic consistency and texture consistency of repair images are improved, thereby generating high-quality repair results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219237A_ABST
    Figure CN120219237A_ABST
Patent Text Reader

Abstract

The invention provides a semantic texture double-branch image restoration method and device and electronic equipment, and belongs to the field of image restoration. The method comprises the following steps: inputting a target damaged image into a semantic texture double-branch image restoration network; respectively extracting texture features and semantic features of the target damaged image; performing mutual constraint processing on the texture features and the semantic features to obtain optimized texture features and optimized semantic features; performing hierarchical enhancement processing on the optimized texture features to obtain enhanced texture features; and according to the optimized semantic features and the enhanced texture features, respectively obtaining a target semantic segmentation image and a target restoration image. According to the method, the image restoration task and the semantic segmentation task are completed through joint training to make full use of texture information and semantic information in the target damaged image, the semantic consistency and texture consistency of the restored image are improved, and therefore a high-quality restoration result is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image restoration, and specifically relates to a semantic texture dual-branch image restoration method, device and electronic equipment. Background Art

[0002] Image restoration refers to filling the mask area in a damaged image with reasonable content inferred from known image content. The goal is to generate the same filling content as the undamaged area, ensuring that the repaired image looks natural without any obvious boundaries or artifacts. This type of technology has been widely used in many fields that require image restoration and reconstruction, such as image editing, old photo restoration, and medical image processing.

[0003] In the prior art, image restoration methods often use deep learning models to learn effective features of damaged images in order to better understand the semantic and structural information of the image, thereby inferring and generating appropriate content to fill the mask area. However, when the damaged area involves targets of multiple different semantic categories, the difficulty of restoration increases significantly. At this time, the restoration is completed solely based on image context information, which lacks sufficient semantic information constraints and is prone to problems such as inconsistent restoration semantics and blurred boundaries between different semantic targets. To address this problem, existing methods have introduced semantic segmentation images to provide semantic information for the image restoration process, constrain the restoration process, and thus effectively improve the image restoration performance.

[0004] However, in the field of image restoration, since the input image is a damaged image, errors are inevitable in the semantic segmentation image predicted based on the input image. In this case, if only the predicted semantic segmentation image is used to provide one-way guidance for the image restoration process, it will mislead the subsequent image restoration process and fail to fully utilize the potential of semantic information. It is easy to cause deficiencies in the semantic consistency and rationality of the restored image, limiting the improvement of image restoration performance.

[0005] In addition, due to the inherent characteristics of convolutional neural networks, their ability to model long-range contextual correlations is weak, resulting in a lack of global consistency in the repair results when dealing with large missing areas, and unsatisfactory visual effects. For example, when the content to be repaired in the damaged area involves objects or scene structures that span the entire image, it is difficult to generate realistic repair results relying solely on local information. In view of this, most image repair methods focus more on the global information of damaged images. However, generating a high-quality repaired image requires not only the integration of global information to ensure overall visual consistency and coherence, but also a deep understanding of the local information of the image to improve the fineness of local textures and structures such as edges and textures.

[0006] Based on the above analysis, to solve these two problems, the present application proposes a semantic-texture dual-branch image inpainting method, apparatus, and electronic device, which can provide sufficient and accurate semantic constraints for the image inpainting process and improve the image inpainting performance. Summary of the Invention

[0007] The objective of the embodiments of the present application is to provide a semantic-texture dual-branch image inpainting method, apparatus, and electronic device, aiming to solve the problems that there are large errors in the effect of the semantic information directly predicted based on the damaged image on image inpainting, and the ability of the convolutional neural network to model long-distance context correlation is weak, resulting in the lack of global consistency in the inpainting results when dealing with large-area missing regions.

[0008] To solve the above technical problems, the present application is implemented as follows: In a first aspect, the embodiments of the present application provide a semantic-texture dual-branch image inpainting method, which includes: Input the target damaged image into the semantic-texture dual-branch image inpainting network. The semantic-texture dual-branch image inpainting network is constructed based on a texture feature branch and a semantic feature branch. Both the texture feature branch and the semantic feature branch are encoder-decoder structures. In the encoding stage, the texture feature branch and the semantic feature branch share weights; Extract the texture features and semantic features of the target damaged image respectively; In the decoding stage, perform mutual constraint processing on the texture features and semantic features to obtain optimized texture features and optimized semantic features. The mutual constraint processing is used to alternately optimize the texture features and semantic features level by level; Perform hierarchical enhancement processing on the optimized texture features to obtain enhanced texture features; the hierarchical enhancement processing performs spatial enhancement and channel enhancement on the features based on the feature hierarchical order; Obtain the target semantic segmentation image and the target inpainting image respectively according to the optimized semantic features and the enhanced texture features.

[0009] Compared with the prior art, the above technical solutions provided by the present application at least include the following beneficial effects: In this application, the target damaged image is first input into a semantic-texture dual-branch image inpainting network. The semantic-texture dual-branch image inpainting network jointly trains the image inpainting task and the image semantic segmentation task, and this network is constructed based on a texture feature branch and a semantic feature branch; both the texture feature branch and the semantic feature branch are encoder-decoder structures, which are respectively used to complete the image inpainting task and the image semantic segmentation task, avoiding mutual interference and effectively extracting the semantic features and texture features of the target damaged image; in the encoding stage, the texture feature branch and the semantic feature branch share weights to ensure the consistency and stability of feature extraction; then, in the decoding stage, mutual constraint processing is performed on the texture features and semantic features to obtain optimized texture features and optimized semantic features. The mutual constraint processing is used to alternately optimize the texture features and semantic features level by level, which can make full use of the complementary relationship between semantic information and texture information, enabling the two types of information to promote and optimize each other, and reducing the uncertainty in the inpainting process; next, hierarchical enhancement processing is performed on the optimized texture features to obtain enhanced texture features; the hierarchical enhancement processing performs spatial enhancement and channel enhancement on the texture features based on the feature hierarchy order, which can enhance the texture features from two different levels of global and local, capture rich texture information to effectively improve the texture feature expression ability, and avoid the limitation of single spatial expression; finally, according to the optimized semantic features and enhanced texture features, the target semantic segmentation image and the target inpainted image are obtained respectively.

[0010] In this application, the image inpainting task and the semantic segmentation task are completed through joint training to make full use of the texture information and semantic information in the target damaged image, improve the semantic consistency and texture consistency of the inpainted image, and thus generate high-quality inpainting results.

[0011] Preferably, mutual constraint processing is performed on the texture features and semantic features to obtain optimized texture features and optimized semantic features. The calculation formula for obtaining the optimized semantic features is: Wherein, OD s is the optimized semantic feature, D s is the said semantic feature, D t is the said texture feature, α t and β t are respectively the scaling factor and offset for the texture feature to constrain the semantic feature, γ t represents the degree to which the texture feature adjusts the semantic feature, and respectively represent calculations using convolutional kernels of two different sizes, 1×1 and 3×3 α andβ process

[0012] In this solution, relevant texture information is learned from texture features to provide guidance for semantic segmentation prediction, correct semantic segmentation bias, and improve the accuracy of semantic segmentation prediction.

[0013] Preferably, mutual constraint processing is performed on texture features and semantic features to obtain optimized texture features and optimized semantic features. The calculation formula for obtaining optimized texture features is: wherein, OD t is the optimized texture feature, α s and β s are respectively the scaling factor and offset for the semantic feature to constrain the texture feature, γ s represents the degree to which the semantic feature adjusts the texture feature.

[0014] In this solution, semantic information is learned from semantic features to provide semantic constraints and guidance for image inpainting tasks, effectively improving image inpainting performance.

[0015] Preferably, hierarchical enhancement processing is performed on the optimized texture features to obtain enhanced texture features. The specific steps include: Dividing the optimized texture features into multiple feature groups; wherein, the multiple feature groups respectively correspond to different spatial levels; Based on convolution operations, calculating spatial enhancement feature groups according to the feature groups; Assembling the spatial enhancement feature groups to generate spatial enhancement features; wherein, the spatial enhancement features are consistent with the optimized texture features in shape; Connecting the spatial enhancement features according to channels and processing them to obtain spatially channel-enhanced features; wherein, the dimensions of the spatially channel-enhanced features are consistent with the dimensions of the spatial enhancement features; Enhancing the optimized texture features according to the spatially channel-enhanced features to obtain enhanced texture features.

[0016] In this solution, based on convolution operations, calculating spatial enhancement feature groups according to the feature groups can collect relevant and effective texture information from the perspective of spatial correlation at two different levels, namely globally and locally. Connecting the spatial enhancement features according to channels and processing them to obtain spatially channel-enhanced features further enhances the effective information and weakens the ineffective information from the perspective of channel correlation on the basis of spatial enhancement, thereby improving the quality of texture features.

[0017] Preferably, based on a convolution operation, the specific steps for calculating a spatially enhanced feature group according to a feature group include: Perform a convolution operation on the feature group to generate a query feature group, a key feature group, and a value feature group; Calculate a correlation matrix based on the query feature group and the key feature group; The calculation formula for the correlation matrix is: Calculate a spatially enhanced feature group based on the correlation matrix and the value feature group; The calculation formula for the spatially enhanced feature group is: Where, Att p represents the correlation matrix, Q p represents the query feature group, K p represents the key feature group, d is the dimension of the query feature group and the key feature group, PF represents the spatially enhanced feature group, V p represents the value feature group, p represents the size of the feature blocks in the current feature group, L is the index of the current feature block in the current feature group, i ∈{1, 2, ……}.

[0018] In this solution, long-distance correlation modeling is performed separately on different feature groups to enhance texture features at both the global and local levels, and relevant and effective texture information is collected from features at different levels to obtain a richer and higher-quality texture feature representation.

[0019] Preferably, the specific steps for connecting spatially enhanced features by channels and processing them to obtain spatially channel-enhanced features include: Connect the spatially enhanced features by channels to obtain a concatenated feature; Obtain the global description information of each channel, and calculate weights based on the global description information; Process the concatenated feature based on the weights to obtain a weighted concatenated feature; Reduce the dimension of the weighted concatenated feature through a convolution operation to obtain a spatially channel-enhanced feature.

[0020] In this solution, it is possible to integrate spatially enhanced features generated at two different levels, global and local, and further enhance the effective information and weaken the ineffective information from the perspective of channel correlation, thereby improving the quality of texture features.

[0021] Preferably, the specific steps of enhancing and optimizing the texture features according to the spatial channel enhancement features to obtain enhanced texture features include: Extracting the relevant parameters of the spatial channel enhancement features; the relevant parameters at least include a scaling factor and an offset; Enhancing and optimizing the texture features based on the relevant parameters to obtain enhanced texture features.

[0022] In this solution, the scaling factor and offset of the spatial channel enhancement features are used to achieve adaptive feature enhancement. While retaining the information of the original input features, that is, the optimized texture features, this enhancement process incorporates the spatial channel enhancement features that are enhanced by spatial correlation and channel correlation successively, which helps to make full use of the effective context semantic information in the spatial channel enhancement features, thereby effectively restoring the detailed texture and local structure information of the image.

[0023] Preferably, the specific steps further include: Constructing a loss function; wherein, the loss function includes an image inpainting loss and a semantic segmentation loss; Training a semantic texture dual-branch image inpainting network based on the loss function to obtain a target semantic texture dual-branch image inpainting network.

[0024] In this solution, constructing the loss function plays a core role in guiding the optimization direction of the semantic texture dual-branch image inpainting network. It clearly defines the difference metric between the predicted inpainted image and the actual inpainted image, as well as the difference metric between the predicted semantic segmentation image and the actual semantic segmentation image, so as to minimize such differences through an optimization algorithm to adjust the model parameters and improve the performance, and obtain a target semantic texture dual-branch image inpainting network.

[0025] In a second aspect, an embodiment of the present application provides a semantic texture dual-branch image inpainting device, including: An image input module, configured to input a target damaged image into a semantic texture dual-branch image inpainting network. The semantic texture dual-branch image inpainting network is constructed based on a texture feature branch and a semantic feature branch. Both the texture feature branch and the semantic feature branch are of an encoder-decoder structure. In the encoding stage, the texture feature branch and the semantic feature branch share weights; A feature extraction module, configured to extract the texture features and semantic features of the target damaged image respectively; A mutual constraint module, configured to perform mutual constraint processing on the texture features and semantic features to obtain optimized texture features and optimized semantic features. The mutual constraint processing is used to alternately optimize the texture features and semantic features step by step; A hierarchical feature enhancement module, configured to perform hierarchical enhancement processing on the optimized texture features to obtain enhanced texture features; the hierarchical enhancement processing performs spatial enhancement and channel enhancement on the features based on the feature hierarchical order; An image output module, configured to respectively obtain a target semantic segmentation image and a target inpainting image according to the optimized semantic features and enhanced texture features.

[0026] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0027] It can be understood that the beneficial effects of the technical solutions provided in the above second aspect and third aspect can refer to the relevant descriptions in the above first aspect, and will not be elaborated here.

[0028] Additional aspects and advantages of the present application will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present application. Description of the Drawings

[0029] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the description of the embodiments in conjunction with the following drawings, where: Figure 1 is a schematic flowchart of a semantic texture dual-branch image inpainting method provided by some embodiments of the present application; Figure 2 is a schematic diagram of a semantic texture dual-branch image inpainting network shown by some embodiments of the present application; Figure 3 is a schematic diagram of a mutual constraint module shown by some embodiments of the present application; Figure 4 is a schematic diagram of a hierarchical feature enhancement module shown by some embodiments of the present application; Figure 5 is a schematic diagram of a block-level feature enhancement module shown by some embodiments of the present application; Figure 6 is a schematic diagram of a dynamic feature integration module shown by some embodiments of the present application; Figure 7 is a block diagram of a semantic texture dual-branch image inpainting method device shown by some embodiments of the present application; Figure 8 is a block diagram of an electronic device shown by some embodiments of the present application.

[0030] The following specific embodiments will further illustrate the present application in conjunction with the above drawings. Specific Embodiments

[0031] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0032] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0033] Next, a semantic texture dual-branch image inpainting method, device, and electronic device provided by the embodiments of the present application will be described in detail with reference to the accompanying drawings through specific embodiments and their application scenarios.

[0034] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a semantic texture dual-branch image inpainting method according to the first embodiment of the present application. The method includes: Step S101: Input the target damaged image into a semantic texture dual-branch image inpainting network. The semantic texture dual-branch image inpainting network is constructed based on a texture feature branch and a semantic feature branch. Both the texture feature branch and the semantic feature branch are encoder-decoder structures. In the encoding stage, the texture feature branch and the semantic feature branch share weights; Specifically, please refer to Figure 2 , the specific structure of the semantic texture dual-branch image inpainting network (Semantictexture dual-branch image inpainting network, STDNet) provided by the present application is as Figure 2 shown. Taking I gt as the ground truth image, M as the mask image to represent the position of the damaged area, where 0 represents the known area and 1 represents the damaged area. Then the damaged image I m can be expressed as As can be seen from the above figure, the STDNet joint training has completed two branches for the image inpainting task and the semantic segmentation task, namely the texture feature branch and the semantic feature branch. Each branch focuses on its respective task, giving full play to its respective advantages and focusing on extracting the most relevant texture features or semantic features. Both of these branches are implemented based on an encoder-decoder structure. In the encoding stage, the texture feature branch and the semantic feature branch share weights.

[0035] Step S102: Extract the texture features and semantic features of the target damaged image respectively; It should be noted that in the feature encoding stage, the semantic features and texture features are extracted with the same weights, so that the two tasks of image inpainting and semantic segmentation are decoded based on the same encoded features, thus ensuring the consistency and stability of feature extraction.

[0036] Step S103: Perform mutual constraint processing on the texture features and semantic features to obtain optimized texture features and optimized semantic features. The mutual constraint processing is used to alternately optimize the texture features and semantic features level by level; In a possible implementation manner, in the decoding stage, the present application first realizes the mutual constraint between the two tasks through the proposed mutual constraint module (MCM) at each level, learns semantic information from the semantic features to provide semantic constraints and guidance for the image inpainting task, and effectively improves the image inpainting performance; at the same time, learns relevant texture information from the texture features to provide guidance for the semantic segmentation prediction and correct the semantic segmentation deviation, thereby improving the accuracy of semantic segmentation.

[0037] It should be noted that the characteristic information contained in the texture features and semantic features extracted by the image inpainting task and the semantic segmentation task is not focused on the same. Simply connecting these two features by channels for processing will lead to chaotic feature representation, unable to make full use of the advantages of both, and may instead introduce noise or redundant information, which will have a negative impact on the performance of both tasks.

[0038] For this reason, the specific structure of the MCM proposed in this embodiment is as Figure 3 shown. Among them, structure (a) is the structure diagram for using the semantic features as the guiding features to optimize the texture features, and structure (b) is the structure diagram for using the texture features as the guiding features to optimize the semantic features. Relu is used to activate the output feature map after each layer of convolution, and Conv is the convolution operation.

[0039] Specifically, MCM learns three parameters, the scaling factor α , the offset β and the aggregation parameter γ . In the learning calculation α and βDuring the process of these two parameters, MCM uses convolutional kernels of two different sizes, 1 and 3, to extract and fuse features at different scales. This helps to obtain useful information at different levels and scales, enhancing the diversity and richness of features, so that the learned parameters are more expressive and adaptable. In β the calculation, the current feature is additionally introduced, effectively avoiding the over-reliance of the offset on the guiding feature, enabling the model to flexibly balance between the guiding feature and the current feature, thus generating more reasonable results. The aggregation parameter γ then dynamically adjusts the degree of optimization of the learned information for the current feature according to the specific requirements of the current task.

[0040] Specifically, referring to structure (a), taking the semantic feature as the guiding feature to optimize the texture feature, the calculation formula for obtaining the optimized texture feature is: In the formula, OD t is the optimized texture feature, D s is the semantic feature, D t is the texture feature, α s and β s are respectively the scaling factor and the offset for the semantic feature to optimize the current texture feature, γ s represents the degree to which the semantic feature adjusts the texture feature, and respectively represent the calculations using convolutional kernels of two different sizes, 1×1 and 3×3, for α and β .

[0041] Specifically, referring to structure (b), taking the texture feature as the guiding feature to optimize the semantic feature, the calculation formula for obtaining the optimized semantic feature is: In the formula, OD s is the optimized semantic feature, α t and β t are respectively the scaling factor and the offset for the texture feature to optimize the semantic feature, γ t represents the degree to which the texture feature adjusts the semantic feature.

[0042] In summary, through MCM, semantic information can be learned from semantic features and used as information supplementation for the image inpainting task branch to deepen its semantic understanding of image context information and improve semantic consistency in the inpainted image. At the same time, texture information can also be learned from texture features through MCM, and this information is used to optimize the semantic features to correct errors and uncertainties in the semantic features, thereby improving the semantic segmentation accuracy.

[0043] MCM is used at three different scales in the decoding stage of STDNet, aiming to alternately optimize and strengthen texture features and semantic features level by level, enabling the two to guide and constrain each other, so as to gradually reduce the uncertainty in the decoding process, which helps to fill in more appropriate texture details while maintaining the overall semantic rationality of the image, thereby improving the image inpainting performance.

[0044] Step S104: Perform hierarchical enhancement processing on the optimized texture features to obtain enhanced texture features; the hierarchical enhancement processing performs spatial enhancement and channel enhancement on the features based on the feature hierarchical order. Specifically, first, the optimized texture features are divided into multiple feature groups; among them, the multiple feature groups respectively correspond to different spatial levels; second, based on the convolution operation, spatial enhancement feature groups are calculated according to the feature groups. It should be noted that due to its superior global context modeling ability, the self-attention mechanism can effectively call long-distance context information during the inpainting process, which helps to generate reasonable filling content. However, its ability to capture local details is relatively weak, and this limitation leads to a lack of fineness in the details of the inpainted result image. To address this problem, the present application proposes a hierarchical feature enhancement module (HFEM) based on feature block groups and the attention mechanism, and its specific structure is as Figure 4 shown.

[0045] In a possible implementation manner, HFEM takes the texture features optimized by semantic features, that is, the optimized texture features OD t as the input. First, with p 1 and p 2 of two different sizes, the input optimized texture features OD t are divided into feature groups at two different spatial levels F 1 and F 2. If , then and , where H and W represent the height and width of the input features, C is the number of channels,L 1 and L 2 respectively refer to F 1 and F the number of feature blocks in 2. p 1 and p 2 are respectively set to 1 and 3.

[0046] Then, please refer to Figure 5 , the Patch-level feature enhancement module (PFEM) based on the design performs long-distance correlation modeling on F 1 and F 2 respectively, to enhance the texture features from both the global and local levels, collect relevant and effective texture information from the features at different levels to obtain a more abundant and high-quality texture feature representation, and generate the spatial enhancement feature groups PF 1 and PF 2.

[0047] Specifically, PFEM takes the feature group F ( F 1 or F 2) as the input, and respectively generates the query feature group Q p , the key feature group K p and the value feature group V p feature groups through three 1×1 convolutions, and calculates the correlation matrix based on the query feature group and the key feature group; the calculation formula of the correlation matrix is: Based on the correlation matrix and the value feature group, calculate the spatial enhancement feature group; The calculation formula of the spatial enhancement feature group is: Among them, Att p represents the correlation matrix, d is the dimension of the query feature group and the key feature group, PF represents the spatial enhancement feature group, p represents the size of the feature block in the current feature group ( p ∈ { p 1, p 2}), L is the index of the current feature block in the current feature group, i ∈ {1, 2,...}.

[0048] It should be noted that when p 1 = 1, the spatial enhancement feature groupPF Each feature block in 1 is actually a single pixel in the feature. At this time, PFEM calculates the correlation between each pixel and all other pixels, effectively capturing the global context information. For the spatially enhanced feature group PF 2, since p 2 = 3, where each feature group contains rich local information. At this time, the attention matrix pays more attention to the correlation between local features.

[0049] Next, assemble the spatially enhanced feature groups to generate spatially enhanced features; among them, the spatially enhanced features have the same shape as the optimized texture features; Specifically, the spatially enhanced feature groups PF 1 and PF 2 are assembled through deformation transformation to obtain spatially enhanced features with the same shape as the optimized texture features OD t : RF 1 and RF 2.

[0050] After that, connect the spatially enhanced features according to the channels and process them to obtain spatially channel-enhanced features; among them, the dimension of the spatially channel-enhanced features is the same as that of the spatially enhanced features; Specifically, connect the spatially enhanced features according to the channels to obtain a concatenated feature; obtain the global description information of each channel, and calculate the weights based on the global description information; process the concatenated feature based on the weights to obtain a weighted concatenated feature; reduce the dimension of the weighted concatenated feature through a convolution operation to obtain spatially channel-enhanced features.

[0051] In a possible implementation manner, please refer to Figure 6 , this embodiment proposes a Dynamic Feature Integration Module (DFIM) to integrate the spatially enhanced features RF 1 and RF 2 generated at two different levels. PFEM starts from the perspective of spatial correlation and collects relevant and effective texture information from different levels. DFIM further enhances the effective information and weakens the invalid information from the perspective of channel correlation on this basis, thereby improving the quality of texture features.

[0052] In a possible implementation manner, DFIM takes the spatially enhanced features RF 1 and RF 2 as inputs, connects these two features according to the channels, and then processes them to obtain the global description and corresponding weights of each channel w , through wAdaptively adjust the stitching features to more accurately capture important information from global and local information. Then use a 1×1 convolution to reduce the dimension of this feature so that it has the same dimension as RF 1 and RF 2, obtaining a feature enhanced by spatial correlation and channel correlation, that is, a spatial-channel enhanced feature RF .

[0053] Enhance and optimize the texture feature according to the spatial-channel enhanced feature to obtain an enhanced texture feature. Specifically, extract the relevant parameters of the spatial-channel enhanced feature; the relevant parameters at least include a scaling factor and an offset; enhance and optimize the texture feature based on the relevant parameters to obtain an enhanced texture feature.

[0054] In a possible implementation manner, enhance the optimized texture feature RF based on the output spatial-channel enhanced feature OD t of DFIM to obtain an enhanced texture feature ROD t , and its corresponding calculation process is described as follows: where m and n respectively represent the scaling factor and the offset learned from the spatial-channel enhanced feature RF through a convolution operation. The adaptively feature enhancement is achieved through the learned scaling factor m and the offset n . While retaining the information of the optimized texture feature OD t , this enhancement process incorporates the feature RF enhanced by spatial correlation and channel correlation successively, which helps to make full use of the effective context semantic information in the spatial-channel enhanced feature RF , thereby effectively restoring the detailed texture and local structure information of the image.

[0055] Step S105: Obtain a target semantic segmentation image and a target restoration image according to the optimized semantic feature and the enhanced texture feature respectively.

[0056] During the training process of the semantic-texture dual-branch image restoration network, STDNet outputs two results: a target restoration image I out and a target semantic segmentation image S out . During the testing process, only the target restoration image I out is output, and take I out the restoration area and the damaged image inI m The known regions in it are stitched together as the final repaired image I , that is .

[0057] In the embodiment of the present application, it further includes: constructing a loss function; wherein, the loss function includes an image repair loss and a semantic segmentation loss; training a semantic texture dual-branch image repair network based on the loss function to obtain a target semantic texture dual-branch image repair network.

[0058] Specifically, the training objective of STDNet is the combined loss Loss , which consists of the image repair loss L img and the semantic segmentation loss L seg , that is . Among them, the semantic segmentation loss is composed of the cross-entropy loss, and its formula is: Among them, n represents the number of semantic target categories in the current dataset.

[0059] The image repair loss L img is composed of the adversarial loss L adv , the L1 loss L 1, the style loss L style , the perceptual loss L perc and the total variation loss L tv .

[0060] The adversarial loss L adv Using the generative-adversarial idea helps to improve the authenticity of the repaired image, and its calculation is as follows: The L1 loss L 1, as a pixel-level loss, is used to evaluate the fitting degree between the repaired image and the ground-truth image at the pixel level. Considering that the key to the repair task is to generate appropriate content to fill the masked area, the loss of the masked area and the known area is calculated separately, and the weight of the masked area loss is increased, so that the repair task focuses on the repair situation of the masked area, and the corresponding calculation is as follows: Among them represents the loss of the known region, is the loss for the corresponding masked region, and is the weight of the masked region loss.

[0061] Perceptual loss L perc is a loss function calculated based on the high-level image features extracted by a pre-trained neural network. It pays more attention to the perceptual differences between images to generate inpainting results that are more in line with the visual effects of the human eye, and is calculated as follows: where the pre-trained VGG-19 network on ImageNet is selected as the pre-trained network, represents the activation map of the corresponding i layer in this pre-trained model.

[0062] Style loss L style is a loss function used to measure the style difference between the generated image and the ground truth image to effectively alleviate the "checkerboard" artifacts caused by transposed convolution. Here, the style loss calculation is completed using the same activation map as L perc and is calculated as follows: where is the Gram matrix of the activation map . The elements in the Gram matrix reflect the correlation between features, and this correlation can be regarded as a quantitative representation of the image style.

[0063] Total variation loss L tv helps to suppress noise in the inpainted image, maintain smoothness, and improve the inpainting quality. If the size of the output inpainted image is , the corresponding calculation formula is as follows: Therefore, the final loss L img on the image inpainting branch is defined as follows: where the loss weights are set to , , , .

[0064] In the experiments of this embodiment, two datasets with semantic segmentation images, CityScapes and CelebAMask-HQ, are selected as the ground truth image datasets, and an irregular mask dataset is selected to simulate damage. For the CityScapes dataset, in this study, a dataset of 5000 high-quality pixel-level annotated images with a size of 2048×1024 is selected as the experimental dataset. Given that the original test set lacks annotations, the validation set is used as the test set, and the other images are used as the training set, resulting in 4500 training images and 500 test images. For the CelebAMask-HQ dataset, it is divided according to the original basic settings, obtaining 27176 training images and 2824 test images. For the irregular mask dataset, it includes 12000 irregular damaged mask images. According to the damage ratio, the mask dataset is divided into six sub-mask datasets with damage ratios of 0-10%, 10-20%, 20-30%, 30-40%, 40-50%, and 50-60% respectively. Each sub-mask dataset has 2000 mask images, and 1750 of them are selected as training images, and the remaining 250 are used as test images.

[0065] During the training process, the images in the ground truth dataset are selected as the ground truth images I gt , the images in the mask dataset are selected as M , and damaged images are constructed in the form of . I m . I m and M are input into STDNet to obtain the training output results I out and S out . In this process, the size of all images is adjusted to 256×256, and the batch size is set to 4. A total of 700,000 training iterations are performed. For the first 500,000 iterations, the training is carried out with a learning rate of 1×10 -4 , and then it is reduced to 5×10 -5 to complete the remaining training.

[0066] This application proposes a semantic-texture dual-branch image inpainting method STDNet, aiming to complete the image inpainting task and the semantic segmentation task through joint training to fully utilize the texture information and semantic information in the damaged image, improve the semantic consistency and texture consistency of the inpainted image, and thus generate high-quality inpainting results. Among them, the proposed MCM extracts semantic information from semantic features to provide semantic information constraints for the image inpainting task, effectively enhancing the semantic rationality in the inpainted image; at the same time, MCM extracts texture information from texture features to provide texture information for the semantic segmentation task, helping this task more accurately distinguish regions of different semantic categories and correct the errors in semantic segmentation, thereby improving the semantic segmentation accuracy. In addition, the proposed HFEM enhances the texture features at each level in the decoding stage. In this process, HFEM first collects the effective information in the texture features from two different levels, the global and the local, in terms of spatial correlation; then, in terms of channel correlation, it further enhances and fuses the effective information, thus generating an accurate correlation matrix. Enhancing the texture features with this correlation matrix can effectively improve the expressive ability of the texture features, contributing to generating more realistic and reliable inpainting results.

[0067] It should be noted that for a semantic-texture dual-branch image inpainting method provided in an embodiment of this application, the execution subject can be a semantic-texture dual-branch image inpainting device, or a control module in the semantic-texture dual-branch image inpainting device for executing and loading a semantic-texture dual-branch image inpainting method. In an embodiment of this application, taking a semantic-texture dual-branch image inpainting device executing and loading a semantic-texture dual-branch image inpainting method as an example, a semantic-texture dual-branch image inpainting method provided in an embodiment of this application is described.

[0068] Please refer to Figure 7 , Figure 7 which is a schematic diagram of a semantic-texture dual-branch image inpainting device shown in the second embodiment of this application. The semantic-texture dual-branch image inpainting device 200 includes: An image input module 201: used to input a target damaged image into a semantic-texture dual-branch image inpainting network. The semantic-texture dual-branch image inpainting network is constructed based on a texture feature branch and a semantic feature branch. Both the texture feature branch and the semantic feature branch are encoder-decoder structures. In the encoding stage, the texture feature branch and the semantic feature branch share weights; A feature extraction module 202: used to extract the texture features and semantic features of the target damaged image respectively; A mutual constraint module 203: used to perform mutual constraint processing on the texture features and semantic features to obtain optimized texture features and optimized semantic features. The mutual constraint processing is used to alternately optimize the texture features and semantic features level by level; Furthermore, it includes: The calculation formula for obtaining the optimized semantic feature is: Wherein, OD s is the optimized semantic feature, D s is the semantic feature, D t is the texture feature, α t and β t are respectively the scaling factor and the offset for the texture feature to constrain the semantic feature, γ t represents the degree to which the texture feature adjusts the semantic feature, and respectively represent the processes of calculating α and β using two different sizes of convolutional kernels, 1×1 and 3×3.

[0069] The calculation formula for obtaining the optimized texture feature is: Wherein, OD t is the optimized texture feature, α s and β s are respectively the scaling factor and the offset for the semantic feature to constrain the texture feature, γ s represents the degree to which the semantic feature adjusts the texture feature.

[0070] Hierarchical feature enhancement module 204: used to perform hierarchical enhancement processing on the optimized texture feature to obtain the enhanced texture feature; the hierarchical enhancement processing performs spatial enhancement and channel enhancement on the feature based on the feature hierarchical order; Furthermore, it includes: Dividing the optimized texture feature into multiple feature groups; wherein, the multiple feature groups respectively correspond to different spatial levels; based on the convolutional operation, calculating the spatially enhanced feature groups according to the feature groups; Furthermore, it includes: Performing a convolutional operation on the feature groups to generate query feature groups, key feature groups, and value feature groups; based on the query feature groups and the key feature groups, calculating the correlation matrix; the calculation formula for the correlation matrix is: Based on the correlation matrix and the value feature groups, calculating the spatially enhanced feature groups; the calculation formula for the spatially enhanced feature groups is: Among them, Att p represents the correlation matrix, Q p represents the query feature group, K p represents the key feature group, d is the dimension of the query feature group and the key feature group, PF represents the spatial enhancement feature group, V p represents the value feature group, p represents the size of the feature block in the current feature group, L is the index of the current feature block in the current feature group, i ∈{1, 2, ……}.

[0071] Assemble the spatial enhancement feature group to generate spatial enhancement features; among them, the spatial enhancement features are consistent with the shape of the optimized texture features; connect the spatial enhancement features according to the channels and process them to obtain spatially channel-enhanced features; among them, the dimension of the spatially channel-enhanced features is consistent with the dimension of the spatial enhancement features; Furthermore, it includes: Connect the spatial enhancement features according to the channels to obtain a spliced feature; obtain the global description information of each channel, and calculate the weights based on the global description information; process the spliced feature based on the weights to obtain a weighted spliced feature; reduce the dimension of the weighted spliced feature through a convolution operation to obtain spatially channel-enhanced features. Enhance the optimized texture features according to the spatially channel-enhanced features to obtain enhanced texture features; Furthermore, it includes: Extract the relevant parameters of the spatially channel-enhanced features; the relevant parameters at least include the scaling factor and the offset; enhance the optimized texture features based on the relevant parameters to obtain enhanced texture features.

[0072] Image output module 205: used to obtain the target semantic segmentation image and the target restoration image according to the optimized semantic features and the enhanced texture features respectively.

[0073] The semantic texture dual-branch image restoration device in the embodiments of the present application can be a device, or a component, an integrated circuit, or a chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a palmtop computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0074] The semantic texture dual-branch image restoration device in the embodiments of the present application can be a device with an operating system. The operating system can be the Android operating system, the iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0075] The semantic texture dual-branch image restoration device provided in the embodiments of the present application can implement Figures 1 to 6 each process implemented by the semantic texture dual-branch image restoration device in the method embodiments. To avoid repetition, it will not be elaborated here.

[0076] Optionally, please refer to Figure 8 , the embodiments of the present application further provide an electronic device 300, including a processor 301, a memory 302, and a computer program 303 stored on the memory 302 and executable on the processor 301. When the computer program 303 is executed by the processor 301, it implements each process of the above-mentioned semantic texture dual-branch image restoration method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0077] The embodiments of the present application further provide a readable storage medium, on which a program or an instruction is stored. When the program or the instruction is executed by a processor, it implements each process of the above-mentioned semantic texture dual-branch image restoration method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0078] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disc.

[0079] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above-mentioned semantic texture dual-branch image restoration method embodiment and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0080] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0081] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the element. In addition, it should be pointed out that the methods and devices in the embodiments of the present application are not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0082] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0083] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A semantic texture dual-branch image restoration method, characterized in that: include: Input the target damaged image into a semantic texture dual-branch image restoration network, wherein the semantic texture dual-branch image restoration network is constructed based on a texture feature branch and a semantic feature branch, wherein both the texture feature branch and the semantic feature branch are encoding-decoding structures, and in the encoding stage, the texture feature branch and the semantic feature branch share weights; Extracting texture features and semantic features of the target damaged image respectively; In the decoding stage, the texture features and the semantic features are subjected to mutual constraint processing to obtain optimized texture features and optimized semantic features, wherein the mutual constraint processing is used to alternately optimize the texture features and the semantic features step by step; Performing hierarchical enhancement processing on the optimized texture features to obtain enhanced texture features; The hierarchical enhancement process performs spatial enhancement and channel enhancement on features based on the hierarchical order of features; According to the optimized semantic features and the enhanced texture features, a target semantic segmentation image and a target restoration image are obtained respectively.

2. A semantic texture dual-branch image restoration method according to claim 1, characterized in that: The texture feature and the semantic feature are mutually constrained to obtain optimized texture features and optimized semantic features. The calculation formula for obtaining the optimized semantic features is: in, OD s To optimize semantic features, D s is the semantic feature, D t is the texture feature, α t and β t are respectively the scaling factor and the offset of the texture feature used to constrain the semantic feature, γ t represents the extent to which the texture feature is used to adjust the semantic feature, and Represents the calculation of convolution kernels of two different sizes: 1×1 and 3×3 α and β process.

3. A semantic texture dual-branch image restoration method according to claim 1, characterized in that: The texture feature and the semantic feature are mutually constrained to obtain an optimized texture feature and an optimized semantic feature. The calculation formula for obtaining the optimized texture feature is: in, OD t To optimize texture features, α s and β s are respectively a scaling factor and an offset of the semantic feature used to constrain the texture feature, γ s Represents the degree to which the semantic feature is used to adjust the texture feature.

4. The semantic texture dual-branch image restoration method according to claim 1, characterized in that: The specific steps of performing hierarchical enhancement processing on the optimized texture features to obtain enhanced texture features include: Dividing the optimized texture features into a plurality of feature groups; wherein the plurality of feature groups correspond to different spatial levels respectively; Based on the convolution operation, a spatial enhancement feature group is calculated according to the feature group; Assembling the spatial enhancement feature group to generate a spatial enhancement feature; wherein the spatial enhancement feature is consistent with the shape of the optimized texture feature; Connecting the spatial enhancement features according to channels and processing them to obtain spatial channel enhancement features; wherein the dimension of the spatial channel enhancement features is consistent with the dimension of the spatial enhancement features; The optimized texture feature is enhanced according to the spatial channel enhancement feature to obtain an enhanced texture feature.

5. A semantic texture dual-branch image restoration method according to claim 4, characterized in that: The specific steps of calculating the spatial enhancement feature group based on the feature group based on the convolution operation include: Performing a convolution operation on the feature group to generate a query feature group, a key feature group, and a value feature group; Based on the query feature group and the key feature group, a correlation matrix is ​​calculated; The calculation formula of the correlation matrix is: Based on the correlation matrix and the value feature group, a spatial enhancement feature group is calculated; The calculation formula of the spatial enhancement feature group is: in, Att p represents the correlation matrix, Q p represents the query feature group, K p represents the key feature group, d are the dimensions of the query feature group and the key feature group, PF represents the spatial enhancement feature group, V p represents the value feature group, p Indicates the size of the feature block in the current feature group, L is the index of the current feature block in the current feature group, i ∈{1,2,…}.

6. A semantic texture dual-branch image restoration method according to claim 4, characterized in that: The specific steps of connecting the spatial enhancement features according to channels and processing them to obtain spatial channel enhancement features include: Connecting the spatial enhancement features according to channels to obtain splicing features; Obtaining global description information of each channel, and calculating a weight based on the global description information; Processing the splicing features based on the weights to obtain weighted splicing features; The dimension of the weighted concatenated feature is reduced through a convolution operation to obtain a spatial channel enhancement feature.

7. A semantic texture dual-branch image restoration method according to claim 4, characterized in that: The specific step of enhancing the optimized texture feature according to the spatial channel enhancement feature to obtain the enhanced texture feature comprises: Extracting relevant parameters of the spatial channel enhancement feature; the relevant parameters at least include a scaling factor and an offset; The optimized texture feature is enhanced based on the relevant parameters to obtain an enhanced texture feature.

8. The semantic texture dual-branch image restoration method according to claim 1, characterized in that: The specific steps also include: Constructing a loss function; wherein the loss function includes image restoration loss and semantic segmentation loss; The semantic texture dual-branch image restoration network is trained based on the loss function to obtain a target semantic texture dual-branch image restoration network.

9. A semantic texture dual-branch image restoration device, characterized in that: include: An image input module is used to input a target damaged image into a semantic texture dual-branch image restoration network, wherein the semantic texture dual-branch image restoration network is constructed based on a texture feature branch and a semantic feature branch, wherein both the texture feature branch and the semantic feature branch are encoding-decoding structures, and in the encoding stage, the texture feature branch and the semantic feature branch share weights; A feature extraction module, used to extract texture features and semantic features of the target damaged image respectively; A mutual constraint module, used for performing mutual constraint processing on the texture feature and the semantic feature to obtain an optimized texture feature and an optimized semantic feature, wherein the mutual constraint processing is used for alternately optimizing the texture feature and the semantic feature step by step; A hierarchical feature enhancement module, used for performing hierarchical enhancement processing on the optimized texture features to obtain enhanced texture features; The hierarchical enhancement process performs spatial enhancement and channel enhancement on features based on the hierarchical order of features; The image output module is used to obtain a target semantic segmentation image and a target restoration image according to the optimized semantic features and the enhanced texture features.

10. An electronic device, characterized in that: include: A memory, a processor, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, is a step of the semantic texture dual-branch image restoration method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Target segmentation system and training method thereof, and target segmentation method and device

    CN113486956A

  • Damaged mural image restoration method and system

    CN116993617A

  • Image edge and local consistency repairing method based on prior semantics

    CN117314779A

  • Remote sensing image semantic segmentation method based on double-branch dynamic attention

    CN117422878A

  • Image restoration method and system based on multi-scale semantic driving

    CN117557474A