Strawberry occlusion image reconstruction method based on contour texture optimization
By introducing contour enhancement blocks, scale adaptation blocks and texture refinement branch blocks into the image reconstruction network, the problem of poor strawberry occluded image restoration in the existing technology is solved, high-quality strawberry occluded image reconstruction is achieved, and structural coherence and texture detail recovery are improved.
Patent Information
- Application Number
- CN202510822731.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-19
AI Technical Summary
Existing GAN-based image restoration methods have difficulty capturing position continuity information when dealing with complex occlusions, and have limited multi-scale feature extraction capabilities, resulting in incoherent strawberry contours and insufficient restoration of texture details. In particular, the restoration results are often blurred or distorted in large-area occlusion scenes.
An image reconstruction network based on an encoder-decoder structure is adopted, and contour enhancement blocks, scale adaptation blocks and texture refinement branch blocks are introduced. Continuous contours are captured through state space modeling and selective scanning mechanism, multi-scale dilated convolution and dynamic grouping strategy are used to extract multi-scale information, and the texture refinement branch block is combined to achieve fine reconstruction of local texture details.
The reconstruction quality of large occluded areas is significantly improved, structural coherence and texture authenticity are maintained, the restoration effect of strawberry occluded images is improved, and the structural consistency and detail preservation capabilities are enhanced.
Smart Images

Figure CN120672624A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a strawberry occluded image reconstruction method based on contour texture optimization. Background Art
[0002] With the rapid development of computer vision technology, image restoration technology has been widely used in fields such as agricultural automation, medical image processing, and cultural relic protection. The goal of image restoration is to reconstruct the content of occluded or missing areas by analyzing the information of unoccluded areas in the image, thereby restoring the integrity and authenticity of the image. Traditional image restoration methods mainly rely on texture synthesis and mathematical diffusion techniques, such as the block matching-based texture synthesis algorithm (BTSA) and the partial differential equation image restoration algorithm (PDE-IIA). These methods are effective when dealing with regular textures or small-scale occlusions, but often cause structural distortion or texture blurring when faced with complex scenes or large-scale occlusions.
[0003] In recent years, the rise of deep learning technology has provided new solutions for image restoration. Generative adversarial networks (GANs) have demonstrated outstanding performance in image restoration due to their adversarial training mechanism and ability to generate realistic images. However, existing GAN-based methods still have shortcomings when dealing with complex occlusions: First, they struggle to effectively capture positional continuity, resulting in discontinuous strawberry outlines; second, their ability to extract multi-scale features is limited, making it difficult to adapt to varying occlusion sizes; and third, they struggle to restore texture details, especially in scenes with large occlusions caused by irregularly shaped objects (such as strawberries), resulting in blurry or distorted restoration results.
[0004] To address the above issues, some studies have attempted to improve the restoration effect by improving the network structure. For example, the Transformer-based self-attention mechanism attempts to link distant pixels, and the deformable convolution-based method adapts to local occlusion. However, these methods only bypass the occlusion problem by enhancing detection capabilities and cannot completely solve the fundamental challenge of information loss. In addition, the shape prior-based restoration method reconstructs regular shapes through geometric assumptions, but has poor adaptability to contours and edge textures under large-area occlusion. Therefore, there is an urgent need for an image restoration method that can simultaneously achieve contextual reasoning, multi-scale feature extraction, and texture detail restoration to improve reconstruction quality and support downstream tasks (such as object detection). Summary of the Invention
[0005] In view of the above defects of the prior art, the present invention provides a strawberry occlusion image reconstruction method based on contour texture optimization to solve the technical problems of poor and inaccurate image reconstruction in the prior art.
[0006] To achieve the above-mentioned purpose and other related purposes, the present invention provides a strawberry occlusion image reconstruction method based on contour texture optimization, comprising: obtaining an original image to be processed; processing the original image using a trained image reconstruction network to obtain a reconstructed unoccluded image; wherein the image reconstruction network is based on an encoder-decoder structure and further comprises: a contour enhancement block, which is introduced at the intermediate jump connection between the encoder and the decoder to capture continuous contours by establishing a distance dependency relationship between features through state space modeling and a selective scanning mechanism; a scale adaptation block, which is introduced between the encoder and the decoder to dynamically extract multi-scale information; and a texture refinement branch block, which is used to process the features output by the scale adaptation block and fuse them into the decoder to enhance texture details.
[0007] In one embodiment of the present invention, the encoder is composed of three convolutional layers with a stride of 2, and the decoder is composed of two transposed convolutional layers and one convolutional layer. The transposed convolutional layer of the decoder is used to restore the spatial resolution, and the convolutional layer is used to refine the features.
[0008] In one embodiment of the present invention, the contour enhancement block performs the following steps on the input feature X: OEB-in Processing to obtain output feature X OEB-out :Use the convolution layer to transform the input feature X OEB-in Processing is performed to obtain the first branch feature; the input feature X OEB-in After layer normalization and reshaping processing, the first branch feature and the second branch feature are fused through element-by-element feature addition operation to obtain the output feature X. OEB-out .
[0009] In one embodiment of the present invention, the selective feature processing module processes the input feature X as follows: SFP-in Processing to obtain output feature X SFP-out :Use the fully connected layer to transform the input feature X SFP-in Expand from length C to 2d, and divide the expanded features into two parts in the channel dimension to obtain feature F x and feature F z According to the feature F x , dynamic parameters are generated by linear transformation, and the dynamic parameters include input mapping parameters B t and the output mapping parameter C t ; Use the selective scanning unit to perform state space modeling according to the following formula:
[0010] x'(t)=Ax(t-1)+B tu(t),y(t)=C t x'(t)+Du(t),
[0011] Where, u(t) is the first feature F x Obtained by projection and segmentation of the fully connected layer, A and D are trainable global parameters, A is initialized by taking the logarithm of the integer sequence, the initial value of D is a constant 1, y(t) is the output feature of the selective scanning unit; the feature F is generated using the SiLU activation function z The gate signal of the selective scanning unit is combined with the output feature y(t) of the selective scanning unit and the feature F z The gate signal is multiplied element by element to obtain the output feature X SFP-out .
[0012] In one embodiment of the present invention, the scale adaptation block performs the following steps on the input feature X SAB-in Processing to obtain output feature X SAB-out :According to the input feature X SAB-in , the channel allocation weights are generated by the learnable matrix W; according to the channel allocation weights, all channel features in each group are weighted and fused to obtain the features of each group X g , g∈{1,2,…,G}, G is the total number of groups; for each group of features X g Apply dilated convolution, activation, and normalization operations with different dilation rate sets, and perform average fusion of the results within the group to obtain the fused feature F of each group. g ; Splice the fused features of each group F g , and then pass through the standard convolution layer, normalization and activation operations to obtain the fusion feature F fused ; Use the channel attention module to the fusion feature F fused Modulate and obtain the modulated characteristic F att ; Using the gating operation to att With the input feature X SAB-in Splicing in the channel dimension to obtain the output feature X SAB-out .
[0013] In one embodiment of the present invention, the output feature X SAB-out The expression of is as follows: SAB-out =M⊙X SAB-in +(1-M)⊙F aat , where M is the spatially adaptive mask generated by a 3×3 convolutional layer and a Sigmoid activation function, and ⊙ is an element-wise multiplication operation.
[0014] In one embodiment of the present invention, the texture refinement branch block performs the following steps on the input feature X: TRB-inProcessing: Use the transposed convolution layer and ReLU activation function to transform the input feature X TRB-in Processing is performed to obtain feature X TRB-1 ; Use the transposed convolution layer and ReLU activation function to transform the feature X TRB-1 Processing is performed to obtain feature X TRB-2 ; Use standard convolution layer and ReLU activation function to transform the feature X TRB-2 Processing is performed to obtain feature X TRB-3 ; The feature X TRB-1 , the feature X TRB-2 And the feature X TRB-3 are respectively fused to the corresponding layers of the decoder.
[0015] In one embodiment of the present invention, an image reconstruction network is trained according to the following steps: obtaining a data set, wherein the samples in the data set are image pairs consisting of an original image and its corresponding unobstructed image; performing enhancement processing on the data set to obtain an enhanced data set; constructing an image reconstruction network; and using the enhanced data set to train the image reconstruction network to obtain the trained image reconstruction network.
[0016] In one embodiment of the present invention, the loss function L during the image reconstruction network training is total The expression is as follows: total =λ1L rec +λ2L percep +λ3L style +λ4L adv , where L rec is the pixel reconstruction loss, L percep is the perceptual loss, L style is the style loss, L adv To combat the loss, λ1, λ2, λ3, and λ4 are the weight coefficients of each loss.
[0017] The beneficial effects of the present invention are as follows: the present invention proposes a strawberry occluded image reconstruction method based on contour texture optimization, wherein the contour enhancement block in the method effectively captures contour continuity through spatial correlation modeling and a selective scanning mechanism; the introduced scale adaptation block adopts multi-scale hole convolution and a dynamic grouping weight strategy to enhance the network's ability to extract features of different scales; the texture refinement branch block realizes fine reconstruction of local texture details through the synergistic effect of three-level texture branches and the decoder main branch; the synergistic effect of these modules enables the improved network to better process large-area occluded areas while maintaining structural coherence and texture authenticity, and ultimately achieves high-quality fruit occlusion removal; compared with the existing technology, the present invention has achieved significant improvements in removal quality, detail preservation, and structural consistency. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 A flowchart provided for an embodiment of the present invention;
[0020] Figure 2 This is a diagram of the architecture of an image reconstruction network provided by one embodiment of the present invention;
[0021] Figure 3 This is a diagram illustrating the architecture of a contour enhancement block provided by an embodiment of the present invention;
[0022] Figure 4 An architectural diagram of a scale adaptation block provided in one embodiment of the present invention;
[0023] Figure 5 This is a diagram illustrating the architecture of a texture refinement branch block according to an embodiment of the present invention;
[0024] Figure 6 A comparison chart of various image reconstruction methods provided in one embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following describes the embodiments of the present invention through specific embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. It should be noted that the following embodiments and the features in the embodiments can be combined with each other unless they conflict. In addition to the specific methods, equipment, and materials used in the embodiments, based on the understanding of the prior art by those skilled in the art and the description of the present invention, any methods, equipment, and materials of the prior art that are similar or equivalent to the methods, equipment, and materials in the embodiments of the present invention can also be used to implement the present invention.
[0026] It should be understood that the terms used in the examples of the present invention are for describing specific embodiments rather than for limiting the scope of protection of the present invention. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as those generally understood by those skilled in the art.
[0027] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In some of the embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.
[0028] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functions and operations that may be implemented by the methods and computer program products of various embodiments disclosed in the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0029] See Figure 1 and Figure 2 , Figure 1 A strawberry occlusion image reconstruction method based on contour texture optimization provided by an embodiment of the present invention includes: obtaining an original image to be processed; processing the original image using a trained image reconstruction network to obtain a reconstructed unoccluded image. Figure 2 As shown, it is based on an encoder-decoder structure and also includes an outline enhancement block (OEB), a scale adaptation block (SAB), and a texture refinement branch block (TRB).
[0030] The contour enhancement block is introduced at the intermediate jump connection between the encoder and decoder to balance computational efficiency and semantic information. It captures continuous contours by establishing distance dependencies between features through state-space modeling and a selective scanning mechanism. The scale adaptation block is introduced between the encoder and decoder, employing multi-scale dilated convolutions and a dynamic grouping strategy to dynamically extract multi-scale information. The texture refinement branch block processes the features output by the scale adaptation block and fuses them into the decoder to enhance texture details and achieve detail enhancement.
[0031] The contour enhancement block in this method effectively captures contour continuity through spatial correlation modeling and selective scanning mechanism; the introduced scale adaptation block adopts multi-scale void convolution and dynamic grouping weight strategy to enhance the network's ability to extract features of different scales; the texture refinement branch block realizes fine reconstruction of local texture details through the synergy of three-level texture branches and the decoder main branch; the synergy of these modules enables the improved network to better handle large-area occlusion areas while maintaining structural coherence and texture authenticity, and ultimately achieves high-quality fruit occlusion removal; compared with the existing technology, the present invention has achieved significant improvements in removal quality, detail preservation and structural consistency.
[0032] In a specific embodiment of the present invention, the encoder consists of three convolutional layers with a stride of 2, which gradually downsample the input features to extract multi-scale features. The decoder consists of two transposed convolutional layers and one convolutional layer. The transposed convolutional layer of the decoder is used to restore spatial resolution, and the convolutional layer is used to refine features. The encoder and decoder are the basic structure of the image reconstruction network. On this basis, the contour enhancement block, scale adaptation block, and texture refinement branch block are introduced to achieve high-quality image reconstruction. Reconstruction here can also be understood as inpainting, that is, repairing and reconstructing the occluded parts of the image.
[0033] See Figure 3 In a specific embodiment of the present invention, the contour enhancement block performs the following steps 1.1 to 1.3 on the input feature X OEB-in Processing to obtain output feature X OEB-out .
[0034] Step 1.1, use the convolution layer to input feature X OEB-in Processing is performed to obtain the first branch feature. This step corresponds to Figure 3 The convolutional layer directly enhances the local structure of the original input features. The convolution operation has the characteristics of fixed receptive field and strong locality, which can effectively preserve and enhance the details and contour features of the strawberry.
[0035] Step 1.2: Input feature X OEB-in After layer normalization and reshaping, the feature is input into the selective feature processing module to obtain the second branch feature. Figure 3 A branch on the right.
[0036] Step 1.3: By performing element-by-element feature summing operations, the first branch features and the second branch features are fused to obtain the output feature X OEB-outThe output features of the two branches are fused through an element-by-element feature summation operation to form the final output features of the contour enhancement block. Through this structural design, the contour enhancement block has the ability to simultaneously extract local structural features and model long-range spatial dependencies, effectively improving the performance of the image inpainting network under complex textures and structures.
[0037] In a specific embodiment of the present invention, the selective feature processing module processes the input feature X according to the following steps 1.2.1 to 1.2.5: SFP-in Processing to obtain output feature X SFP-out .
[0038] Step 1.2.1, use the fully connected layer to transform the input feature X SFP-in Expand from length C to 2d, and divide the expanded features into two parts in the channel dimension to obtain feature F x and feature F z .
[0039] In this step, input feature X SFP-in After the input projection (Input Proj), a fully connected layer, the feature vector at each position is expanded from length C to a longer vector with a length of 2d. This operation is implemented by a weight matrix (shape C×2d) and a bias vector (length 2d), resulting in a tensor of dimension [B, L, 2d] (i.e., the expanded feature). The output tensor is then split into two parts along the channel dimension, with the first d dimensions constituting the feature F. x , used to generate the core parameters required for state space modeling (including time step, input mapping matrix and output mapping matrix); the last d dimensions constitute the feature F z , used for subsequent gate activation control to decide which features should be retained or suppressed.
[0040] Step 1.2.2, according to feature F x , dynamic parameters are generated through linear transformation, and the dynamic parameters include input mapping parameters B t and the output mapping parameter C t .
[0041] In this step, feature F x It is sent to the parameter generation module (Param Gen), which contains a linear transformation layer. The weight matrix size of this layer is d×(r+2s), the length of the bias vector b is r+2s, and the length of the last bit of the output tensor is r+2s. The first r elements constitute the step control parameter Δt, which is used to control the propagation speed of the state in space; the middle s elements constitute the input mapping parameter B t, used to compress the input into the state space; the following s elements constitute the output mapping parameter C t , which is used to remap the state vector back to the output feature space.
[0042] Step 1.2.3: Use the selective scanning unit to perform state space modeling according to the following formula:
[0043] x'(t)=Ax(t-1)+B t u(t),
[0044] y(t)=C t x'(t)+Du(t),
[0045] Where u(t) is the first characteristic F x The state transfer matrix A and the direct transmission vector D obtained by projection and segmentation of the fully connected layer are trainable global parameters, A∈R d×s , whose value is initialized by taking the logarithm of the integer sequence and set as a trainable variable through nn.Parameter in PyTorch. It is used to control how the state information is propagated from the previous position to the current position during training. d , whose initial value is a constant 1, serves as a bypass from the input feature directly to the output to preserve local structural details. y(t) is the output feature of the selective scanning unit.
[0046] In the first formula, x(t-1) represents the state at the previous position (one vector per channel); A·x(t-1) represents the propagation of the state to the current position, and the matrix A controls the propagation ratio and attenuation mode; B t u(t) represents how the input features of the current position are compressed and added to the current state; the sum of the two, x'(t), represents the intermediate state after the current position is updated.
[0047] In the second formula, the output map C t x'(t) generates the output feature of the current position by weighted fusion of the state vector x'(t); this output is added to the original input feature under the weight of the global vector D to obtain the final spatial feature that takes into account both context modeling and local preservation.
[0048] Step 1.2.4: Generate feature F using SiLU activation function z The gate signal can be expressed as: Gate = SiLU (F z ). The gating mechanism uses another feature flow F z Based on the above, the SiLU activation function is used for modulation to determine the retention ratio of each feature, so that the network can adaptively adjust the expression strength of the feature.
[0049] Step 1.2.5: Combine the output feature y(t) of the selective scanning unit with the feature F z The gate signal is multiplied element by element to obtain the output feature X SFP-out , that is, X SFP-out =y(t)⊙F z .
[0050] See Figure 4 In a specific embodiment of the present invention, the scale adaptation block performs the following steps 2.1 to 2.6 on the input feature X SAB-in Processing to obtain output feature X SAB-out .
[0051] Step 2.1: Based on the input feature X SAB-in , generate channel assignment weights through the learnable matrix W. Input feature X SAB-in ∈R B×C×H×W , represents the set of input feature maps in a batch, where B is the batch size, C is the number of channels, and H and W are the spatial dimensions (i.e., length and width). In order to achieve dynamic grouping of input channels, the scale adaptation block introduces a learnable parameter matrix W∈R C×G , which can be automatically updated through back propagation during the training process, where G represents the number of groups of convolution branches, that is, the total number of groups. For example, in a specific embodiment of the present invention, G = 4, corresponding to 4 convolution branches. Channel allocation weight A = Softmax(W), where A∈R C×G , A c,g Indicates the weight of the c-th channel assigned to the g-th convolution branch.
[0052] Step 2.2: According to the channel assignment weights, weighted fusion of all channel features in each group is performed to obtain the feature X of each group. g , g∈{1,2,…,G}, G is the total number of groups. For each convolution branch g∈{1,2,…,G}, the input feature X SAB-in ∈R B×C×H×W Divide into C sub-tensors along the channel dimension: X=[X1,X2,…,X C ], X c ∈R B×1×H×W , each X c Represents the cth channel of the original feature tensor. All channels are grouped according to the weight A c,g Weighted fusion to obtain each group of features X g , which can be expressed as: Where the dot represents the element-wise multiplication between a scalar and a tensor.
[0053] Step 2.3: For each group of features X gApply dilated convolution, activation, and normalization operations with different dilation rate sets, and perform average fusion of the results within the group to obtain the fused feature F of each group. g .
[0054] In this step, each group g corresponds to an expansion rate set R g = {r1, r2}, for the input feature X of group g g Perform dilated convolution, activation, and normalization operations to extract local structural information with a receptive field of r. The expression is as follows:
[0055]
[0056] Represents a two-dimensional convolution with a dilation rate r, ReLU is a nonlinear activation function, and IN stands for Instance Normalization to ensure a stable distribution.
[0057] Then, the convolution output results under different expansion rates in the same group are averaged and fused to obtain the final response characteristics of group g:
[0058]
[0059] Step 2.4: Concatenate the fused features F of each group g , and then pass through the standard convolution layer, normalization and activation operations to obtain the fusion feature F fused , which can be expressed as:
[0060] F cat =Concat(F1,F2,…,F G ), F fused =ReLU(IN(Conv 3×3 (F cat ))),
[0061] In the fusion stage, the output features F of all groups g That is {F1,F2,…,F G} will be concatenated in the channel dimension to form a tensor F cat ∈R B×G×H×W , the tensor contains the response features from convolution groups of different scales, and is then sent to a feature fusion module for integration. The structure includes a 3×3 standard convolution layer (without expansion), instance normalization operation (IN) and ReLU activation function. The final fusion feature F fused ∈R B×G×H×W .
[0062] Step 2.5: Use the channel attention module to fusion feature F fused Modulate and obtain the modulated characteristic Fatt .
[0063] In this step, in order to further strengthen the structural relevance, a channel attention module is used to fusion feature F fused The channel attention operation consists of three steps: (1) Global average pooling (GAP) performs spatial averaging on each channel to obtain the channel description vector v = GAP (F fused )∈R B×G×1×1 ; (2) Then two 1×1 convolutional layers are used to generate the weight of each channel to control the importance of each channel (with ReLU and Sigmoid in the middle): α = σ(W2(ReLU(W1(v))))∈R B×G×1×1 , where W1 and W2 represent the learnable weights of the two layers of 1×1 convolution in the attention module, and σ is the Sigmoid function; (3) Attention weighted modulation, the above attention weight α is applied to the fusion feature map F fused , enhance the response of key semantic channels and suppress redundant features: Fatt = α·F fused ∈R B×G×H×W The purpose of channel attention is to allow the model to automatically emphasize important semantic group responses, suppress interference features, and improve discrimination ability.
[0064] Step 2.6, use the gating operation to transform the feature F att With input feature X SAB-in Splicing in the channel dimension to get the output feature X SAB-out .
[0065] In a specific embodiment of the present invention, the output feature X SAB-out The expression is as follows:
[0066] X SAB-out =M⊙X SAB-in +(1-M)⊙F aat ,
[0067] Where M is a spatially adaptive mask generated by a 3×3 convolutional layer and a sigmoid activation function, and ⊙ is an element-by-element multiplication operation. This mask controls the fusion ratio of the two features at each spatial location, completing the final feature fusion in a point-by-point weighted manner. This not only preserves the original structure but also introduces the ability to extract multi-scale information, effectively improving feature extraction in occluded areas.
[0068] See Figure 5 In a specific embodiment of the present invention, the texture refinement branch block performs the following steps 3.1 to 3.4 on the input feature X TRB-in to be processed.
[0069] Step 3.1, use the transposed convolution layer and ReLU activation function to transform the input feature XTRB-in Processing is performed to obtain feature X TRB-1 Specifically, the transposed convolution layer can have a kernel size of 4×4, a stride of 2, and a padding of 1. In the expanded feature map, the convolution kernel covers the input area in a sliding manner. The gray block represents the receptive field range, and the green grid corresponds to the upsampled output. The size H×W is upsampled to 2H×2W. At the same time, the ReLU activation function is used to enhance the feature expression capability.
[0070] Step 3.2: Use the transposed convolution layer and ReLU activation function to transform the feature X TRB-1 Processing is performed to obtain feature X TRB-2 This step again uses the same configuration of transposed convolution and activation operations to further increase the spatial resolution to 4H×4W. At each upsampling stage, the module fuses the intermediate features layer by layer into the backbone generation path, suppressing artifacts that may be introduced during the upsampling process while preserving rich local texture information.
[0071] Step 3.3: Use the standard convolution layer and ReLU activation function to transform the feature X TRB-2 Processing is performed to obtain feature X TRB-3 After completing two upsampling steps, the texture refinement branch uses a standard convolutional layer with a kernel size of 3×3 to perform channel compression and feature fusion to generate the final three-channel RGB output, and continues to access ReLU activation to maintain nonlinear modeling capabilities.
[0072] Step 3.4: transform feature X TRB-1 , Feature X TRB-2 and feature X TRB-3 They are fused to the corresponding layers of the decoder. TRB-1 Input to the first layer of the decoder, X TRB-2 Input to the second layer of the decoder, X TRB-3 Input to the third layer of the decoder.
[0073] All convolution kernel parameters within the texture refinement branch block are learnable variables and are automatically optimized through backpropagation during the end-to-end training process, thus giving the texture refinement branch block good texture adaptability and cross-sample generalization capabilities.
[0074] See Figure 1 ,In a specific embodiment of the present invention, the image reconstruction network is trained ,according to the following steps.
[0075] (1) Obtain a dataset. The samples in the dataset are image pairs consisting of the original image and its corresponding unobstructed image. When constructing the dataset, a detailed data collection and preprocessing plan should be developed. The acquired occluded image data should be strictly screened to remove samples that are blurred, noisy, or have other quality issues to ensure the quality and representativeness of the training dataset.
[0076] (2) Enhance the dataset to obtain an enhanced dataset. Enhancement processing includes random occlusion generation, random translation, rotation, scaling, and other geometric transformations. These data enhancement operations increase the diversity of training samples and improve the model's adaptability to different types of occlusion and deformation.
[0077] (3) Construct an image reconstruction network. The network structure is as follows Figure 2 As shown, each sub-block has been described in detail above.
[0078] (4) Use the enhanced data set to train the image reconstruction network to obtain a trained image reconstruction network.
[0079] In a specific embodiment of the present invention, the loss function L during image reconstruction network training is total The expression is as follows:
[0080] L total =λ1L rec +λ2L percep +λ3L style +λ4L adv ,
[0081] Where, L rec is the pixel reconstruction loss, L percep is the perceptual loss, L style is the style loss, L adv To combat the loss, λ1, λ2, λ3, and λ4 are the weight coefficients of each loss.
[0082] The pixel reconstruction loss calculates the pixel differences between the generated image and the real image, using the L1 norm. The perceptual loss uses a pretrained classification network (such as VGG19) to extract multi-layer features and calculate differences in the feature space to enhance semantic expression. The style loss compares the Gram matrix of the image feature map to constrain the consistency of the generated image's texture and style with the original image. The adversarial loss uses a discriminator to determine whether an image is real or fake, guiding the generator to improve the naturalness and realism of the image.
[0083] The above-mentioned weight coefficients are adjustable weight coefficients, and their values can be set according to the task scenario, data set characteristics and experimental requirements. In the present invention, they are specified by the user through command line parameters during the training startup phase, and there is no mandatory requirement to meet the normalization conditions (such as the sum is 1).
[0084] It should be noted that the step division of the various methods above is only for the purpose of clear description. During implementation, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they contain the same logical relationship, they are all within the scope of protection of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs without changing the core design of the algorithm and process are all within the scope of protection of this patent.
[0085] It is understood that when training the image reconstruction network, the peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) can be used to evaluate the model after training. Texture boundary similarity (TBSIM) can also be introduced, using local binary patterns (LBP) and gray-level co-occurrence matrix (GLCM) to measure texture similarity, while multi-scale gradient analysis is introduced to measure the naturalness of edge transitions in the restored area.
[0086] Figure 6 The figure shows the reconstruction effect comparison between the method in the present invention and some existing methods. There are four scenes from a to d, where A represents the original image, B represents the occluded binary mask, C to F represent some advanced network models in the prior art, and G represents the image reconstruction network in the present invention. Figure 6 It is not difficult to see that the strawberry reconstruction results on the four samples of the present invention not only take into account the successful restoration of large-area occlusions, but also optimize the spatial position in (a) and the texture details and structural reasoning in (bd), and minimize edge artifacts and blurring.
[0087] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical principles disclosed herein are intended to be covered by the claims of the present invention.
Claims
1. A strawberry occlusion image reconstruction method based on contour texture optimization, characterized in that: include: Get the original image to be processed; Processing the original image using a trained image reconstruction network to obtain a reconstructed unobstructed image; The image reconstruction network is based on an encoder-decoder structure and further comprises: A contour enhancement block is introduced at the intermediate skip connection between the encoder and the decoder to capture continuous contours by establishing a distance dependency relationship between features through state space modeling and a selective scanning mechanism; A scale adaptation block is introduced between the encoder and the decoder to dynamically extract multi-scale information; The texture refinement branch block is used to process the features output by the scale adaptation block and fuse them into the decoder to enhance texture details.
2. The strawberry occlusion image reconstruction method based on contour texture optimization according to claim 1, wherein The encoder consists of three convolutional layers with a stride of 2, and the decoder consists of two transposed convolutional layers and one convolutional layer. The transposed convolutional layer of the decoder is used to restore the spatial resolution, and the convolutional layer is used to refine the features.
3. The strawberry occlusion image reconstruction method based on contour texture optimization according to claim 1, wherein The contour enhancement block performs the following steps on the input feature X OEB-in Processing to obtain output feature X OEB-out : Use the convolution layer to transform the input feature X OEB-in Processing is performed to obtain the first branch feature; For the input feature X OEB-in After layer normalization and reshaping, the layers are input into the selective feature processing module to obtain the second branch features. By performing element-by-element feature summing operation, the first branch feature and the second branch feature are fused to obtain the output feature X OEB-out .
4. The strawberry occlusion image reconstruction method based on contour texture optimization according to claim 3, wherein The selective feature processing module performs the following steps on the input feature X SFP-in Processing to obtain output feature X SFP-out : The input feature X is transformed into SFP-in Expand from length C to 2d, and divide the expanded features into two parts in the channel dimension to obtain feature F x and feature F z ; According to the feature F x , dynamic parameters are generated by linear transformation, and the dynamic parameters include input mapping parameters B t and the output mapping parameter C t ; The state space modeling is performed using the selective scanning unit according to the following formula: x'(t)=Ax(t-1)+B t u(t), y(t)=C t x'(t)+Du(t), Where, u(t) is the first feature F x Obtained by projection and segmentation of the fully connected layer, A and D are trainable global parameters, A is initialized by taking the logarithm of the integer sequence, the initial value of D is a constant 1, and y(t) is the output feature of the selective scanning unit; The feature F is generated using the SiLU activation function z The gating signal; The output feature y(t) of the selective scanning unit is compared with the feature F z The gate signal is multiplied element by element to obtain the output feature X SFP-out .
5. The strawberry occlusion image reconstruction method based on contour texture optimization according to claim 1, wherein The scale adaptation block performs the following steps on the input feature X SAB-in Processing to obtain output feature X SAB-out : According to the input feature X SAB-in , generate channel assignment weights through the learnable matrix W; According to the channel allocation weights, all channel features in each group are weighted and fused to obtain the features of each group X g , g∈{1,2,…,G}, G is the total number of groups; For each set of features X g Apply dilated convolution, activation, and normalization operations with different dilation rate sets, and perform average fusion of the results within the group to obtain the fused feature F of each group. g ; Splice the fused features F of each group g , and then pass through the standard convolution layer, normalization and activation operations to obtain the fusion feature F fused ; The channel attention module is used to fused Modulate and obtain the modulated characteristic F att ; The feature F is transformed into att With the input feature X SAB-in Splicing in the channel dimension to obtain the output feature X SAB-out .
6. The strawberry occlusion image reconstruction method based on contour texture optimization according to claim 5, wherein The output feature X SAB-out The expression is as follows: X SAB-out =M⊙X SAB-in +(1-M)⊙F aat , Where M is the spatially adaptive mask generated by a 3×3 convolutional layer and a Sigmoid activation function, and ⊙ is an element-wise multiplication operation.
7. The strawberry occlusion image reconstruction method based on contour texture optimization according to claim 1, wherein The texture refinement branch block performs the following steps on the input feature X TRB-in Processing: The input feature X is transformed using the transposed convolution layer and the ReLU activation function. TRB-in Processing is performed to obtain feature X TRB-1 ; The feature X is transformed using the transposed convolution layer and the ReLU activation function. TRB-1 Processing is performed to obtain feature X TRB-2 ; The feature X is transformed using a standard convolutional layer and ReLU activation function. TRB-2 Processing is performed to obtain feature X TRB-3 ; The feature X TRB-1 , the feature X TRB-2 And the feature X TRB-3 are respectively fused to the corresponding layers of the decoder.
8. The strawberry occlusion image reconstruction method based on contour texture optimization according to claim 1, wherein The image reconstruction network is trained as follows: Obtain a data set, where samples in the data set are image pairs consisting of an original image and its corresponding unobstructed image; Performing enhancement processing on the data set to obtain an enhanced data set; Build an image reconstruction network; The image reconstruction network is trained using the enhanced data set to obtain the trained image reconstruction network.
9. The strawberry occlusion image reconstruction method based on contour texture optimization according to claim 8, wherein The loss function L during the image reconstruction network training total The expression is as follows: L total =λ1L rec +λ2L percep +λ3L style +λ4L adv , Where, L rec is the pixel reconstruction loss, L percep is the perceptual loss, L style is the style loss, L adv To combat the loss, λ1, λ2, λ3, and λ4 are the weight coefficients of each loss.
Citation Information
Cited By
Image restoration method and electronic equipment
CN121213426A