A method and system for image semantic segmentation based on few samples

By preprocessing and feature fusion of images, combined with recursive enhancement mechanism, the problems of inaccurate segmentation results, large computing resources consumption and long training time in image semantic segmentation are solved, and efficient and accurate image segmentation is achieved.

CN119339381BActive Publication Date: 2025-05-20SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411895972.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-20
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

The prior art has problems in the semantic segmentation of images, such as inaccurate segmentation results, large computing resources consumption and long training time, especially when dealing with high-resolution images or videos.

Method used

The semantic segmentation method based on few samples is adopted, and the segmentation model is optimized to improve segmentation efficiency and accuracy by pre-processing, feature extraction and fusion of images, combined with recursive enhancement mechanism and feature interaction.

Benefits of technology

It realizes the real-time and accuracy of image segmentation under the conditions of resource limitation and short training time, and can effectively deal with challenges such as object occlusion and complex road conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119339381B_ABST
    Figure CN119339381B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for image semantic segmentation based on a few samples, and relates to the technical field of semantic segmentation in computer vision. The method comprises the steps of: obtaining a known image sample set, and using the preprocessed image to train an image semantic segmentation model; wherein the training process comprises: performing feature extraction operations on the image, enhancing the high-frequency components and combining them with the low-frequency components, capturing the correlation between the information in the preliminary fusion features through feature interaction and multi-layer fusion, and iteratively optimizing the feature representation using a recursive enhancement mechanism to obtain the final fusion feature, restoring the image information according to the final fusion feature, and obtaining the final output image; performing image semantic segmentation on the image to be segmented using the image semantic segmentation model, and obtaining the image segmentation result. The present invention can overcome the problems of object occlusion, complex road conditions, resource limitations, and excessive training time, and ensure the real-time and accuracy of image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semantic segmentation in computer vision, and particularly to a few-shot based image semantic segmentation method and system. Background Art

[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Scene segmentation is an important technology in the field of computer vision. Its main purpose is to perform pixel-level classification on images, divide the images into different specific regional attributes, so as to achieve semantic understanding of the image content. Scene segmentation not only involves the recognition of objects in the image, but also includes in-depth understanding of the spatial relationships, shapes and semantic information of the objects. It enables machines to perceive like humans, recognizing the relationship and classification between the background and the foreground in a painting, and thus identifying the features of the image. Scene segmentation is used in many scenarios in practical applications. In autonomous driving, scene segmentation can achieve real-time recognition of various elements on the road, providing accurate environmental perception information for the autonomous driving system to help the vehicle perform path planning and decision-making. In augmented reality and virtual reality, by accurately identifying and segmenting objects in the real world, the interaction between virtual objects and the real environment is enhanced. In industrial automation, scene segmentation can effectively identify defects on items on the production line, thus ensuring the intelligence of the entire pipeline process.

[0004] Although scene segmentation technology has made significant progress in multiple fields, there are still problems to be solved in practical applications. First, interference from various factors during the image collection process causes the images to be unclear, such as factors like light and angle, resulting in inaccurate segmentation results and thus affecting the effect of practical applications. Second, scene segmentation models based on deep learning often require a large amount of computing resources. When processing high-resolution images or videos, the computational overhead and storage requirements are relatively large, which poses a challenge to real-time systems. Therefore, how to improve the segmentation efficiency, reduce the computational cost and resource consumption during the image segmentation process is an urgent problem to be solved in the prior art. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a few-shot based image semantic segmentation method and system, which can overcome problems such as object occlusion, complex road conditions, resource limitations and excessive training time, ensuring the real-time performance and accuracy of image segmentation.

[0006] To achieve the above purpose, the present invention is implemented through the following technical solutions:

[0007] The first aspect of the present invention provides a few-shot based image semantic segmentation method, including the following steps:

[0008] Obtain a set of known image samples and perform preprocessing operations on the sample set images;

[0009] Construct an image semantic segmentation model and use the preprocessed images to train the image semantic segmentation model;

[0010] Among them, the training process includes: performing feature extraction operations on the images, dividing the extracted features into high-frequency and low-frequency components, enhancing the high-frequency components and then combining them with the low-frequency components to obtain preliminary fused features, capturing the correlation between the information in the preliminary fused features through feature interaction and multi-layer fusion, and using a recursive enhancement mechanism to iteratively optimize the feature representation to obtain the final fused features, and restoring the image information according to the final fused features to obtain the final output image;

[0011] Use a loss function to evaluate the difference between the output image and the actual label image, optimize the parameters of the image semantic segmentation model, and obtain a trained image semantic segmentation model;

[0012] Use the trained image semantic segmentation model to perform image semantic segmentation on the image to be segmented to obtain an image segmentation result.

[0013] Furthermore, the preprocessing operations on the sample set images include data augmentation, image cropping, and normalization processing of the images, and after labeling the sample set images, they are divided into a training set and a test set.

[0014] Even further, before each training, randomly crop the sample set images to increase the number of images and ensure image diversity.

[0015] Furthermore, construct an image semantic segmentation model through an encoding-decoding method. The encoding part uses ResNet50 as the basic backbone network to extract image features, and fixes all the parameters of the backbone network without updating them during the training process.

[0016] Even further, the decoding part restores the information of the input image according to the final fused features through convolution and bilinear interpolation operations.

[0017] Furthermore, the specific steps for enhancing the high-frequency components and then combining them with the low-frequency components are:

[0018] Guide the high-frequency components through a mask, and perform enhancement processing through pooling, convolution, ReLU activation function, and DFA to capture the details of the information to obtain enhanced high-frequency components;

[0019] Perform normalization operations on the enhanced high-frequency components and fuse them with the low-frequency components.

[0020] Furthermore, the specific steps for capturing the correlation between the information in the preliminary fused features through feature interaction and multi-layer fusion are:

[0021] The preliminary fusion features are combined with the mask to generate target-specific support features, and the query image is processed in a two-branch manner. Among them, one branch enhances the direction information through the DFA module to obtain the direction features of the query image, and the other branch extracts the global features;

[0022] The support features interact with the direction features of the query image, and after weighting, they are fused with the global features of the query image;

[0023] The fusion result of each layer is combined with the output of the previous layer to generate the final feature representation.

[0024] The second aspect of the present invention provides a few-shot based image semantic segmentation system, including:

[0025] A data acquisition module, configured to acquire a set of known image samples and perform preprocessing operations on the sample set images;

[0026] A model training module, configured to construct an image semantic segmentation model and train the image semantic segmentation model using the preprocessed images;

[0027] Among them, the training process includes: performing feature extraction operations on the images, dividing the extracted features into high-frequency and low-frequency components, enhancing the high-frequency components and then combining them with the low-frequency components to obtain preliminary fusion features, capturing the correlation between the information in the preliminary fusion features through feature interaction and multi-layer fusion, and using a recursive enhancement mechanism to iteratively optimize the feature representation to obtain the final fusion features, and restoring the image information according to the final fusion features to obtain the final output image;

[0028] A model optimization module, configured to evaluate the difference between the output image and the actual label image using a loss function, optimize the parameters of the image semantic segmentation model to obtain a trained image semantic segmentation model;

[0029] An image semantic segmentation module, configured to perform image semantic segmentation on the image to be segmented using the trained image semantic segmentation model to obtain an image segmentation result.

[0030] The third aspect of the present invention provides a medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in the few-shot based image semantic segmentation method described in the first aspect of the present invention.

[0031] The fourth aspect of the present invention provides a device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the few-shot based image semantic segmentation method described in the first aspect of the present invention.

[0032] The above one or more technical solutions have the following beneficial effects:

[0033] The present invention discloses a few-shot based image semantic segmentation method and system. First, the input image is preprocessed, and the pixel values are normalized and mapped to the range of 0 to 1 to reduce the influence of external factors. Then, the fixed backbone network resnet is used to extract features from the picture. In the encoding stage, convolution is used for feature extraction. Then, the HLFFFA module enhances the representation of target features by separating high-frequency and low-frequency components, calculates attention weights independently for the high-frequency part, and retains global context information at the same time. The RCFM module evaluates the correlation between support and query features through cosine similarity and iteratively optimizes the feature representation through a recursive enhancement mechanism. In the decoding stage, the information of the input image is restored through convolution and bilinear interpolation operations to generate an output image with semantic information. Finally, the difference between the prediction result and the actual label image is evaluated through the DICE loss function to further optimize the segmentation model and output the final segmentation result. The semantic segmentation algorithm of the present invention improves the unmanned driving technology, solves problems such as object occlusion, complex road conditions, resource limitations, and long training time, and ensures real-time performance and accuracy.

[0034] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0036] Figure 1 is the framework diagram of the few-shot based image semantic segmentation method in the first embodiment of the present invention;

[0037] Figure 2 is the schematic diagram of the picture preprocessing method in the first embodiment of the present invention;

[0038] Figure 3 is the schematic diagram of HLFFFA in the first embodiment of the present invention;

[0039] Figure 4 is the schematic diagram of RCFM in the first embodiment of the present invention;

[0040] Figure 5 is the schematic diagram of DFA in the first embodiment of the present invention;

[0041] Figure 6 is the schematic diagram of FWB in the first embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0043] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof;

[0044] Embodiment 1:

[0045] Embodiment 1 of the present invention provides a few-shot based image semantic segmentation method, including the following steps:

[0046] S1: Obtain a set of known image samples and perform preprocessing operations on the sample set images.

[0047] S1.1: Receive a video stream and extract an image sequence according to a preset frame rate to obtain continuous image frame data from the same scene.

[0048] In this embodiment, taking the image semantic segmentation in the process of unmanned driving as an example, by collecting the video of the autonomous driving scene and extracting each frame of image, the video is composed of a series of image frames, which are segmented into single-frame images, and one image is extracted every 10 frames to obtain the input data.

[0049] S1.2: Perform preprocessing operations on the sample set images.

[0050] In a specific embodiment, as Figure 2 shown, the preprocessing operations on the sample set images include data augmentation, image cropping and normalization processing of the images, and after labeling the sample set images, they are divided into a training set and a test set. After the preprocessing is completed, the training set is divided into a support set and a query set, which are respectively input into the network for training. The images in the support set are support images, and the images in the query set are query images. The support set is used to train and adjust the model, while the query set is used to evaluate the generalization ability of the model.

[0051] Collect the video of the autonomous driving scene and extract the key frame images from it. Perform pixel-by-pixel annotation on the images to clarify the class labels of each image, and finally generate a training set and a test set. Before each training, randomly crop the sample set images to increase the number of images and ensure image diversity.

[0052] S1.2.1: Assign class labels to each pixel to generate a corresponding binary map.

[0053] When dealing with a large amount of image data, especially for pixel-by-pixel annotation, manual annotation is not only time-consuming but also error-prone, resulting in a waste of a large amount of human resources and easily introducing inconsistencies, which affects the quality of the training set. To improve the accuracy and reliability of the model, a large amount of high-quality annotated data is usually required. However, the traditional manual annotation method is not only cumbersome and error-prone, but also due to the large amount of data and long annotation cycle, it has become a bottleneck in training deep learning models. To solve this problem, optimizing data utilization through preprocessing techniques can significantly reduce the dependence on a large amount of annotated data. Through intelligent processing methods, such as data augmentation, image cropping, and normalization, not only the data quality is maintained, but also the required number of images can be greatly reduced, improving the training effect and accuracy of the model. In addition, the intelligent label generation method can effectively reduce annotation errors and improve the accuracy of training data. In summary, the innovation of preprocessing techniques not only saves the cost of manual annotation, avoids the deviation caused by incorrect annotation, but also improves the network performance and training efficiency by reducing the dependence on a large amount of annotated data, thus significantly improving the accuracy of the model and providing strong support for the popularization and development of deep learning technology in practical applications.

[0054] A device for establishing a background model by calculating the mean and variance of each pixel point in a picture, normalizing the pixel points, and extracting scene features;

[0055] S1.2.2: Before each training, randomly crop the input image and the annotated image to ensure image diversity.

[0056] If the cropped image is larger than the original image, crop it from a random position; if it is smaller than the original image, fill it to complete; at the same time, perform random flipping. Repeat the above steps to ensure that each training image has changes, effectively increasing the amount of training data.

[0057] S1.2.3: Calculate the similarity of all pixels in the image and start building a network model.

[0058] S1.2.4: Use the normalization method to standardize the image.

[0059] In a specific implementation, by subtracting the mean and dividing by the standard deviation, the image data is converted into a form that conforms to the normal distribution. This process effectively removes the difference in image brightness, reduces the contrast fluctuation between different images, ensures the consistency of the input data, and avoids the interference of brightness changes on the training process. After standardization, the neural network can converge more quickly and stably, and improves the sensitivity to image details.

[0060] In this embodiment, by calculating the mean and variance of each pixel in the image, a detailed background model is constructed, and normalization technology is utilized to effectively extract the key features of the image, improving the segmentation ability and performance of the model.

[0061] S2: Construct Figure 1 the image semantic segmentation model shown in the figure, and use the preprocessed image to train the image semantic segmentation model.

[0062] During this process, feature extraction is performed on the image through a pre-trained backbone network, the backbone parameters in the network are frozen, and a model architecture with a fixed background is established. By performing preprocessing operations on the images in the training set, the training process is optimized. The image semantic segmentation model designs the HLFFFA module to enhance the representation of target features by separating high-frequency and low-frequency components, calculates attention weights independently for the high-frequency part, and retains global context information at the same time; the RCFM module evaluates the correlation between support and query features through cosine similarity, and iteratively optimizes the feature representation through a recursive enhancement mechanism; finally, the prediction result is output by a 2D convolution. This process effectively improves the processing efficiency of the network and reduces the error rate at the same time, thus optimizing the model performance.

[0063] In a specific implementation manner, this embodiment adopts an efficient lightweight network model to perform real-time scene segmentation on the image frames extracted from the video of the autonomous driving scene.

[0064] The image semantic segmentation model is constructed in an encoding-decoding manner. The encoding part uses ResNet50 as the basic backbone network to extract image features, and all parameters of the backbone network are fixed and not updated during the training process.

[0065] Specifically, ResNet50 is used as the basic backbone network, and the query image and the support image are respectively input into the network that has been pre-trained on ImageNet to extract the features of the image. To ensure that the network can adapt to a wider range of scenarios and avoid overfitting, all parameters of the backbone network are fixed and not updated during the training process.

[0066] Among them, the training process of the image semantic segmentation model includes:

[0067] S2.1: Perform feature extraction operations on the image, and divide the extracted features into high-frequency and low-frequency components.

[0068] S2.1.1: During the feature extraction process, this embodiment designs a high-low frequency feature fusion attention module (HLFFFA), as Figure 3As shown, first, preliminary processing is performed using a 7×7 convolutional kernel with a stride of 2, resulting in a feature map of size 119×119×128, which can extract the basic spatial information of the image. Then, the feature map extracted in the middle layer contains richer spatial information and has a size of 60×60×1024. Finally, deeper semantic features are obtained through advanced feature extraction, and the feature map has a size of 60×60×2048. To reduce the computational complexity, 1×1 convolution is used to compress the number of channels to 64, reducing unnecessary computational volume.

[0069] S2.1.2: Through the masked average pooling method, the middle-level features and high-level features of the support image are fused to form support features, and the support features are finally abstracted into a prototype vector of size 1×1×256. This prototype vector represents the core information of the support image.

[0070] The generation process can be described by the following formula:

[0071] .

[0072] Where, represents the prototype vector, represents the average pooling operation, is the interpolation and padding operation to reshape the mask into the same shape as the support feature . In 5-shot, the average of 5 prototype vectors is taken as the new prototype vector.

[0073] Support feature is separated into high-frequency components and low-frequency components .

[0074] ,

[0075] .

[0076] Where, represents the Fourier transform, represents the function of separating high and low frequencies.

[0077] S2.2: After enhancing the high-frequency components and combining them with the low-frequency components, preliminary fused features are obtained.

[0078] S2.2.1: Guide the high-frequency components through a mask, and perform enhanced processing through pooling, convolution, ReLU activation function, and DFA to capture the details of the information, obtaining enhanced high-frequency components.

[0079] In a specific implementation, these enhanced high-frequency components are then normalized and fused with the low-frequency components, thereby enriching the detailed information while retaining the global context information, as shown in Figure 3.

[0080] For the high-frequency components and the support mask , the generation process of the enhanced high-frequency components is as follows:

[0081] ,

[0082] , .

[0083] Among them, is obtained by element-wise multiplication of the mask with , which can highlight the features of the target area; is generated by combining the outputs of max pooling and average pooling, and is used to capture multi-scale context information; is the final refined high-frequency component generated by the directional feature aggregation module.

[0084] S2.2.2: The enhanced high-frequency components are normalized and fused with the low-frequency components.

[0085] The high-frequency normalization process is defined as:

[0086] , ,

[0087] .

[0088] Among them, is obtained by concatenating the outputs of applying max pooling and average pooling operations to , is the high-frequency component obtained by normalizing the high-frequency components through the activation function, which is the enhanced feature. After the high-frequency components are feature-enhanced, the normalized high-frequency components are combined with the low-frequency components through element-wise multiplication.

[0089] S2.3: Capture the correlation between information in the initially fused features through feature interaction and multi-layer fusion.

[0090] S2.3.1: Combine the initially fused features with the mask to generate target-specific support features, and process the query image in a two-branch manner.

[0091] Among them, in this embodiment, a Recursive Cosine Fusion Module (RCFM) is designed. As Figure 4 shown, the RCFM module generates target-specific support features by combining the support image features with the mask, and processes the query image in a two-branch manner: one branch enhances the directional information through a Directional Feature Aggregation (DFA) module to obtain the directional features of the query image, and the other branch extracts global features. The structure of the DFA module is as Figure 5 shown. The module structure is divided into two paths. One path includes a 3×3 convolution, and the other path includes two 3×1 convolutions. The query image extracts directional information after passing through the two paths, and then passes through a 1×1 convolution to serve as the output of the DFA module. In this embodiment, by setting convolution kernels of different sizes in the two paths to extract directional features of the query image respectively, richer directional information can be extracted.

[0092] S2.3.2: The support features interact with the directional features of the query image, and after weighting, they are fused with the global features of the query image.

[0093] S2.3.3: Combine the fusion result of each layer with the output of the previous layer to generate the final feature representation.

[0094] Through feature interaction and multi-layer fusion, this module effectively improves the model's ability to capture local details and global semantics, thereby enhancing the representation of target features, as Figure 4 shown. Its formula is expressed as:

[0095] ,

[0096] ,

[0097] ,

[0098] ,

[0099] .

[0100] Among them, represents the product of the support mask and the support features, is the final output feature after fusion and element-wise multiplication, is the query feature enhanced by convolution, is the query feature extracted by a 3×3 convolution, represents element-wise multiplication, represents the concatenation operation along the channel dimension, Represents the features output by the previous layer.

[0101] S2.4: Use the recursive enhancement mechanism to iteratively optimize the feature representation to obtain the final fused features.

[0102] This embodiment proposes a Feature Weight Module (FWM). As Figure 6 shown, FWM extracts multi-scale receptive fields through dilated convolutions, avoiding information loss that may be caused by pooling layers. Dilated convolutions can expand the receptive field and preserve detailed information, thereby enhancing the model's ability to capture features at different scales and improving the speed and accuracy in image segmentation tasks. Its formula is:

[0103] .

[0104] Among them, represents the prior mask, refers to the prototype vector, is the concatenation of the intermediate layer features of the query image, represents the FWM operation, is the attention feature, emphasizing the important regions in the feature space, represents the recursively enhanced features refined through iterative processing.

[0105] S2.5: Restore the image information according to the final fused features to obtain the final output image.

[0106] In this embodiment, the decoding part restores the information of the input image according to the final fused features through convolution and bilinear interpolation operations.

[0107] S2.5.1: In the decoding stage, first reduce the number of channels through a 1×1 convolutional layer (stride = 1) to remove redundant information. Then, use another 1×1 convolutional layer (also with a stride of 1) to enhance the non-linear features and output a feature map with a size of 60×60×64.

[0108] S2.5.2: Pool the input feature map through pooling kernels of different sizes (such as 1x1, 2x2, 3x3, 6x6, etc.) to obtain features at different scales. Subsequently, upsample these pooling results to be the same size as the original feature map and concatenate them along the channel dimension. The concatenated multi-scale feature map contains rich global information, helping the network better understand objects of different sizes, and finally improving the semantic segmentation accuracy through convolution processing.

[0109] S2.5.3: Finally, upsample the image through the bilinear interpolation method so that the size of the predicted output is Figure 1 consistent with the original, and the size of the feature map is 400×400×number of classes.

[0110] S3: Evaluate the difference between the output image and the actual label image using a loss function, optimize the parameters of the image semantic segmentation model, and obtain a trained image semantic segmentation model.

[0111] In a specific implementation, in the lightweight network, the optimization process is carried out by comparing the difference between the prediction map and the original map. To achieve this goal, this embodiment uses a loss function to measure the gap between the two. The formula of the loss function is:

[0112] 。

[0113] Among them, is the weighting coefficient for each category, represents the probability distribution of the predicted output, represents the probability distribution of the true label, is a very small constant used to prevent the denominator from being zero.

[0114] In the final evaluation stage, to ensure that the performance of the network on the test set is consistent with that in the training stage, the processing flow of the test set is the same as that of the training set. The only difference is that no random cropping and data augmentation operations are performed. This is because in the test stage, it is hoped that the network can perform inference in real scenarios that have not been seen before and evaluate its generalization ability. Data augmentation is usually used in the training process to improve the robustness of the model. Therefore, the original image data needs to be used for prediction in the test stage.

[0115] In addition, to comprehensively evaluate the performance of the model, this embodiment uses multiple evaluation metrics to quantify its segmentation effect. Two important metrics are the mean intersection over union and the foreground-background intersection over union.

[0116] S4: Use the trained image semantic segmentation model to perform image semantic segmentation on the image to be segmented, and obtain the image segmentation result.

[0117] The semantic segmentation algorithm of the present invention can effectively optimize the unmanned driving technology, ensure accuracy while guaranteeing the real-time performance of the system, and can cope with challenges such as object occlusion and complex road conditions. In addition, the algorithm is specially designed to overcome the problems of tight computing resources and long training time, enabling the unmanned driving system to process image data more efficiently during actual operation. Through in-depth analysis of different scenarios and environments, the algorithm can provide stable performance under various road conditions, ensure that autonomous driving vehicles make fast and accurate decisions in complex dynamic environments, and significantly improve the application reliability and intelligent level of the unmanned driving technology.

[0118] Embodiment 2:

[0119] Embodiment 2 of the present invention provides a few-shot based image semantic segmentation system, including:

[0120] A data acquisition module, configured to acquire a set of known image samples and perform preprocessing operations on the sample set images;

[0121] A model training module, configured to construct an image semantic segmentation model and train the image semantic segmentation model using the preprocessed images;

[0122] Among them, the training process includes: performing feature extraction operations on the images, dividing the extracted features into high-frequency and low-frequency components, enhancing the high-frequency components and then combining them with the low-frequency components to obtain preliminary fusion features, capturing the correlation between information in the preliminary fusion features through feature interaction and multi-layer fusion, and iteratively optimizing the feature representation using a recursive enhancement mechanism to obtain the final fusion features, and restoring the image information according to the final fusion features to obtain the final output image;

[0123] A model optimization module, configured to evaluate the difference between the output image and the actual label image using a loss function, optimize the parameters of the image semantic segmentation model, and obtain a trained image semantic segmentation model;

[0124] An image semantic segmentation module, configured to perform image semantic segmentation on the image to be segmented using the trained image semantic segmentation model to obtain an image segmentation result.

[0125] Embodiment 3:

[0126] Embodiment 3 of the present invention provides a medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in the few-shot based image semantic segmentation method described in Embodiment 1 of the present invention.

[0127] Embodiment 4:

[0128] Embodiment 4 of the present invention provides a device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the few-shot based image semantic segmentation method described in Embodiment 1 of the present invention.

[0129] The steps involved in the above Embodiments 2, 3, and 4 correspond to those in Method Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1.

[0130] Those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0131] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A method for image semantic segmentation based on few samples, characterized in that: The following steps are involved: Obtain a known image sample set and perform preprocessing operations on the sample set images; Build an image semantic segmentation model and use the preprocessed images to train the image semantic segmentation model; The training process includes: performing feature extraction on the image, dividing the extracted features into high-frequency and low-frequency components, enhancing the high-frequency components and combining them with the low-frequency components to obtain preliminary fusion features, capturing the correlation between information in the preliminary fusion features through feature interaction and multi-layer fusion, and iteratively optimizing feature representation using a recursive enhancement mechanism to obtain the final fusion features, restoring image information based on the final fusion features, and obtaining the final output image; The specific steps of capturing the correlation between information in the preliminary fusion features through feature interaction and multi-layer fusion are: The preliminary fusion features are combined with the mask to generate target-specific support features, and the query image is processed in a dual-branch manner. One branch enhances the directional information through the DFA module to obtain the directional features of the query image, and the other branch extracts the global features. The DFA module structure is divided into two paths, one including a 3×3 convolution and the other including two 3×1 convolutions. The query image extracts the directional information after passing through the two paths, and after fusion and addition, it passes through a 1×1 convolution as the output of the DFA module. Supports interaction between features and directional features of query images, and fusion with global features of query images after cosine similarity calculation; Combine the fusion results of each layer with the output of the previous layer to generate the final feature representation; The loss function is used to evaluate the difference between the output image and the actual label image, and the parameters of the image semantic segmentation model are optimized to obtain a trained image semantic segmentation model. The trained image semantic segmentation model is used to perform image semantic segmentation on the image to be segmented to obtain the image segmentation result.

2. The image semantic segmentation method based on few samples according to claim 1, characterized in that: The sample set image preprocessing operations include data augmentation, image cropping and normalization, and the sample set images are labeled and divided into training sets and test sets.

3. The image semantic segmentation method based on few samples according to claim 2, characterized in that: Before each training, the sample set images are randomly cropped to increase the number of images and ensure image diversity.

4. The image semantic segmentation method based on few samples according to claim 1, characterized in that: The image semantic segmentation model is constructed by encoding and decoding. The encoding part uses ResNet50 as the basic backbone network to extract image features. All parameters of the backbone network are fixed and not updated during the training process.

5. The image semantic segmentation method based on few samples according to claim 4, characterized in that: The decoding part restores the information of the input image based on the final fusion features through convolution and bilinear interpolation operations.

6. The image semantic segmentation method based on few samples according to claim 1, characterized in that: The specific steps of enhancing the high-frequency components and combining them with the low-frequency components are: The high-frequency components are guided by the mask, and the details of the information are captured through pooling, convolution, ReLU activation function and DFA enhancement processing to obtain enhanced high-frequency components; The enhanced high-frequency components are normalized and fused with the low-frequency components.

7. A few-sample based image semantic segmentation system, characterized in that: include: A data acquisition module is configured to acquire a known image sample set and perform a preprocessing operation on the sample set images; A model training module is configured to build an image semantic segmentation model and train the image semantic segmentation model using the preprocessed images; The training process includes: performing feature extraction on the image, dividing the extracted features into high-frequency and low-frequency components, enhancing the high-frequency components and combining them with the low-frequency components to obtain preliminary fusion features, capturing the correlation between information in the preliminary fusion features through feature interaction and multi-layer fusion, and iteratively optimizing feature representation using a recursive enhancement mechanism to obtain the final fusion features, restoring image information based on the final fusion features, and obtaining the final output image; The specific steps of capturing the correlation between information in the preliminary fusion features through feature interaction and multi-layer fusion are: The preliminary fusion features are combined with the mask to generate target-specific support features, and the query image is processed in a dual-branch manner. One branch enhances the directional information through the DFA module to obtain the directional features of the query image, and the other branch extracts the global features. The DFA module structure is divided into two paths, one including a 3×3 convolution and the other including two 3×1 convolutions. The query image extracts the directional information after passing through the two paths, and after fusion and addition, it passes through a 1×1 convolution as the output of the DFA module. Supports interaction between features and directional features of query images, and fusion with global features of query images after cosine similarity calculation; Combine the fusion results of each layer with the output of the previous layer to generate the final feature representation; A model optimization module is configured to use a loss function to evaluate the difference between the output image and the actual label image, optimize the image semantic segmentation model parameters, and obtain a trained image semantic segmentation model; The image semantic segmentation module is configured to perform image semantic segmentation on the image to be segmented using the trained image semantic segmentation model to obtain an image segmentation result.

8. A computer-readable storage medium, characterized in that: A plurality of instructions are stored therein, and the instructions are suitable for being loaded by a processor of a terminal device and executing the image semantic segmentation method based on a small number of samples as described in any one of claims 1-6.

9. A terminal device, characterized in that: It includes a processor and a computer-readable storage medium, the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are suitable for being loaded by the processor and executing the image semantic segmentation method based on few samples described in any one of claims 1-6.

Citation Information

Patent Citations

  • Infrared image generation method and device, equipment and medium

    CN115861753A

  • Training method and device of small sample semantic segmentation model, equipment and medium

    CN118096800A