Interactive image segmentation algorithm based on algorithm computer
By optimizing the interactive image segmentation algorithm through lightweight neural networks and adaptive graph cut technology, the problems of complex network structure and tedious user interaction are solved, and high-precision and efficient image segmentation is achieved, which is suitable for medical imaging and industrial quality inspection.
Patent Information
- Application Number
- CN202510799813.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-30
AI Technical Summary
The existing interactive image segmentation algorithm has complex network structure design and is difficult to train. Users need to interact with the computer multiple times, the operation process is cumbersome, the image accuracy is low and the applicability is poor.
The lightweight neural network MobileNetV3-large is used for feature extraction. It combines the two-way interactive propagation mechanism and label propagation algorithm. Through user interaction and adaptive graph cutting technology, it optimizes segmentation boundaries and handles uncertain areas, reducing the number of user interactions.
It significantly reduces the number of user interactions while ensuring accuracy, is suitable for small sample scenarios, improves the accuracy and applicability of image segmentation, and is particularly suitable for medical imaging and industrial quality inspection.
Smart Images

Figure CN120726324A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer image segmentation, in particular to an interactive image segmentation algorithm based on an algorithmic computer. Background Art
[0002] Image segmentation, a core task in computer vision, aims to divide an image into regions or objects with specific semantic meaning. Accurate image segmentation is an essential prerequisite for many advanced vision applications, such as object recognition, scene understanding, medical image analysis, autonomous driving, image editing, and augmented reality. While deep learning, particularly fully automatic segmentation methods based on convolutional neural networks, have made significant progress and achieved near-human-level accuracy on specific datasets, they still face fundamental challenges in practical applications, including insufficient generalization to complex scenes, a lack of prior knowledge and user intent, and difficulty achieving pixel-level accuracy. Interactive image segmentation is a key technology that has emerged and continues to develop in this context. It cleverly combines human visual cognition, domain knowledge, and specific intent with the computational power and pattern recognition capabilities of computer algorithms. Users can provide simple and intuitive interactive cues such as points, scribble lines, and bounding boxes to guide the algorithm to focus on the target region and iteratively refine the segmentation results in real time.
[0003] After searching, the closest existing technology is an algorithmic computer-based interactive image segmentation method with application number CN202210034769.X. The method uses an operation module, a signal transmission module, and a signal reception module to import the image to be segmented into a computer. The analysis module and the recognition module then automatically analyze and process the initial size of the image to be segmented, perform segmentation and recognition on the initial image, and preview it. However, the above-mentioned existing interactive image segmentation algorithm still has the following shortcomings in practical applications:
[0004] The network structure design of the interactive image segmentation algorithm is complex and the training is more difficult. Users need to conduct multiple human-computer interactions. The operation process is relatively cumbersome and complicated, and the image accuracy is low and the applicability is poor. Summary of the Invention
[0005] The purpose of the present invention is to provide an interactive image segmentation algorithm based on an algorithmic computer to solve the problems of the existing interactive image segmentation algorithm proposed in the above background technology, such as complex network structure design, greater training difficulty, multiple human-computer interactions required by users, relatively cumbersome and complicated operation procedures, low image accuracy and poor applicability.
[0006] To this end, the present invention provides an interactive image segmentation algorithm based on an algorithmic computer, comprising the following steps:
[0007] S1. Image upload and preprocessing: The user uploads the image to be segmented to the computer. The computer initializes the uploaded image, generates an empty mask M, and then performs feature extraction on the uploaded image.
[0008] S2. Build a two-way interactive propagation mechanism: User clicks trigger interactions, perform intelligent contour inference, establish a two-way propagation mechanism, and introduce uncertainty assessment to guide users to obtain a label propagation probability map;
[0009] S3, Adaptive Image Segmentation: Use the probability map obtained by the label propagation algorithm to optimize the segmentation boundary and handle uncertain areas;
[0010] S4. Image detection after segmentation: The user adds correction points in the error area and performs feedback training until the user is satisfied.
[0011] Preferably, in step S1, the specific steps of extracting features from the uploaded image are:
[0012] (1) Select the lightweight neural network MobileNetV3-large as the feature extractor. First, resize the input image to the fixed size required by the feature extractor. The fixed size of the input is H×W, and the scaling ratio is recorded.
[0013] (2) Input the image into the network and obtain the output feature maps of the three specified layers, including the image spatial resolution and semantic information;
[0014] (3) Feature processing: Feature maps of different scales are upsampled to the same spatial size and upsampled using bilinear interpolation to obtain 1 / 4 of the original image size. The three upsampled feature maps are then concatenated in the channel dimension and fused through a 1x1 convolutional layer to reduce the number of channels and integrate multi-scale information. Finally, each feature map is normalized so that the feature vector has unit length.
[0015] (4) Associated to superpixels: Use the SLIC algorithm to divide the original image into superpixels. For each superpixel, calculate its corresponding feature vector and take the average of the feature vectors of all pixels covered by it as the feature vector of the superpixel.
[0016] (5) Storing features: The feature vector of each superpixel is stored in the graph node for subsequent label propagation and graph cutting.
[0017] Preferably, the three specified layers are the first layer feature map size of the input image The second layer feature map size is the input image The feature map size of the third layer is the input image The first designated layer has high spatial resolution and contains more edge and texture detail information, while the third designated layer contains more semantic information.
[0018] Preferably, in step S2: the specific steps of the two-way interactive propagation mechanism are:
[0019] (1) User click: The user clicks once on the target area and once on the background area, generating a seed set S = {s f , s b};
[0020] (2) Local response propagation: Based on the seed point, calculate the spatial distance R of the Gaussian diffusion response map f Similarity to color R b , fuse CNN features, train a shallow random forest classifier, and predict the initial probability P of each superpixel f =(v i );
[0021] (3) Global semantic propagation: Constructing image model, edge weight w ij Can be positioned as: In the formula, c represents the color mean of the image, f represents the feature of CNN, Represents the boundary strength, and the seed label is spread to the whole image through the label propagation algorithm to obtain the probability map P 全 .
[0022] Preferably, the random forest classifier is used to quickly convert the click conditions of the target area and background area of the sparse interactive information input by the user into pixel-level or super-pixel-level probability predictions, and use the random forest to learn the characteristic patterns of the local area of the seed point including color, texture and position, and predict the category probability of the unlabeled area.
[0023] Preferably, the specific steps of the label propagation algorithm are:
[0024] (1) Construct the weight matrix: For each pair of adjacent superpixel nodes v i and v j , calculate the similarity w between the two ij , use the Gaussian kernel function to define the similarity: Where d(v i ,v j ) represents node v i and v j The Euclidean distance between feature vectors, σ represents the parameter that controls the decay rate of similarity;
[0025] (2) Construct the probability transfer matrix: Calculate the diagonal matrix D ii =∑ j w ij, represents the sum of all outgoing edge weights of node i, and the expression of the probability transfer matrix is represents the probability of transferring from node i to node j;
[0026] (3) Initialize the label matrix: construct an n×2 matrix Y, where each row corresponds to the label distribution of a node. For the seed node s, if Y s =1, it indicates the target area. If Y s = 0 indicates the background area. For unmarked node u, initialize Y u =[0.5, 0.5] means uncertainty;
[0027] (4) Iterative label propagation: In each iteration t, the labels of the unlabeled nodes are updated. After each iteration, the labels of the seed nodes are reset to their initial values. The iteration is repeated until the label matrix convergence condition is met or the maximum number of iterations is reached. The label matrix convergence condition is ||Y t -Y t-1 ||F≤ζ, where ζ represents the minimum threshold;
[0028] (5) Output result: Output the feature vector of superpixel.
[0029] Preferably, in step S3, the specific steps of adaptive image segmentation are:
[0030] (1) Calculate the uncertainty of the node according to the probability graph. The specific calculation expression is: U(v i )=1-|2P(v i )-1|, in the uncertain region U(v i )≥θ, enhance the weight of adjacent edges and keep the original weight in the determined area;
[0031] (2) Solving the minimum segmentation: According to the energy function Where D(v i ) represents the negative logarithmic probability based on the image probability, V represents the smoothing term, the Max-Flow algorithm is used to solve the minimum segmentation, and the binary mask M1 is output.
[0032] Preferably, the required hardware needs to support GPU acceleration, ordinary CPU can run graph segmentation, and it needs to support click / brush input and real-time display of segmentation results.
[0033] The present invention proposes an interactive image segmentation algorithm based on an algorithmic computer, which has the following beneficial effects:
[0034] The interactive image segmentation algorithm combines lightweight semantic propagation, adaptive graph cutting, and a closed-loop user feedback system to significantly reduce user interaction while ensuring accuracy. It is more adaptable to small sample scenarios than pure deep learning solutions and is suitable for scenarios such as medical imaging and industrial quality inspection that require both accuracy and efficiency.
[0035] The algorithm can quickly adapt to new fields, new categories or unknown objects, reduce dependence on large amounts of training data in specific fields, minimize the number and complexity of user interactions while ensuring high segmentation accuracy, and improve user experience. It has the characteristics of minimalist interaction, ultra-high precision, super robustness and real-time response. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flow chart of the interactive image segmentation algorithm of the present invention;
[0037] Figure 2 Figure of the present invention. DETAILED DESCRIPTION
[0038] In order to make the purpose, technical solutions and beneficial technical effects of the present invention clearer, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described in this specification are merely for the purpose of explaining the present invention and are not intended to limit the present invention. The terms used in this specification are merely for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0039] Example:
[0040] See also Figure 1-2 , the present invention provides an interactive image segmentation algorithm based on an algorithmic computer, comprising the following steps:
[0041] S1. Image upload and preprocessing: The user uploads the image to be segmented to the computer. The computer initializes the uploaded image, generates an empty mask M, and then performs feature extraction on the uploaded image.
[0042] The specific steps for feature extraction of uploaded images are as follows: select the lightweight neural network MobileNetV3-large as the feature extractor, first adjust the input image to the fixed size required by the feature extractor, the fixed size of the input is H×W, and record the scaling ratio; input the image into the network to obtain the output feature maps of the three specified layers, including the image spatial resolution and semantic information; feature processing: upsample the feature maps of different scales to the same spatial size, use bilinear interpolation for upsampling, upsample to 1 / 4 the size of the original image, then splice the three upsampled feature maps in the channel dimension, and then fuse them through a 1x1 convolution layer to reduce the number of channels and integrate multi-scale information. information, and finally normalize each feature map so that the feature vector has unit length; associate to superpixels: use the SLIC algorithm to divide the original image into superpixels, for each superpixel, calculate its corresponding feature vector, and take the average of the feature vectors of all pixels covered by it as the feature vector of the superpixel; store features: store the feature vector of each superpixel in a graph node for subsequent label propagation and graph cutting; the random forest classifier is used to quickly convert the click status of the target area and background area of the sparse interaction information input by the user into pixel-level or superpixel-level probability prediction, use random forest to learn the feature pattern of the local area of the seed point including color, texture and position, and predict the category probability of the unlabeled area.
[0043] S2. Build a two-way interactive propagation mechanism: User clicks trigger interactions, perform intelligent contour inference, establish a two-way propagation mechanism, and introduce uncertainty assessment to guide users to obtain a label propagation probability map;
[0044] The specific steps of the two-way interactive propagation mechanism are as follows: User click: The user clicks once on the target area and once on the background area to generate a seed set S = {s f , s b Local response propagation: Based on the seed point, calculate the spatial distance R of the Gaussian diffusion response map f Similarity to color R b , fuse CNN features, train a shallow random forest classifier, and predict the initial probability P of each superpixel f =(v i ); Global semantic propagation: build image model, edge weight w ij Can be positioned as: In the formula, c represents the color mean of the image, f represents the feature of CNN, Represents the boundary strength, and the seed label is spread to the whole image through the label propagation algorithm to obtain the probability map P 全The random forest classifier is used to quickly convert the sparse interactive information input by the user into click conditions of the target area and background area into pixel-level or super-pixel-level probability predictions. The random forest is used to learn the feature patterns of the local area of the seed point, including color, texture and position, and predict the category probability of the unlabeled area.
[0045] S3, Adaptive image segmentation: Use the probability map obtained by the label propagation algorithm to optimize the segmentation boundary and process the uncertain area; the specific steps of adaptive image segmentation are: calculate the uncertainty of the node according to the probability map, and the specific calculation expression is: U(v i )=1-|2P(v i )-1|, in the uncertain region U(v i )≥θ, enhance the adjacent edge weights, and keep the original weights in the determined area; solve the minimum segmentation: according to the energy function Where D(v i ) represents the negative logarithmic probability based on the image probability, V represents the smoothing term, the Max-Flow algorithm is used to solve the minimum segmentation, and the binary mask M1 is output; image detection after segmentation: the user adds correction points in the error area and performs feedback training until the user is satisfied.
[0046] In this embodiment: taking a medical CT image as an example, a user uploads an image to be segmented with a size of 1024×1024 pixels and PNG format through a web interface. The system receives the image file, verifies the format and size, and then stores it in a temporary storage area;
[0047] Initialize an empty mask. The system creates a completely black matrix with the same size as the original image as the initial mask. Each pixel value of this mask matrix is 0 (background), waiting to be filled with subsequent segmentation results.
[0048] Image preprocessing, color space conversion: convert the BGR format to LAB color space to better preserve tissue contrast, size normalization: maintain the aspect ratio, scale the long side to 512 pixels (generate a 512×512 image) to enhance contrast, apply the CLAHE algorithm to the brightness channel to enhance the contrast between nodules and surrounding tissues, noise suppression: use a 3×3 median filter to remove CT scan noise;
[0049] Multi-scale feature extraction, loading a lightweight CNN model: three-level feature capture shallow features (128×128 resolution): capture nodule edge texture, mid-level features (64×64 resolution): extract vascular branching structure, deep features (32×32 resolution): identify the overall morphology of nodules, edge feature enhancement: calculate tissue boundary intensity map using Sobel operator, feature fusion: align the feature samples at all levels and splice them into a comprehensive feature map;
[0050] Superpixel segmentation, adaptive superpixel generation: Calculate the appropriate number of segments based on image complexity (420 superpixels are generated in this example), use the SLIC algorithm to divide the image, keep tissue boundaries aligned, and calculate superpixel features: spatial features: centroid coordinates of each superpixel, color features: average color value in LAB space, texture features: local binary pattern (LBP) variance, CNN features: multi-scale feature mean of corresponding regions;
[0051] Data structure encapsulation, the system builds a session object containing the following elements: original image: stores the original 1024×1024 image, processed image: stores the preprocessed 512×512 image, empty mask: the initial all-black mask matrix, feature set: multi-scale feature map (light / medium / deep / edge), obtain superpixel data: label map and feature vector of each superpixel
[0052] User interface preparation, front-end interface display: preprocessed CT image (with contrast enhancement), semi-transparent overlay showing superpixel boundaries, interactive toolbar (foreground / background marking buttons), system status prompt: "Initialization completed, please mark nodule areas", performance (actual data in this example), output data: feature map size: 512×512×32 (1.8MB after compression), superpixel data: 420 regions × 25-dimensional features.
[0053] In this embodiment, the system can quickly complete the conversion from the original CT image to the interactive ready state. The doctor can then click once on the nodule area and once on the surrounding tissue, and the system can perform real-time segmentation based on the features extracted in the initialization phase.
[0054] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. An interactive image segmentation algorithm based on an algorithmic computer, characterized by: The following steps are involved: S1. Image upload and preprocessing: The user uploads the image to be segmented to the computer. The computer initializes the uploaded image, generates an empty mask M, and then performs feature extraction on the uploaded image. S2. Build a two-way interactive propagation mechanism: User clicks trigger interactions, perform intelligent contour inference, establish a two-way propagation mechanism, and introduce uncertainty assessment to guide users to obtain a label propagation probability map; S3, Adaptive Image Segmentation: Use the probability map obtained by the label propagation algorithm to optimize the segmentation boundary and handle uncertain areas; S4. Image detection after segmentation: The user adds correction points in the error area and performs feedback training until the user is satisfied.
2. The interactive image segmentation algorithm based on an algorithmic computer according to claim 1, characterized in that: In step S1, the specific steps for feature extraction of the uploaded image are: (1) Select the lightweight neural network MobileNetV3-large as the feature extractor. First, resize the input image to the fixed size required by the feature extractor. The input fixed size is H×W, and the scaling ratio is recorded. (2) Input the image into the network and obtain the output feature maps of the three specified layers, including the image spatial resolution and semantic information; (3) Feature processing: Feature maps of different scales are upsampled to the same spatial size and upsampled using bilinear interpolation to obtain 1 / 4 of the original image size. The three upsampled feature maps are then concatenated in the channel dimension and fused through a 1x1 convolutional layer to reduce the number of channels and integrate multi-scale information. Finally, each feature map is normalized so that the feature vector has unit length. (4) Associated to superpixels: Use the SLIC algorithm to divide the original image into superpixels. For each superpixel, calculate its corresponding feature vector and take the average of the feature vectors of all pixels covered by it as the feature vector of the superpixel. (5) Storing features: The feature vector of each superpixel is stored in the graph node for subsequent label propagation and graph cutting.
3. The interactive image segmentation algorithm based on an algorithmic computer according to claim 2, characterized in that: The three specified layers are: the first layer feature map size is the input image The second layer feature map size is the input image The third layer feature map size is the input image The first designated layer has high spatial resolution and contains more edge and texture detail information, while the third designated layer contains more semantic information.
4. The interactive image segmentation algorithm based on an algorithmic computer according to claim 1, characterized in that: In step S2: the specific steps of the two-way interactive propagation mechanism are: (1) User click: The user clicks once on the target area and once on the background area, generating a seed set S = {s f , s b }; (2) Local response propagation: Based on the seed point, calculate the spatial distance R of the Gaussian diffusion response map f Similarity to color R b , fuse CNN features, train a shallow random forest classifier, and predict the initial probability P of each superpixel f =(v i ); (3) Global semantic propagation: Constructing image model, edge weight w ij Can be positioned as: In the formula, c represents the color mean of the image, f represents the feature of CNN, Represents the boundary strength, and the seed label is spread to the whole image through the label propagation algorithm to obtain the probability map P 全 .
5. The interactive image segmentation algorithm based on an algorithmic computer according to claim 4, characterized in that: The random forest classifier is used to quickly convert the click conditions of the target area and background area of the sparse interactive information input by the user into pixel-level or superpixel-level probability predictions, and use the random forest to learn the feature patterns of the local area of the seed point, including color, texture and position, to predict the category probability of the unlabeled area.
6. The interactive image segmentation algorithm based on an algorithmic computer according to claim 5, characterized in that: The specific steps of the label propagation algorithm are: (1) Construct the weight matrix: For each pair of adjacent superpixel nodes v i and v j , calculate the similarity w between the two ij , use the Gaussian kernel function to define the similarity: Where d(v i ,v j ) represents node v i and v j The Euclidean distance between feature vectors, σ represents the parameter that controls the decay rate of similarity; (2) Construct the probability transfer matrix: Calculate the diagonal matrix D ii =∑ j w ij , represents the sum of all outgoing edge weights of node i, and the expression of the probability transfer matrix is represents the probability of transferring from node i to node j; (3) Initialize the label matrix: construct an n×2 matrix Y, where each row corresponds to the label distribution of a node. For the seed node s, if Y s =1, it indicates the target area. If Y s = 0 indicates the background area. For unmarked node u, initialize Y u =[0.5, 0.5] means uncertainty; (4) Iterative label propagation: In each iteration t, the labels of the unlabeled nodes are updated. After each iteration, the labels of the seed nodes are reset to their initial values. The iteration is repeated until the label matrix convergence condition is met or the maximum number of iterations is reached. The label matrix convergence condition is ||Y t -Y t-1 ||F≤ζ, where ζ represents the minimum threshold; (5) Output result: Output the feature vector of superpixel.
7. The interactive image segmentation algorithm based on an algorithmic computer according to claim 1, characterized in that: In step S3, the specific steps of adaptive image segmentation are: (1) Calculate the uncertainty of the node according to the probability graph. The specific calculation expression is: U(v i )=1-|2P(v i )-1|, in the uncertain region U(v i )≥θ, enhance the weight of adjacent edges and keep the original weight in the determined area; (2) Solving the minimum segmentation: According to the energy function Where D(v i ) represents the negative logarithmic probability based on the image probability, V represents the smoothing term, the Max-Flow algorithm is used to solve the minimum segmentation, and the binary mask M1 is output.
8. The interactive image segmentation algorithm based on an algorithmic computer according to claim 1, characterized in that: The required hardware must support GPU acceleration, while ordinary CPUs can run graph segmentation. It must support click / brush input and display segmentation results in real time.
Citation Information
Patent Citations
Interactive image segmentation method based on algorithm computer
CN114549565A
Cited By
Visual guidance positioning method and system for radiator assembly line
CN121120783A
Interactive image segmentation method and system, computer equipment and storage medium
CN121280725A
Incomplete multi-mode brain tumor segmentation method based on double-level uncertainty guidance network
CN121811043A
An incomplete multi-modal brain tumor segmentation method based on a double-layer uncertainty guided network
CN121811043B