A method for improving the stability of video image segmentation based on loss function

By simulating the motion characteristics of the video sequence and introducing a stability loss adjustment model, the problem of insufficient stability in video image segmentation is solved, and a more accurate and consistent video segmentation result is achieved.

CN112949529BActive Publication Date: 2025-05-09HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110271743.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-12
Publication Date
2025-05-09
Estimated Expiration
2041-03-12

AI Technical Summary

Technical Problem

The existing video image segmentation method is difficult to achieve stable segmentation results in video analysis, and there are often problems of jitter or inaccuracy of edge pixels, and the existing stability improvement method takes a long time or has great limitations.

Method used

By simulating the motion characteristics of the video sequence, the model is trained using a self-supervised method, and the stability loss adjustment model is introduced. The network is fine-tuned by introducing perturbations to the images, reducing holes and leaks segmentation, and enhancing the consistency of segmentation boundaries.

Benefits of technology

It effectively alleviates the jitter problem of video segmentation, improves the segmentation stability of the static image training model under video data, and achieves more accurate and consistent segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112949529B_ABST
    Figure CN112949529B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for improving the stability of target segmentation based on a loss function, comprising the following steps: S1, simulating a video sequence; S2, training a model; S3, fine-tuning a pre-trained model; S4, introducing a stability loss adjustment model. The detector constructed by the present invention example for a deep learning solution not only improves the accuracy, but also significantly improves the consistency of the same target at the boundary in consecutive frames under visual verification, reduces a large number of mis-segmentations and hole missed segmentations, improves the segmentation stability of a model trained on static images under video data, and effectively alleviates the jitter problem of video segmentation. The stability loss is introduced into the loss function, which improves the accuracy of segmentation and achieves the optimization of the segmentation accuracy of video data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of video image target segmentation and relates to a method for improving the stability of video image segmentation based on a loss function. Background Art

[0002] The iterative update of segmentation technology has made rapid progress in image segmentation. The effects of traditional segmentation methods that used digital image processing, topology, and mathematics have lagged far behind those based on convolutional neural networks and deep learning. Deep learning segmentation methods have gradually replaced traditional methods and applied them to various industries in display life. Image segmentation, as a basic image understanding task, has penetrated into daily life.

[0003] With the emergence of more and more large-scale manually annotated datasets, such as ImageNet classification, convolutional neural networks have shown excellent results in many basic areas of computer vision. However, for video analysis tasks, the cost of establishing a complete pixel-by-pixel intensive manually annotated dataset is very high. Given the segmentation mask estimation of a target instance in one or several frames, the same instance needs to be accurately segmented in the next video clip. Compared with target tracking, the difficulty of dataset annotation and pixel-by-pixel tracking undoubtedly bring huge challenges. Expensive annotation labor, narrow coverage, and rough annotation quality lead to the inability of trained models to obtain fine and stable segmentation results, which are often reflected in the jitter or inaccuracy of edge pixels. Some existing segmentation stability improvement methods, on the one hand, from the perspective of post-processing, algorithms such as CRF optimize the edge by weakening the correlation between features in the color space, but the long time consumption and algorithm limitations restrict the application of such methods. Many algorithms try to adjust the downsampling layer and loss layer to obtain accurate boundaries. Even if these algorithms have achieved good accuracy and performance, when predicting each frame in the video in sequence form, jitter problems or misaligned distortion usually occur at the boundaries. Summary of the invention

[0004] To solve the above problems, the technical solution of the present invention is a method for improving the stability of video image segmentation based on a loss function, comprising the following steps:

[0005] S1, simulated video sequence: simulate the motion of the target through linear geometric transformation, simulate the motion characteristics between video sequences, use self-supervision method, and do not introduce additional annotation information;

[0006] S2, training model: Use ImageNet data to train the parameters of the basic model so that the model has image classification prior information, and then fine-tune it on this basis. Then use the validation set to adjust the relevant hyperparameters of the model to determine the optimal hyperparameters of the final detector;

[0007] S3, fine-tuning of pre-trained models: Use test data to verify the performance of the trained model, visualize the detection effect of the model, and take corresponding optimization measures for hole and leak segmentation in different situations and scenarios to optimize the target segmentation effect;

[0008] S4 introduces a stability loss adjustment model: perturbations are introduced into the image to simulate target motion and generate images from different perspectives to fine-tune the network obtained in S3, achieve the underlying segmentation goals, reduce holes and missed segmentations, and enhance the consistency of segmentation boundaries.

[0009] Preferably, the linear geometric transformation includes rotation, scaling and translation.

[0010] Preferably, the step of introducing the stability loss adjustment model to perform perspective alignment on the disturbed image comprises the following steps:

[0011] The original image is transformed by matrix to achieve a large disturbance effect, and the matrix is ​​recorded as T b ;

[0012] The original image is transformed by matrix to achieve a small disturbance effect, and the matrix is ​​recorded as T S ;

[0013] The augmented images with different perturbations are superimposed on the original images in the batch_size dimension to form an input image of size 3*batch_size. The image is input into the network for forward propagation and the prediction results of the large perturbations are propagated through After the transformation matrix is ​​mapped to the corresponding perspective of the original image, T S The transformation matrix aligns the small perturbation view.

[0014] Preferably, after aligning the prediction result of the large disturbance image with the small disturbance perspective, the loss L is calculated using the small disturbance as the reference label. bs , and calculate the small perturbation and true label loss L at the same time sy , multiplied by the loss weights and added together to get the total loss L stab .

[0015] The beneficial effects of the present invention are as follows: In view of the stability problem of video image segmentation, the present invention proposes to use perturbations to simulate video frames and introduce stability loss, thereby improving the segmentation stability of the static image training model under video data and effectively alleviating the jitter problem of video segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a flowchart of the steps of a method for improving the stability of video image segmentation based on a loss function according to a specific embodiment of the method of the present invention. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0018] On the contrary, the present invention covers any substitution, modification, equivalent method and scheme made on the essence and scope of the present invention as defined by the claims. Further, in order to make the public have a better understanding of the present invention, some specific details are described in detail in the detailed description of the present invention below. Those skilled in the art can fully understand the present invention without the description of these details.

[0019] See also Figure 1 , is a flowchart of a method for improving the stability of video image segmentation based on a loss function according to an embodiment of the present invention, comprising the following steps:

[0020] S1, simulated video sequence: By using simple linear geometric transformations such as rotation, scaling and translation to simulate the motion of the target, the motion characteristics that may occur between video sequences are simulated, and a self-supervised approach is used to enable the network to learn temporal stability without introducing additional annotation information.

[0021] S2, training model: The parameters of the basic model trained using ImageNet data enable the model to have large-scale image classification prior information, and then fine-tune it on this basis, and then use the validation set to adjust the relevant hyperparameters of the model to determine the optimal hyperparameters of the final detector.

[0022] S3, pre-trained model fine-tuning: Use test data to verify the performance of the trained model, visualize the detection effect of the model, take different optimization measures for holes and missed segmentation in different situations and scenarios, and achieve targeted optimization of the target segmentation effect.

[0023] S4 introduces a stability loss adjustment model: perturbations are introduced into the image to simulate target motion and generate images from different perspectives to fine-tune the network obtained in S3, achieve better underlying segmentation targets, reduce holes and missed segmentations, and enhance the consistency of segmentation boundaries.

[0024] In a specific embodiment, S1 simulates the motion of the target by using simple linear geometric transformations such as rotation, scaling and translation, simulating the motion characteristics that may occur between video sequences, and adopts a self-supervised approach to enable the network to learn temporal stability without introducing additional annotation information.

[0025] Taking the rotation operation as an example, the original image coordinate system uses the upper left corner of the image as the origin, while the image center or custom coordinates are used as the far point during rotation. Therefore, the coordinate system needs to be transformed first, and the principle of keeping the distance between the point before and after rotation and the image focal point consistent must be followed during rotation. The expression is as follows:

[0026] Coordinate system conversion calculation method:

[0027]

[0028] Among them, center x Represents the distance between the left and right edges of the image and the center of the image, center y represents the distance between the upper and lower boundaries of the original image and the center point of the image, x, y represent the coordinates of the point on the image in the coordinate system with the upper left corner of the image as the origin, and x', y' represent the coordinates of the point on the original image in the Cartesian coordinate system with the center of the image as the origin.

[0029] After the rotation angle is determined, in the Cartesian coordinate system with the center of the image as the origin, for any pixel point (x0, y0), its distance to the far point is calculated, and the corresponding position (x1, y1) obtained after rotating by the set angle is obtained.

[0030] Coordinate rotation calculation method:

[0031]

[0032] Wherein, x″, y″ represent the coordinates of the corresponding point in the coordinate system with the center of the image as the origin after a rotation transformation, and α represents the angle of image rotation.

[0033] S2, training model: The parameters of the basic model trained using ImageNet data enable the model to have large-scale image classification prior information, and then fine-tune it on this basis, and then use the validation set to adjust the relevant hyperparameters of the model to determine the optimal hyperparameters of the final detector.

[0034] For model training, in order to prevent instability in the early stages of training, the stability loss L is not used first. bs , only the main loss L is used sy Train the basic model so that the model has basic segmentation capabilities.

[0035] Main loss L sy Calculation method:

[0036] L sy =||f(T s (n))-T s (t)||2

[0037] Where f represents the forward propagation of the network, R is the input image, T s represents weak perturbation, and t represents the label mask corresponding to the input image.

[0038] S3, pre-trained model fine-tuning: Use test data to verify the performance of the trained model, visualize the detection effect of the model, take different optimization measures for holes and missed segmentation in different situations and scenarios, and achieve targeted optimization of the target segmentation effect.

[0039] S4, optimizes the holes and missed segmentation in S3, introduces two different degrees of perturbations to the training image, and the large perturbation transformation matrix is ​​recorded as T b , the small perturbation matrix is ​​recorded as T s , and superimpose them in the batch_size dimension to form an input image of size 3*batch_size, which is fed into the network for forward propagation. The result of the large disturbance is passed through T b -1 After the transformation matrix is ​​mapped to the corresponding perspective of the original image, T s The transformation matrix aligns the small perturbation perspective and uses the small perturbation as the reference label to calculate the loss L bs , and calculate the small perturbation and true label loss L at the same time sy , perform a weighted summation of the two losses, multiply them by the loss weights respectively, and then calculate the total loss L stab .

[0040] The calculation method of aligning large disturbance with small disturbance is as follows:

[0041]

[0042] Total loss calculation method:

[0043] L stab =w1·L bs +w2·L sy

[0044] Where w1 and w2 represent the weights corresponding to the stability loss and the main loss respectively.

[0045] In order to prevent instability in the early stages of training, we first do not use stability loss and only use the main loss L sy Train the basic model so that it has basic segmentation capabilities. After adding stability loss, use the same network structure to load the basic model as the training initialization network parameters. This strategy can improve the consistency of the model's segmentation results under different time sequences. At the same time, the loss between weak perturbations and labels avoids the learning of erroneous features.

[0046] In order to balance the effects of different losses, it has been experimentally verified that setting the weight w1 to 0.3 is the most appropriate.

[0047] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for improving the stability of video image segmentation based on a loss function, characterized in that: The following steps are involved: S1, simulated video sequence: simulate the motion of the target through linear geometric transformation, simulate the motion characteristics between video sequences, use self-supervision method, and do not introduce additional annotation information; S2, training model: Use ImageNet data to train the parameters of the basic model so that the model has image classification prior information, and then fine-tune it on this basis. Then use the validation set to adjust the relevant hyperparameters of the model to determine the optimal hyperparameters of the final detector; For model training, in order to prevent instability in the early stages of training, only the main loss L is used. sy Train the basic model so that the model has basic segmentation capabilities, and the main loss L sy Calculation method: L sy =||f(T s (n))-T s (t)||2 Where f represents the forward propagation of the network, n is the input image, T s represents weak perturbation, t represents the label mask corresponding to the input image; S3, fine-tuning of pre-trained models: Use test data to verify the performance of the trained model, visualize the detection effect of the model, and take corresponding optimization measures for hole and leak segmentation in different situations and scenarios to optimize the target segmentation effect; S4, introduces a stability loss adjustment model: introduces perturbations to the image to simulate target motion and generate images from different perspectives to fine-tune the network obtained in S3, achieve the underlying segmentation goals, reduce holes and missed segmentations, and enhance the consistency of segmentation boundaries; The method of introducing the stability loss adjustment model to perform perspective alignment on the disturbed image comprises the following steps: The original image is transformed by matrix to achieve a large disturbance effect, and the matrix is ​​recorded as T b ; The original image is transformed by matrix to achieve a small disturbance effect, and the matrix is ​​recorded as T S ; The augmented images with different perturbations are superimposed on the original images in the batch_size dimension to form an input image of size 3*batch_size. The image is input into the network for forward propagation and the prediction results of the large perturbations are transmitted through After the transformation matrix is ​​mapped to the corresponding perspective of the original image, T S The transformation matrix is ​​aligned with the small perturbation perspective. After aligning the large perturbation image prediction result with the small perturbation perspective, the loss L is calculated using the small perturbation as the reference label. bs , and calculate the small perturbation and true label loss L at the same time sy , multiplied by the loss weights and added together to get the total loss L stab ; The calculation method of aligning large disturbance with small disturbance is: x′, y′ represent the coordinates of a point on the original image in a Cartesian coordinate system with the center of the image as the origin; x″, y″ represent the coordinates of the corresponding point in a coordinate system with the center of the image as the origin after a rotation transformation, and α represents the angle of image rotation; Total loss calculation method: 50 stab =w1·L bs +w2·L sy Where w1 and w2 represent the weights corresponding to the stability loss and the main loss respectively.

2. The method according to claim 1, characterized in that The linear geometric transformation includes rotation, scaling and translation.