Image segmentation method, device, electronic device and storage medium

By using a pre-trained image segmentation model and weak label target box in image segmentation, combined with the reprocessing of binary graphs and the optimization of LOSS function based on probability graphs, the problems of poor image segmentation effect and dependence on labels are solved, and efficient image segmentation under weak supervision conditions are achieved.

CN119516199BActive Publication Date: 2025-06-27ZHIDAO NETWORK TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510066127.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-06-27
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The prior art has problems in image segmentation with poor segmentation effect and heavy dependence on labels, especially in the absence of segmentation capabilities in the absence of label data, and the application scenarios are limited.

Method used

An image segmentation method is adopted to input a pre-trained image segmentation model by obtaining image data containing the target to be segmented and a weak label target box. The binary graph obtained in this model is reprocessed according to the target's own characteristics and optimized using a LOSS function based on the probability graph and reprocessing.

Benefits of technology

Effective image segmentation is achieved under weak supervision, the accuracy of predicted binary graphs is optimized, and the LoRA parameters are updated through gradients, improving the model's segmentation ability and application scope of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119516199B_ABST
    Figure CN119516199B_ABST
Patent Text Reader

Abstract

The present application discloses an image segmentation method, apparatus, electronic device, and storage medium. The method includes obtaining image data containing a target to be segmented and a weak-label target box; inputting the image data containing the target to be segmented and the weak-label target box into a pre-trained image segmentation model, where the binary map obtained in the image segmentation model is reprocessed according to the self-characteristics of the target, and the LOSS function in the image segmentation model is calculated based on the probability map and the reprocessed binary map; and obtaining an image segmentation result according to the pre-trained image segmentation model. Through the present application, on the one hand, effective segmentation in the weak supervision scenario is achieved, and on the other hand, the universality of image segmentation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of image segmentation for autonomous driving, and in particular, to an image segmentation method, apparatus, electronic device, and storage medium. Background Art

[0002] Image segmentation is to divide an image into several specific regions with unique properties, and assign the same number to the pixels of the same connected region. Image segmentation mainly includes traditional methods and deep learning methods.

[0003] Traditional image segmentation includes algorithms such as threshold-based segmentation, region-based segmentation, and edge-based segmentation. However, based on traditional image segmentation algorithms, the segmentation effect is often poor, and the edge details are often not ideal.

[0004] Segmentation methods relying on deep learning include: single-modal segmentation and multi-modal segmentation algorithms. For the segmentation algorithm based on single-modal deep learning, although the effect is good and fine-grained segmentation can be achieved, it is highly dependent on labeled tags and does not have the ability to segment unlabeled data, so the application scenario is severely limited. For the multi-modal segmentation large model, it can not only achieve accurate segmentation results on labeled data, but also achieve segmentation ability on open-set data sets, greatly improving the universality of the segmentation model. Therefore, the segmentation algorithm based on single-modal deep learning has good effect and can achieve fine-grained segmentation, but it is highly dependent on labeled tags and does not have the ability to segment unlabeled data, and the application scenario is severely limited. Summary of the Invention

[0005] Embodiments of this application provide an image segmentation method, apparatus, electronic device, and storage medium to achieve effective segmentation in the case of weak supervision.

[0006] Embodiments of this application adopt the following technical solutions:

[0007] In a first aspect, embodiments of this application provide an image segmentation method, where the method includes:

[0008] Obtain image data containing the target to be segmented and a weak-label target box;

[0009] Input the image data containing the target to be segmented and the weak-label target box into a pre-trained image segmentation model. The binary map obtained in the image segmentation model is reprocessed according to the characteristics of the target itself, and the LOSS function in the image segmentation model is calculated based on the probability map and the reprocessed binary map;

[0010] Obtain an image segmentation result according to the pre-trained image segmentation model.

[0011] In some embodiments, the binary map obtained from the image segmentation model is reprocessed according to the self-characteristics of the target, which at least includes one of the following:

[0012] Perform morphological processing on the binary map obtained from the image segmentation model to remove noise points;

[0013] Extract the outer contour of each connected region from the binary map obtained from the image segmentation model, obtain the convex hull region and the area of each contour, compare the area with a preset threshold, and set it as the background if it is less than the preset threshold, otherwise set it as the foreground.

[0014] In some embodiments, the pre-trained image segmentation model at least includes an Image Encoder, a Prompt Encoder, and a Mask Decoder;

[0015] The Image Encoder is used to extract features after encoding the image data;

[0016] The Prompt Encoder is used to extract features after encoding the weak label target box information;

[0017] The Mask Decoder is used to fuse the features extracted after encoding by the Image Encoder and the Prompt Encoder, and obtain a prediction result;

[0018] The Image Encoder includes three branches, namely the Anchor branch, the Teacher branch, and the Student branch. The weak data augmentation method is used for the Anchor branch and the Teacher branch, and the strong data augmentation method is used for the Student branch.

[0019] In some embodiments, the binary map obtained from the image segmentation model is reprocessed according to the self-characteristics of the target, including:

[0020] Based on the binary map of the Anchor branch and the binary map of the Teacher branch , reprocess according to the self-characteristics of the target to obtain a new binary map of the Teacher branch and the binary map of the Anchor branch ,

[0021] The self-characteristics of the target at least include one of the following: the closed region characteristics of the target, the noise characteristics of the connected region of the target.

[0022] In some embodiments, the Teacher branch and the Student branch respectively output probability maps and , and the LOSS function in the image segmentation model is calculated based on the probability maps and the reprocessed binary maps, including:

[0023] Using and to calculate the focal loss and dice loss of the corresponding branch;

[0024] Using respectively and 、 to calculate the dice loss of the corresponding branch;

[0025] Based on the binary map obtain each instance feature on the feature map , based on the binary map obtain each instance feature on the feature map , calculate the similarity loss of the Teacher branch and the Anchor branch on the instance features and , the feature map is generated after encoding by the Teacher branch in the Image Encoder, and the feature map is obtained by the Anchor branch using the original structure of the Image Encoder and first freezing the structure and then encoding;

[0026] Sum all the losses after adding a specified coefficient, calculate the gradient to determine the LOSS function.

[0027] In some embodiments, the method further includes:

[0028] Add LoRA parameters to the Teacher branch and the Student branch in the Image Encoder for training, and at the same time share the LoRA parameters and freeze the original Image Encoder structure;

[0029] Input the weak-label target box into the Prompt Encoder for feature extraction;

[0030] The Mask Decoder respectively receives the feature information output by the three branches of the Image Encoder and the Prompt Encoder branch for feature fusion, and outputs the feature maps of the three branches after fusing the features.

[0031] In some embodiments, obtaining an image segmentation result according to the pre-trained image segmentation model includes:

[0032] The pre-trained image segmentation model merges the trained LoRA parameters into the ImageEncoder;

[0033] Performing a preset processing operation on the input image based on the longest side and the shortest side, and simultaneously performing a preset adjustment operation on the weak-label target box;

[0034] Inputting the image into the merged Image Encoder network, inputting the weak-label target box into the Prompt Encoder network structure, and outputting a binary Mask map through the Mask Decoder.

[0035] In some embodiments, after obtaining the image segmentation result according to the pre-trained image segmentation model, it further includes:

[0036] Obtaining target pixel coordinate information according to the target information in the image segmentation result;

[0037] Converting the target pixel coordinate information to the world coordinate system to obtain the coordinate information of the target in the world coordinate system;

[0038] Calibrating or correcting the extrinsic parameters of the roadside camera according to the coordinate information of the target in the world coordinate system and the target pixel coordinate information in the roadside camera.

[0039] In a second aspect, an image segmentation device is further provided in an embodiment of the present application, where the device includes:

[0040] An acquisition module for acquiring image data including a target to be segmented and a weak-label target box;

[0041] A processing module for inputting the image data including the target to be segmented and the weak-label target box into a pre-trained image segmentation model, where the binary map obtained in the image segmentation model is reprocessed by the self-characteristics of the target, and the LOSS function in the image segmentation model is calculated based on the probability map and the reprocessed binary map;

[0042] A segmentation module for obtaining an image segmentation result according to the pre-trained image segmentation model.

[0043] In a third aspect, an electronic device is further provided in an embodiment of the present application, including: a processor; and a memory arranged to store computer-executable instructions, where the executable instructions, when executed, cause the processor to execute the above method.

[0044] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, cause the electronic device to execute the above method.

[0045] The above at least one technical solution adopted in the embodiment of the present application can achieve the following beneficial effects: acquiring image data containing a target to be segmented and a weakly labeled target box, and then inputting the image data containing the target to be segmented and the weakly labeled target box into a pre-trained image segmentation model. The binary map obtained from the image segmentation model is reprocessed according to the characteristics of the target itself, and the LOSS function in the image segmentation model is calculated based on the probability map and the reprocessed binary map. Through the reprocessing of the binary image, the accuracy of the predicted binary map is optimized. At the same time, through the calculation of the loss function, the gradient only updates the LoRA parameters in the Image Encoder. Finally, according to the pre-trained image segmentation model, an image segmentation result is obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0047] Figure 1 is a schematic flowchart of the image segmentation method in an embodiment of the present application;

[0048] Figure 2 is a schematic structural diagram of the image segmentation device in an embodiment of the present application;

[0049] Figure 3 is a schematic processing flowchart of the image segmentation model in the image segmentation method in an embodiment of the present application;

[0050] Figure 4 is a schematic structural diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0052] Traditional image segmentation algorithms include threshold-based segmentation, region-based segmentation, edge-based segmentation, etc. Segmentation methods relying on deep learning include single-modal segmentation and multi-modal segmentation algorithms. Single-modal segmentation algorithms include algorithms such as U-Net and Mask R-CNN. Multi-modal segmentation algorithms include SAM, weSAM, segGPT, etc.

[0053] Based on traditional image segmentation algorithms, the segmentation effect is often poor, especially at the edge details. Based on single-modal deep learning segmentation algorithms, the effect is better and can also achieve refined segmentation. However, they rely heavily on labeled tags and do not have the ability to segment unlabeled data, severely limiting the application scenarios.

[0054] To address the deficiencies in the above image segmentation methods, in the embodiments of this application, a multi-modal segmentation large model is provided, which can not only achieve accurate segmentation results on labeled data but also achieve segmentation capabilities on open-set datasets.

[0055] The following will detail the technical solutions provided by each embodiment of this application in conjunction with the accompanying drawings.

[0056] The embodiments of this application provide an image segmentation method. As Figure 1 shown, the following is a schematic diagram of the image segmentation method process in the embodiments of this application. The method at least includes the following steps S110 to step S130:

[0057] Step S110, obtain image data containing the target to be segmented and weak label bounding boxes.

[0058] In the model prediction stage, obtain image data containing the target to be segmented. In particular, for weakly supervised segmentation, weak label bounding boxes also need to be input, which is consistent with the input in the model training stage.

[0059] Step S120, input the image data containing the target to be segmented and the weak label bounding boxes into a pre-trained image segmentation model. The binary map obtained in the image segmentation model is reprocessed according to the characteristics of the target itself. The LOSS function in the image segmentation model is calculated based on the probability map and the reprocessed binary map.

[0060] The pre-trained image segmentation model preferably adopts the weSAM segmentation large model. weSAM inherits the overall idea of SAM and mainly has three parts of the structure: Image Encoder, Prompt Encoder, Mask Decoder.

[0061] Furthermore, the binary map obtained in the image segmentation model is reprocessed according to the self-characteristics of the target, mainly aiming to further optimize the accuracy of the predicted binary map, so it depends on the geometric characteristics of the target itself. For example, when the target is a zebra crossing, it depends on the self-geometric characteristics of the zebra crossing.

[0062] Furthermore, the LOSS function in the image segmentation model is calculated based on the probability map and the reprocessed binary map. The LOSS function in the image segmentation model is obtained by statistically calculating multiple losses.

[0063] Specifically, in machine learning or deep learning, when it is necessary to perform a weighted sum of multiple loss functions to calculate the total loss and then perform gradient descent optimization, it can be achieved by assigning a coefficient to each loss function. This is usually used in multi-task learning, where the model needs to learn multiple related or unrelated tasks simultaneously. Each task's loss function can have a specific weight, so that their contributions to the total loss can be adjusted according to the importance or difficulty of the tasks.

[0064] It can be understood that the loss function, also known as the cost function, is an important concept for evaluating the performance of models in machine learning and deep learning. It is used to quantify the difference or error between the predicted value and the true value of the model. The main purpose of the loss function is to guide the training process of the model: by minimizing the loss function, we can find the optimal values of the model parameters, thereby improving the prediction accuracy of the model. During the training process, by continuously adjusting the model parameters (such as the weights in a neural network) to reduce the value of the loss function, the model gradually learns the internal laws and features of the data and can thus make more accurate predictions.

[0065] Step S130, obtain the image segmentation result according to the pre-trained image segmentation model.

[0066] The pre-trained image segmentation model performs corresponding processing operations on the input image data containing the target to be segmented based on the longest side and the shortest side. At the same time, the corresponding weak-label target boxes are adjusted accordingly. Then, in the merged Image Encoder network of the image input in the weSAM segmentation large model, the weak-label rectangular boxes are input into the Prompt Encoder network structure, and finally, a binary map Mask map is output through the Mask Decoder.

[0067] Through the above steps, the image data containing the target to be segmented and the weak-label target bounding box are input into a pre-trained image segmentation model. And according to the pre-trained image segmentation model, an image segmentation result is obtained. For the image segmentation model, the calculation of the binary map and the loss function is optimized, so that the effective segmentation of the image target can be achieved under the weak supervision condition by the above method.

[0068] Different from the related technologies, the method in the embodiment of the present application has general universality, providing a solution idea and an optimization method for the weak supervision image segmentation algorithm. Compared with the single-modal deep learning segmentation algorithm, it does not rely on labeled tags, and since the weak supervision tags are used, it does not have the ability to segment unlabeled data pages, thus improving the applicable range of the application scenarios.

[0069] In an embodiment of the present application, the post-processing of the binary map obtained in the image segmentation model according to the self-characteristics of the target at least includes one of the following: performing morphological processing on the binary map obtained in the image segmentation model to remove noise points; extracting the outer contour of each connected region of the binary map obtained in the image segmentation model, obtaining the convex hull region and the region area of each contour, comparing the region area with a preset threshold, and setting it as the background if it is less than the preset threshold, otherwise setting it as the foreground.

[0070] As Figure 2 shown, in order to further optimize the accuracy of the predicted binary map, relying on the geometric characteristics of the target itself, taking the target as a zebra crossing as an example, the set characteristics mainly include:

[0071] First, each zebra crossing is a single closed region, so there will be no problem of regional vacancy within the closed region of the zebra crossing.

[0072] Second, considering that in the process of image segmentation, there may be very small regions near the zebra crossing, which have a very small area relative to the zebra crossing, so there will be no noise points in the very small connected regions.

[0073] Based on the above two points, the Teacher binary map and the Anchor binary map are processed.

[0074] Case 1: Perform morphological opening operation to remove noise points.

[0075] Case 2: Extract the outer contour of each connected region from the binary image, and obtain the convex hull region and the region area of each contour. Specifically, in step b1), the convex hull region and the region area are obtained by using, including but not limited to, the convex hull algorithm or the connected region search algorithm. b2) By comparing the region area with a specified threshold, if it is less than the threshold, it is set as the background (non-zebra crossing), otherwise it is set as the foreground (zebra crossing). That is, specifically in the model, a new Teacher binary image is obtained. and the Anchor binary image .

[0076] In an embodiment of the present application, the pre-trained image segmentation model at least includes an Image Encoder, a Prompt Encoder, and a Mask Decoder; the Image Encoder is used to extract features after encoding the image data; the Prompt Encoder is used to extract features after encoding the weak label target box information; the Mask Decoder is used to fuse the features extracted after encoding by the Image Encoder and the Prompt Encoder, and obtain a prediction result; the Image Encoder includes three branches, namely the Anchor branch, the Teacher branch, and the Student branch. The weak data augmentation method is used for the Anchor branch and the Teacher branch, and the strong data augmentation method is used for the Student branch.

[0077] Specifically, in the image segmentation model, the Image Encoder encodes the image information to extract features. The Prompt Encoder encodes the rectangular box information of the data to extract features. In addition, points or rough segmentation masks can also be used as inputs. The Mask Decoder fuses the features encoded by the Image Encoder and the Prompt Encoder and gives a prediction result. weSAM has three Image Encoder branches, namely the Anchor branch, the Teacher branch, and the Student branch. It should be noted that during the training process, LoRA parameters are added and learned and updated on the basis of the Teacher branch and the Student branch, and the original Image Encoder parameters are frozen.

[0078] Preferably, as Figure 2As shown, it also includes two forms of data augmentation for image data, namely weak data augmentation (weak aug) and strong data augmentation (strong aug). Weak data augmentation uses operations including but not limited to horizontal flipping and vertical flipping with a probability of 0.5. Strong data augmentation includes but not limited to: formatting data to the uint8 type, histogram equalization, image sharpening, randomly changing image brightness and contrast, simulating image shadows, etc.

[0079] It can be understood that the above weak data augmentation (weak aug) and strong data augmentation (strong aug) are only examples and are not used to limit the protection scope of the embodiments of the present application. Specifically, the Anchor branch and the Teacher branch use weak data augmentation, and the Student branch uses strong data augmentation.

[0080] In an embodiment of the present application, the binary map obtained in the image segmentation model is reprocessed based on the self-characteristics of the target, including: the binary map based on the Anchor branch and the binary map of the Teacher branch , and a new binary map of the Teacher branch is obtained through reprocessing according to the self-characteristics of the target and the binary map of the Anchor branch ; the self-characteristics of the target at least include one of the following: the closed region characteristics of the target, the noise characteristics of the connected regions of the target.

[0081] According to the closed region characteristics of the target, such as the closed region of the zebra crossing. And according to the noise characteristics of the connected regions of the target, such as the characteristic that there will be no noise points in extremely small connected regions in the zebra crossing.

[0082] Based on the binary map of the Anchor branch and the binary map of the Teacher branch , a new binary map of the Teacher branch and the binary map of the Anchor branch are obtained after reprocessing according to the self-characteristics of the target and the binary map of the Anchor branch .

[0083] In an embodiment of the present application, the Teacher branch and the Student branch respectively output probability maps and , and the LOSS function in the image segmentation model is calculated based on the probability maps and the reprocessed binary maps, including: using and to calculate the focal loss and dice loss of the corresponding branch; using respectively and 、 Calculate the dice loss for the corresponding branch; based on the binary map Obtain on the feature map each instance feature , based on the binary map Obtain on the feature map each instance feature , calculate the similarity loss between the Teacher branch and the Anchor branch on the instance feature and . The feature map is generated after encoding by the Teacher branch in the Image Encoder, and the feature map is obtained by the Anchor branch using the original structure of the Image Encoder and first freezing the structure and then encoding; sum all the losses after adding a specified coefficient, and calculate the gradient to determine the LOSS function.

[0084] (1) Calculate the focal loss and dice loss according to the binary map of the Teacher branch and the probability map of the Student branch. As Figure 2 shown, L st .

[0085] (2) Calculate the dice loss according to the binary map of the Anchor branch and the probability map of the Student branch and the probability map of the Teacher branch. As Figure 2 shown, L anchor .

[0086] (3) Based on the binary map Obtain each instance feature on the feature map , based on the binary map Obtain each instance feature on the feature map and each instance feature on the feature map . That is, before entering the Mask Decoder, calculate the similarity loss between the Teacher branch and the Anchor branch on the instance feature.

[0087] Specifically, please refer to Figure 2 , use and to calculate the focal loss and dice loss. Use respectively and , to calculate the dice loss. Based on the binary map Obtain on the feature map Each instance feature above , based on the binary image obtain each instance feature on the feature map Each instance feature above . Sum all losses after multiplying them by a specified coefficient, and calculate the gradient. The gradient only updates the LoRA parameters in the Image Encoder. Based on the LoRA parameters, backpropagation is performed. Specifically, when updating, only the LoRA parameters are updated, and the frozen parameters are not updated.

[0088] It can be understood that Focal Loss is a loss specifically designed for class imbalance and was proposed in the paper "Focal Loss for Dense Object Detection" and applied to the training of object detection tasks. In fact, this is an idea independent of specific fields and can be applied to any data with class imbalance. Dice Loss comes from the paper "Dice Loss for Data-imbalanced NLP Tasks".

[0089] In an embodiment of the present application, the method further includes: adding LoRA parameters to the Teacher branch and the Student branch in the Image Encoder for training, while sharing the LoRA parameters and freezing the original Image Encoder structure; inputting the weak-label object boxes into the Prompt Encoder for feature extraction; the Mask Decoder respectively receives the feature information output from the three branches of the Image Encoder and the Prompt Encoder branch for feature fusion, and outputs the feature maps of the three branches after fusing the features.

[0090] As Figure 2 shown, the Image Encoder includes, from top to bottom in sequence: a Student branch, a Teacher branch, and an Anchor branch. In the embodiment of the present application, LoRA parameters are added to the Student branch and the Teacher branch, and no LoRA parameter is added to the Anchor branch.

[0091] First, add LoRA parameters to the Teacher branch and the Student branch in the Image Encoder network for training, and share the LoRA parameters, while freezing the original Image Encoder structure.

[0092] Then, the Teacher branch generates a feature map after encoding in the Image Encoder 。The Anchor branch uses the original Image Encoder structure with the structure frozen, and the encoded feature map is 。

[0093] Finally, the three branches also share parameters for the frozen Image Encoder.

[0094] Specifically, taking the zebra crossing as an example for illustration.

[0095] For the zebra crossing rectangle information, it is input into the Prompt Encoder structure for feature extraction; the MaskDecoder network receives the feature information output from the three Image Encoder branches and the Prompt Encoder branch respectively, and performs feature fusion. After fusing the features, the feature maps of the three branches are output.

[0096] Among them, the Teacher branch and the Student branch output probability maps and , the Teacher branch additionally outputs a binary map (foreground with a probability value greater than 0.5, otherwise background), and the Anchor branch outputs a binary map (foreground with a probability value greater than 0.5, otherwise background).

[0097] In an embodiment of the present application, according to the pre-trained image segmentation model, an image segmentation result is obtained, including: the pre-trained image segmentation model combines the trained LoRA parameters onto the ImageEncoder; performs a preset processing operation on the input image based on the longest side and the shortest side and simultaneously performs a preset adjustment operation on the weak label target box; inputs the image into the combined Image Encoder network, inputs the weak label target box into the Prompt Encoder network structure, and outputs a binary map Mask map through the Mask Decoder.

[0098] It can be understood that the image segmentation model includes three branches in the training stage. The Anchor branch is used to generate the ground truth, the Teacher branch is used for supervision, and the Student branch is for learning and is biased towards the ground truth. In the inference stage, the image segmentation model only contains the Student branch, processes the input image data including the target to be segmented and the weak label target box, and only needs to update the LoRA parameters in the Image Encoder during iteration.

[0099] When the image segmentation model performs inference, the image segmentation model combines the trained LoRA parameters onto the ImageEncoder. For the input image, including but not limited to, resize the longest side to a specified size and perform padding on the shortest side. At the same time, make corresponding adjustments to the weak-label bounding boxes input together. Finally, the image is input into the combined Image Encoder network, and the weak-label bounding boxes are input into the Prompt Encoder network structure, and a binary Mask image is output through the MaskDecoder.

[0100] In an embodiment of the present application, after obtaining the image segmentation result according to the pre-trained image segmentation model, it further includes: obtaining target pixel coordinate information according to the target information in the image segmentation result; converting the target pixel coordinate information to the world coordinate system to obtain the coordinate information of the target in the world coordinate system; calibrating or correcting the extrinsic parameters of the roadside camera according to the coordinate information of the target in the world coordinate system and the target pixel coordinate information in the roadside camera.

[0101] When applied to a specific scenario, it solves the problem that in the case of weak supervision, the image segmentation method can effectively segment based on the zebra crossing image data obtained by the vehicle-end or roadside camera.

[0102] Taking the target as the zebra crossing as an example, to achieve accurate calibration of the extrinsic parameters of the roadside camera, it relies on the RTK information on the vehicle side and the zebra crossing information recognized in the perception data of the on-vehicle camera, and converts the zebra crossing pixel coordinates to the world coordinate system by combining the VSLAM technology (Visual Simultaneous Localization and Mapping).

[0103] Combined with the zebra crossing pixel information in the roadside camera and the corresponding zebra crossing world coordinate information provided by the vehicle side, it assists in the calibration and correction of the extrinsic parameters of the roadside camera.

[0104] The VSLAM technology is a vision-based SLAM technology that uses a camera to obtain environmental image information while performing self-localization of the robot and environmental map construction.

[0105] The embodiment of the present application also provides an image segmentation device 300, as Figure 3 shown, which provides a schematic structural diagram of the image segmentation device in the embodiment of the present application. The image segmentation device 300 at least includes: an acquisition module 310, a processing module 320, and a segmentation module 330, where:

[0106] In one embodiment of the present application, the obtaining module 310 is specifically configured to: obtain image data including the target to be segmented and weak-label bounding boxes.

[0107] In the model prediction stage, image data including the target to be segmented is obtained. In particular, for weakly supervised segmentation, weak-label bounding boxes also need to be input, which is consistent with the input in the model training stage.

[0108] In one embodiment of the present application, the processing module 320 is specifically configured to: input the image data including the target to be segmented and the weak-label bounding boxes into a pre-trained image segmentation model. The binary map obtained from the image segmentation model is reprocessed according to the inherent characteristics of the target. The LOSS function in the image segmentation model is calculated based on the probability map and the reprocessed binary map.

[0109] The pre-trained image segmentation model preferably adopts the weSAM large model. weSAM inherits the overall idea of SAM and mainly has three parts of the structure: Image Encoder, Prompt Encoder, and Mask Decoder.

[0110] Furthermore, the binary map obtained from the image segmentation model is reprocessed according to the inherent characteristics of the target. The main purpose is to further optimize the accuracy of the predicted binary map, so it depends on the geometric characteristics of the target itself. For example, when the target is a zebra crossing, it depends on the inherent geometric characteristics of the zebra crossing.

[0111] Furthermore, the LOSS function in the image segmentation model is calculated based on the probability map and the reprocessed binary map. The LOSS function in the image segmentation model is obtained by statistically aggregating multiple losses.

[0112] Specifically, in machine learning or deep learning, when it is necessary to perform weighted summation of multiple loss functions to calculate the total loss and then perform gradient descent optimization, it can be achieved by assigning a coefficient to each loss function. This is usually used in multi-task learning, where the model needs to learn multiple related or unrelated tasks simultaneously. Each task's loss function can have a specific weight, so that their contributions to the total loss can be adjusted according to the importance or difficulty of the task.

[0113] It can be understood that the loss function, also known as the cost function, is an important concept for evaluating the performance of models in machine learning and deep learning. It is used to quantify the difference or error between the predicted values and the true values of the model. The main purpose of the loss function is to guide the training process of the model: by minimizing the loss function, we can find the optimal values of the model parameters, thereby improving the prediction accuracy of the model. During the training process, by continuously adjusting the model parameters (such as the weights in a neural network) to reduce the value of the loss function, the model gradually learns the internal laws and features of the data, and thus can make more accurate predictions.

[0114] In an embodiment of the present application, the segmentation module 330 is specifically configured to: obtain an image segmentation result according to the pre-trained image segmentation model.

[0115] The pre-trained image segmentation model performs corresponding processing operations on the input image data containing the target to be segmented based on the longest side and the shortest side. At the same time, the corresponding adjustment is made to the input weak-label target boxes. Then, in the merged Image Encoder network of the image input in the weSAM segmentation large model, the weak-label rectangular boxes are input into the Prompt Encoder network structure, and finally a binary Mask map is output through the Mask Decoder.

[0116] In an embodiment of the present application, the processing module 320 is further configured to: the binary map obtained from the image segmentation model is reprocessed by the self-characteristics of the target, at least including one of the following:

[0117] Perform morphological processing on the binary map obtained from the image segmentation model to remove noise points;

[0118] Extract the outer contour of each connected region from the binary map obtained from the image segmentation model, obtain the convex hull region and the region area of each contour, compare the region area with a preset threshold, and set it as the background if it is less than the preset threshold, otherwise set it as the foreground.

[0119] In an embodiment of the present application, the processing module 320 is further configured to divide the pre-trained image segmentation model into an Image Encoder, a Prompt Encoder, and a Mask Decoder;

[0120] The Image Encoder is used to encode and extract features from the image information;

[0121] The Prompt Encoder is used to encode and extract features from the target box information of the data;

[0122] The Mask Decoder is used to fuse the features encoded by the Image Encoder and the Prompt Encoder and obtain a prediction result;

[0123] The Image Encoder includes three branches, namely the Anchor branch, the Teacher branch, and the Student branch. A weak data augmentation method is used for the Anchor branch and the Teacher branch, and a strong data augmentation method is used for the Student branch.

[0124] In an embodiment of the present application, the processing module 320 is further configured to

[0125] Based on the binary map of the Anchor branch and the binary map of the Teacher branch , reprocess according to the self-characteristics of the target to obtain a new binary map of the Teacher branch and the binary map of the Anchor branch ;

[0126] The self-characteristics of the target include at least one of the following: the closed area characteristics of the target, the noise characteristics of the connected areas of the target.

[0127] In an embodiment of the present application, the Teacher branch and the Student branch output probability maps and , and the processing module 320 is further configured to

[0128] Use and to calculate the focal loss and dice loss of the corresponding branches;

[0129] Use respectively and , to calculate the dice loss of the corresponding branches;

[0130] Based on the binary map obtain each instance feature on the feature map , based on the binary map obtain each instance feature on the feature map , and calculate the similarity loss of the Teacher branch and the Anchor branch on the instance features and ;

[0131] Sum all losses after increasing them by a specified coefficient, calculate the obtained gradients, and determine the LOSS function.

[0132] In one embodiment of the present application, it further includes an update module for

[0133] During the training of the Teacher branch and the Student branch in the Image Encoder, add LoRA parameters, share the LoRA parameters, and freeze the original Image Encoder structure;

[0134] Input the weak-label target bounding box into the Prompt Encoder for feature extraction;

[0135] The Mask Decoder respectively receives the feature information output from the three branches of the Image Encoder and the Prompt Encoder branch for feature fusion, and outputs the feature maps of the three branches after fusing the features.

[0136] In one embodiment of the present application, the segmentation module 330 is further used for

[0137] The pre-trained image segmentation model combines the trained LoRA parameters onto the Image Encoder;

[0138] Perform a preset processing operation on the input image based on the longest side and the shortest side, and at the same time perform a preset adjustment operation on the weak-label target bounding box;

[0139] Input the image into the merged Image Encoder network and input the weak-label target bounding box into the Prompt Encoder network structure, and output a binary Mask map through the Mask Decoder.

[0140] In one embodiment of the present application, it further includes a post-processing module for

[0141] Obtain the target pixel coordinate information according to the target information in the image segmentation result;

[0142] Convert the target pixel coordinate information to the world coordinate system to obtain the coordinate information of the target in the world coordinate system;

[0143] Calibrate or correct the extrinsic parameters of the roadside camera according to the coordinate information of the target in the world coordinate system and the target pixel coordinate information in the roadside camera.

[0144] It can be understood that the above image segmentation device can implement each step of the image segmentation method provided in the foregoing embodiments. The relevant explanations regarding the image segmentation method are applicable to the image segmentation device and will not be elaborated here.

[0145] Figure 4 This is a schematic structural diagram of an electronic device according to an embodiment of the present application. Please refer to Figure 4 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.

[0146] The processor, network interface, and memory can be interconnected through an internal bus, and the internal bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 4 only a bidirectional arrow is used in

[0147] The memory is used to store programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory can include a memory and a non-volatile memory, and provides instructions and data to the processor.

[0148] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming an image segmentation device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:

[0149] Obtain image data containing the target to be segmented and weak label target boxes;

[0150] Input the image data containing the target to be segmented and the weak label target boxes into a pre-trained image segmentation model. The binary map obtained in the image segmentation model is reprocessed according to the characteristics of the target itself, and the LOSS function in the image segmentation model is calculated based on the probability map and the reprocessed binary map;

[0151] Obtain an image segmentation result according to the pre-trained image segmentation model.

[0152] The above as in this application Figure 1 The method executed by the image segmentation device disclosed in the embodiments shown above can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or by instructions in the form of software. The above processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0153] The electronic device can also execute Figure 1 the method executed by the image segmentation device in Figure 1 the embodiments shown, and implement the functions of the image segmentation device in

[0154] The embodiments of the present application also propose a computer-readable storage medium, which stores one or more programs, and the one or more programs include instructions that, when executed by an electronic device including a plurality of application programs, can enable the electronic device to execute Figure 1 the method executed by the image segmentation device in the embodiments shown, and specifically used to execute:

[0155] Obtain image data containing the target to be segmented and weak label bounding boxes;

[0156] Input the image data containing the target to be segmented and the weak-label target bounding box into a pre-trained image segmentation model. The binary map obtained from the image segmentation model is reprocessed according to the characteristics of the target itself, and the LOSS function in the image segmentation model is calculated based on the probability map and the reprocessed binary map.

[0157] Obtain an image segmentation result according to the pre-trained image segmentation model.

[0158] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0159] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0160] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, so that the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0162] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0163] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0164] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0165] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0166] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0167] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for image segmentation, wherein: The method comprises: Obtain image data containing the target to be segmented and the weakly labeled target frame; The image data containing the target to be segmented and the weakly labeled target frame are input into a pre-trained image segmentation model, the binary image obtained in the image segmentation model is reprocessed according to the characteristics of the target itself, and the LOSS function in the image segmentation model is calculated based on the probability map and the reprocessed binary image; The pre-trained image segmentation model at least includes an Image Encoder, a Prompt Encoder, and a Mask Decoder, wherein the Image Encoder includes three branches, namely, an Anchor branch, a Teacher branch, and a Student branch; The binary image obtained in the image segmentation model is processed again according to the characteristics of the target itself, including: Binary graph based on Anchor branch and the binary graph of the Teacher branch , according to the characteristics of the target itself, the binary graph of the new Teacher branch is obtained by reprocessing And the binary graph of the Anchor branch ; The target's own characteristics include at least one of the following: closed area characteristics of the target, noise characteristics of the connected area of ​​the target; The Teacher branch and the Student branch output probability maps respectively and The LOSS function in the image segmentation model is calculated based on the probability map and the reprocessed binary map, including: use and Calculate the focal loss and dice loss of the corresponding branch; use Respectively and , Calculate the dice loss of the corresponding branch; Based on binary image Get the feature map Each instance feature , based on binary graph Get the feature map Each instance feature , calculate the instance features of the Teacher branch and the Anchor branch and The similarity loss on the feature map It is generated by the Teacher branch after encoding in the Image Encoder. The feature map The Anchor branch uses the original structure of the Image Encoder and first freezes the structure and then encodes it; All losses are increased by a specified coefficient and summed up, and the gradient is calculated to determine the LOSS function; An image segmentation result is obtained according to the pre-trained image segmentation model.

2. The method of claim 1, wherein: The binary image obtained in the image segmentation model is processed according to the characteristics of the target and includes at least one of the following: The binary image obtained in the image segmentation model is subjected to morphological processing to remove noise points; The outer contour of each connected area is extracted from the binary image obtained in the image segmentation model, and the convex hull area and area of ​​each contour are obtained. The area is compared with a preset threshold. If it is less than the preset threshold, it is set as the background, otherwise it is set as the foreground.

3. The method of claim 1, wherein: The Image Encoder is used to encode the image data and then extract features; The Prompt Encoder is used to encode the weakly labeled target frame information and then extract features; The Mask Decoder is used to fuse the features extracted after encoding by the Image Encoder and the Prompt Encoder to obtain a prediction result; A weak data enhancement method is used for the Anchor branch and the Teacher branch, and a strong data enhancement method is used for the Student branch.

4. The method of claim 3, wherein: The method further comprises: In the Image Encoder, LoRA parameters are added to the Teacher branch and the Student branch for training, and the LoRA parameters are shared and the original Image Encoder structure is frozen; Input the weakly labeled target frame into the Prompt Encoder for feature extraction; The Mask Decoder receives feature information output from the three branches of the Image Encoder and the Prompt Encoder branch respectively to perform feature fusion, and outputs feature maps of the three branches after the feature fusion.

5. The method of claim 3, wherein: According to the pre-trained image segmentation model, an image segmentation result is obtained, including: The pre-trained image segmentation model merges the trained LoRA parameters into the Image Encoder; Perform preset processing operations on the input image based on the longest and shortest sides, and perform preset adjustment operations on the weakly labeled target box; The image is input into the merged Image Encoder network, the weak label target box is input into the Prompt Encoder network structure, and the binary image Mask image is output through the Mask Decoder.

6. The method of claim 1, wherein: After obtaining the image segmentation result according to the pre-trained image segmentation model, the method further includes: Obtaining target pixel coordinate information according to the target information in the image segmentation result; Convert the target pixel coordinate information into a world coordinate system to obtain the target coordinate information in the world coordinate system; The roadside camera external parameters are calibrated or corrected according to the coordinate information of the target in the world coordinate system and the target pixel coordinate information in the roadside camera.

7. An image segmentation device, wherein: The device comprises: An acquisition module is used to acquire image data containing the target to be segmented and a weakly labeled target frame; A processing module, used for inputting the image data containing the target to be segmented and the weakly labeled target frame into a pre-trained image segmentation model, wherein the binary image obtained in the image segmentation model is reprocessed according to the characteristics of the target itself, and the LOSS function in the image segmentation model is calculated based on the probability map and the reprocessed binary image; The pre-trained image segmentation model at least includes an Image Encoder, a Prompt Encoder, and a Mask Decoder, wherein the Image Encoder includes three branches, namely, an Anchor branch, a Teacher branch, and a Student branch; The binary image obtained in the image segmentation model is processed again according to the characteristics of the target itself, including: Binary graph based on Anchor branch and the binary graph of the Teacher branch , according to the characteristics of the target itself, the binary graph of the new Teacher branch is obtained by reprocessing And the binary graph of the Anchor branch ; The target's own characteristics include at least one of the following: closed area characteristics of the target, noise characteristics of the connected area of ​​the target; The Teacher branch and the Student branch output probability maps respectively and The LOSS function in the image segmentation model is calculated based on the probability map and the reprocessed binary map, including: use and Calculate the focal loss and dice loss of the corresponding branch; use Respectively and , Calculate the dice loss of the corresponding branch; Based on binary image Get the feature map Each instance feature , based on binary graph Get the feature map Each instance feature , calculate the instance features of the Teacher branch and the Anchor branch and The similarity loss on the feature map It is generated by the Teacher branch after encoding in the Image Encoder. The feature map The Anchor branch uses the original structure of the Image Encoder and first freezes the structure and then encodes it; All losses are increased by a specified coefficient and summed up, and the gradient is calculated to determine the LOSS function; The segmentation module is used to obtain an image segmentation result according to the pre-trained image segmentation model.

8. An electronic device, comprising: processor; as well as A memory arranged to store computer executable instructions, which, when executed, cause the processor to perform the method of any one of claims 1 to 5.

9. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, enables the electronic device to execute any one of the methods of claims 1 to 5.

Citation Information

Patent Citations

  • Stalk tissue anatomy characteristic parameter high-throughput extraction method based on deep learning

    CN113344008A

  • Box number identification method based on OCR technology

    CN115690807A