Identification method, measurement method and identification device based on tight frame mark
By constructing a deep learning network module based on tight-frame labels and utilizing the backbone network, segmentation network, and regression network, accurate recognition and measurement of targets are achieved, solving the problems of high pixel-level annotation data consumption and inaccurate boundary recognition in existing technologies and improving the accuracy of target measurement.
Patent Information
- Application Number
- CN202211058151.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-10-11
- Filing Date
- 2021-10-19
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2041-10-19
AI Technical Summary
Existing deep learning-based target recognition and measurement methods require precise pixel-level annotation data, which results in a large consumption of manpower and material resources, and the boundary recognition is not accurate enough, making it difficult to meet the needs of precise measurement.
A deep learning method based on tight-box labels is adopted. By constructing a network module including a backbone network, an image segmentation network with weak supervision learning, and a regression network for bounding box regression, the tight-box labels are used to train and identify the target, and the minimum enclosing rectangle of the target is obtained to achieve accurate measurement.
It reduces the time and labor cost of pixel-level data annotation and improves the accuracy of target recognition and measurement, especially for targets with little size change, and can achieve high-precision measurement.
Smart Images

Figure CN115423818B_ABST
Abstract
Claims
1. A recognition method based on tight frame mark, characterized in that: The invention relates to a recognition method for identifying a target using a network module trained based on a tight bounding box of the target, wherein the tight bounding box is the minimum bounding rectangle of the target, and the network module includes a segmentation network for image segmentation and a regression network based on bounding box regression. The recognition method includes: obtaining an input image including at least one target, wherein the at least one target belongs to at least one category of interest; inputting the input image into the network module to obtain a first output output by the segmentation network and a second output output by the regression network, wherein the first output includes the probability that each pixel in the input image belongs to each category, and the second output includes the offset between the position of each pixel in the input image and the tight bounding box of each category of the target; identifying the target based on the first output and the second output, wherein the network module also includes a backbone network, wherein the backbone network is used to extract a feature map of the input image, the segmentation network uses the feature map as input to obtain the first output, and the regression network uses the feature map as input to obtain the second output.
2. The identification method according to claim 1, wherein: The feature map has the same resolution as the input image.
3. The identification method according to claim 1, wherein: The offset in the second output is used as the target offset, where the target offset is normalized based on the average size of the targets of each category.
4. The identification method according to claim 1, wherein: The network module is trained by the following method: Constructing a training sample, wherein the input image data of the training sample includes a plurality of to-be-trained images, the plurality of to-be-trained images include images containing objects belonging to at least one category, and the label data of the training sample includes a gold standard for the category to which the object belongs and a gold standard for a tight-box label of the object; obtaining, by the network module, predicted segmentation data output by the segmentation network and predicted offset output by the regression network corresponding to the training sample based on the input image data of the training sample; The training loss of the network module is determined based on the label data corresponding to the training sample, the predicted segmentation data, and the predicted offset; and the network module is trained based on the training loss to optimize the network module.
5. The identification method according to claim 4, characterized in that: The determining of the training loss of the network module based on the label data corresponding to the training sample, the predicted segmentation data and the predicted offset includes: obtaining the segmentation loss of the segmentation network based on the predicted segmentation data and the label data corresponding to the training sample; obtaining the regression loss of the regression network based on the predicted offset corresponding to the training sample and the true offset corresponding to the label data, wherein the true offset is the offset between the position of the pixel point of the image to be trained and the gold standard of the tight frame mark of the target in the label data; and obtaining the training loss of the network module based on the segmentation loss and the regression loss.
6. The identification method according to claim 5, characterized in that: Using multi-instance learning, multiple to-be-trained bags are obtained based on the gold standard of the tight-frame labels of the targets in each to-be-trained image by category, and the segmentation loss is obtained based on the multiple to-be-trained bags of each category, wherein the multiple to-be-trained bags include multiple positive bags and multiple negative bags, and all pixel points on each of the multiple straight lines connecting the two opposite sides of the gold standard of the tight-frame label of the target are divided into a positive bag, the multiple straight lines include at least one group of mutually parallel first parallel lines and mutually parallel second parallel lines respectively perpendicular to each group of first parallel lines, and the negative bag is a single pixel point in the area outside the gold standard of the tight-frame labels of all targets of a category.
7. The identification method according to claim 4, characterized in that: According to the category and by using the expected intersection-and-union ratio corresponding to the pixel points of the to-be-trained image, the pixel points having the expected intersection-and-union ratio greater than the preset expected intersection-and-union ratio are screened out from the pixel points of the to-be-trained image to optimize the regression network.
8. A measurement method based on a tight frame mark, characterized in that: The invention relates to a measurement method for identifying a target based on the identification method according to any one of claims 1 to 7 to obtain a tight frame of each category of targets so as to measure the target, wherein the tight frame is the minimum circumscribed rectangle of the target.
9. A recognition device based on tight frame mark, characterized in that: The invention relates to a recognition device for recognizing a target using a network module trained based on a tight bounding box of the target, wherein the tight bounding box is the minimum bounding rectangle of the target. The recognition device comprises an acquisition module, a network module, and a recognition module. The acquisition module is configured to acquire an input image comprising at least one target, wherein the at least one target belongs to at least one category of interest. The network module is configured to receive the input image and obtain a first output and a second output based on the input image, wherein the first output comprises a probability that each pixel in the input image belongs to each category, and the second output comprises an offset between the position of each pixel in the input image and the tight bounding box of each category of the target. The network module comprises a segmentation network for image segmentation and a regression network based on bounding box regression, wherein the segmentation network is configured to output the first output, and the regression network is configured to output the second output. The recognition module is configured to recognize the target based on the first and second outputs. The network module further comprises a backbone network, wherein the backbone network is configured to extract a feature map of the input image, the segmentation network uses the feature map as input to obtain the first output, and the regression network uses the feature map as input to obtain the second output.
Citation Information
Patent Citations
Automatic driving target identification method based on improved Mask R-CNN
CN113111722A
Dynamic resolution instance segmentation method and computer readable storage medium
CN113111885A