Labeling method, device and equipment and computer medium
Through the combination of the instance segmentation model and reference semantic information, efficient annotation of object areas in the image is achieved, the problem of low manual annotation efficiency is solved, and the accuracy and efficiency of the annotation are improved.
Patent Information
- Application Number
- CN202311842783.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, the label data of manually labeling images takes a long time and the labeling efficiency is low.
A labeling method is provided, by obtaining the image to be analyzed, inputting an instance segmentation model to determine the object area, and marking the target semantic category of each object area in the image in combination with reference semantic information.
The accuracy of determining the edge of the object mask is improved, and the accuracy and labeling speed and efficiency of target semantic categories of each object area in the image are enhanced.
Smart Images

Figure CN120236071A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the field of machine learning, and particularly relates to a labeling method, apparatus, device, and computer medium. Background Art
[0002] In the related art, with the rapid development of computer vision models, the training data of computer vision models has become increasingly important. In the process of training a computer vision model based on training data, the label data in the training data is generally manually labeled. For example, for the label of the category corresponding to the pixels in an image, the annotator needs to spend a lot of manpower and energy to label the category of each pixel in an image. Manually labeling data takes a long time and has low labeling efficiency. There is a lack of an efficient image labeling method. Summary of the Invention
[0003] Embodiments of the present disclosure provide an implementation different from the related art to solve the technical problem that it takes a long time and has low labeling efficiency to manually label the label data of images in the related art.
[0004] In a first aspect, the present disclosure provides a labeling method, which includes:
[0005] Obtain an image to be analyzed;
[0006] Input the image to be analyzed into an instance segmentation model to obtain at least one object region in the image to be analyzed;
[0007] Determine reference semantic information;
[0008] Based on the at least one object region and the reference semantic information, label the target semantic category of each object region in the image to be analyzed.
[0009] In a second aspect, the present disclosure provides a labeling apparatus, including:
[0010] An obtaining unit, configured to obtain an image to be analyzed;
[0011] An input unit, configured to input the image to be analyzed into an instance segmentation model to obtain at least one object region in the image to be analyzed;
[0012] A determining unit, configured to determine reference semantic information;
[0013] A labeling unit, configured to label the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information.
[0014] In a third aspect, the present disclosure provides an electronic device, including:
[0015] A processor; and
[0016] A memory for storing executable instructions of the processor;
[0017] Wherein, the processor is configured to execute any method in the first aspect or any possible implementation manner of the first aspect by executing the executable instructions.
[0018] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any method in the first aspect or any possible implementation manner of the first aspect is implemented.
[0019] The present disclosure provides a solution for obtaining an image to be analyzed; inputting the image to be analyzed into an instance segmentation model to obtain at least one object region in the image to be analyzed; determining reference semantic information; and annotating the target semantic class of each object region in the image to be analyzed based on the at least one object region and the reference semantic information. The solution of the present application annotates the target semantic class of each object region in the image to be analyzed by analyzing at least one object region in the image to be analyzed and the reference semantic information. The present application determines a mask of an object, that is, an object region, through an instance segmentation model, which can improve the accuracy of determining the edge of the object mask, thereby improving the accuracy of determining the target semantic class of each object region in the image to be analyzed, and also improving the speed and efficiency of annotating the semantics of each object region in the image to be analyzed. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following will briefly introduce the drawings required for the description of the embodiments or related technologies. Obviously, the following drawings are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0021] Figure 1 It is a schematic structural diagram of a system provided by an embodiment of the present disclosure;
[0022] Figure 2 It is a schematic flowchart of a labeling method provided by an embodiment of the present disclosure;
[0023] Figure 3 It is a schematic structural diagram of a labeling device provided by an embodiment of the present disclosure;
[0024] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] Embodiments of the present disclosure will be described in detail below, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, and should not be construed as limiting the present disclosure.
[0026] The terms "first", "second", etc. in the description, claims and drawings of the embodiments of the present disclosure are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0027] The inventors have found through research that when training a computer vision model for identifying image content or classifying images, the corresponding label data relies on a large amount of manual annotation, which is not only inefficient but also costly. When annotating personnel annotate an image in units of pixels, it takes a annotating personnel 20-30 minutes to annotate one image. The annotating personnel need to spend a lot of manpower and energy on annotating the category of each pixel of a picture. For a picture with a 2K resolution, it takes a annotating personnel 10 days to annotate 2,000 pictures. The acquisition of label data consumes very high annotation costs and takes a lot of time, and the acquisition efficiency of label data is relatively low.
[0028] The technical solutions of the present disclosure and how the technical solutions of the present disclosure solve the above technical problems will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present disclosure will be described below with reference to the accompanying drawings.
[0029] Figure 1 FIG. 13 is a schematic structural diagram of a system provided for an exemplary embodiment of the present disclosure. The system includes: a terminal 10 and a server 11.
[0030] The terminal 10 is used to collect a plurality of images to be analyzed and store the plurality of images to be analyzed in the server 11; the server 11 can execute the annotation method in the present application to annotate the images to be analyzed.
[0031] Optionally, the above system may further include an annotation device 12 and a training device 13. In this case, the server 11 may be only used to store the aforementioned multiple images to be analyzed. The annotation device 12 can access each image to be analyzed stored in the server 11 through a URL (Uniform Resource Locator), and annotate each image to be analyzed. After the annotation is completed, the annotated image to be analyzed is fed back to the server 11. When the training device 13 performs model training based on the annotated image to be analyzed, it obtains the annotated image to be analyzed from the server 11.
[0032] Optionally, when the aforementioned system only includes the terminal 10 and the server 11, the server 11 includes a data hosting module and a training module, where: the data hosting module can access each image to be analyzed stored in the server 11 through a URL (Uniform Resource Locator), and annotate each image to be analyzed. After the annotation is completed, the annotated image to be analyzed is fed back to the server 11. When the training module 13 performs model training based on the annotated image to be analyzed, it obtains the annotated image to be analyzed from the server 11. Through this system, the trained model is kept synchronized with the annotated data and continuously improved. By gradually iterating and feedback loops, the biases and errors in the data are gradually corrected, improving the performance and accuracy of the model, while correcting the biases and errors in the data.
[0033] Among them, the aforementioned data hosting module is a resident service running on the server 11, used to support data annotation, facilitating the generation of annotated data. The server 11 can use workflows and triggers to control the training module to achieve automatic iteration of the model, easily tracing the relationship between the data and the model.
[0034] Experimental results have shown that this application effectively improves the performance metrics of downstream tasks. The solution proposed in this application can ensure that the model has broad adaptability and generalization ability by continuously collecting, integrating, and expanding data, so as to obtain more accurate and comprehensive training data.
[0035] Figure 2 FIG. is a schematic diagram of an annotation method provided by an exemplary embodiment of the present disclosure. This annotation method can be executed by the aforementioned server 11 or the aforementioned annotation device 12, and can also be executed by any electronic device other than Figure 1 This application does not make any limitations. This method at least includes the following steps S201 - S204:
[0036] S201. Obtain an image to be analyzed;
[0037] Optionally, the image to be analyzed is an image to be annotated. After being annotated, the image to be analyzed can be used as training data for a model used to classify images or determine image content. Specifically, the annotated image to be analyzed can be the label data in the training data, used to determine the loss information between the prediction result of the model for the image and the label data (actual result), facilitating the adjustment of the model's parameters. The format of the image to be analyzed is not limited in this application.
[0038] S202. Input the image to be analyzed into an instance segmentation model to obtain at least one object region in the image to be analyzed;
[0039] Optionally, the image to be analyzed may include one or more objects, and each object corresponds to an object region.
[0040] Optionally, each object region can be an instance mask.
[0041] Optionally, an object refers to the content in the image to be analyzed, such as: trees, birds, rivers, etc.
[0042] Optionally, for an object region corresponding to an object, the object region may include the mask value corresponding to the pixels corresponding to the instance of the object in the image to be analyzed, and the mask value can be 0 or 1.
[0043] S203. Determine the reference semantic information;
[0044] In some alternative embodiments of the present application, the reference semantic information is obtained by inputting the image to be analyzed into a preset semantic segmentation model, and the reference semantic information includes the pixel semantic categories of each pixel in the image to be analyzed.
[0045] Specifically, the semantic segmentation model can classify each pixel in the image to be analyzed into corresponding pixel semantic categories, and the pixel semantic categories are used to represent the semantics of the image content to which the corresponding pixels belong, such as, trees, birds, rivers, background information in the image, etc.
[0046] S204. Based on the at least one object region and the reference semantic information, label the target semantic categories of each object region in the image to be analyzed.
[0047] In some alternative embodiments of the present application, the instance segmentation model in the present application can perform instance segmentation on the image to be analyzed to obtain the instance mask of each object, that is, the aforementioned object region. The instance segmentation model can segment each object in the image to be analyzed into independent instances and generate a mask for each instance to represent the pixel-level position of the object. The aforementioned instance segmentation model can be SAM (Segment Anything Model).
[0048] Specifically, let I be the image to be analyzed, and let M n be the object region of the nth object, that is, the instance mask, and let f be an instance segmentation model (such as the SAM model). Then the instance mask M n of the nth object can be expressed as: M n = f(I, n).
[0049] For each object region, the target semantic category of the object region can be determined by performing pixel indexing in the semantic map and using a voting method. Specifically, according to the pixel positions in the object region, the corresponding pixels in the semantic map can be found, and the semantic categories to which these pixels belong are counted. Finally, through a voting mechanism, the target semantic category of the object region can be determined. Among them, the semantic map refers to the set of pixel semantic categories of each pixel in the image to be analyzed obtained by inputting the image to be analyzed into a preset semantic segmentation model.
[0050] Optionally, the set of pixel semantic categories of the foregoing pixels is the foregoing reference semantic information.
[0051] In some alternative embodiments of the present application, the reference semantic information includes at least one semantic category. In S204, the step of labeling the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information includes the following S2041-S2042:
[0052] S2041. For each object region among the at least one object region, count the number of pixels whose pixel semantic categories in the image to be analyzed belong to various semantic categories in the object region, and obtain the number of pixels corresponding to various semantic categories;
[0053] S2042. Take the semantic category corresponding to the largest number of pixels among the number of pixels as the target semantic category of the object region.
[0054] In some alternative embodiments of the present application, the method further includes the following S01-S02:
[0055] S01. Obtain the ratio of the largest number of pixels to the total number of pixels in the object region;
[0056] S02. If the ratio is greater than a preset threshold, then trigger the execution of taking the semantic category corresponding to the largest number of pixels among the number of pixels as the target semantic category of the object region.
[0057] In some alternative embodiments of the present application, the pixel semantic category of the pixel (i1, j1) in the image I to be analyzed can be expressed by the following formula:
[0058] c(i1, j1) = S(I)(i, j)
[0059] Among them, c(i1, j1), that is, S(I)(i1, j1), refers to the pixel semantic category of S(I) at the pixel position (i1, j1). S(I) is the semantic map output by the semantic analysis model. 0 < i1 < W, 0 < j1 < H, where W and H respectively represent the width and height of the image to be analyzed, and the value range of c is C. C is a set of at least one semantic category included in the reference semantic information.
[0060] In some alternative embodiments of the present application, for the instance mask M n The method for determining the target semantic category can be implemented by the following formula for the foregoing S2041 - S2042:
[0061]
[0062] Among them, Class(M n ) is the target semantic category of the instance mask M n . S(I) is the semantic map of the image to be analyzed. argmax c∈C represents a function that returns the semantic category c corresponding to the highest value. (i2, j2) represents a pixel coordinate in the instance mask M n . S(I)(i2, j2) represents the pixel semantic category at the pixel (i2, j2) in the semantic map S(I). I1(condition) is an indicator function that returns 1 if the condition is true and 0 otherwise.
[0063] Optionally, in the foregoing S204, the step of labeling the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information may further include, for each object region in the at least one object region, performing the following S20411 - S20423:
[0064] S20411. For each semantic category in at least one semantic category in the reference semantic information, count the number of pixels in the object region whose reference semantic information includes the semantic category, to obtain the number of pixels corresponding to the semantic category;
[0065] S20412. Determine the ratio of the number of pixels to the total number of pixels in the object region, to obtain the ratio corresponding to the semantic category;
[0066] S20423. Among the at least one ratio corresponding to the at least one semantic category, use the semantic category corresponding to the largest ratio as the target semantic category corresponding to the object region;
[0067] In some alternative embodiments of the present application, for the instance mask Mn The determination method of the corresponding target semantic category can be implemented by the following formula for the foregoing S20411-S20413:
[0068]
[0069] Wherein, Class(M n ) is the target semantic category corresponding to the instance mask M n , S(I) is the semantic map of the image to be analyzed, C is the set of the foregoing at least one semantic category, argmax c∈C represents a function that returns the semantic category c corresponding to the highest value, |M| is the total number of pixels in the instance mask M n , (i2, j2) represents the coordinate of a pixel in the instance mask M n , S(I)(i2, j2) represents the pixel semantic category at the pixel (i2, j2) in the semantic map S(I), and I1(condition) is an indicator function that returns 1 if the condition is true and 0 otherwise.
[0070] In some alternative embodiments of the present application, the reference semantic information includes an image feature information set and the semantic category corresponding to the image feature information set.
[0071] In some alternative embodiments of the present application, the image feature information set includes multiple groups of first image feature information;
[0072] Specifically, the image feature information set includes multiple groups of first image feature information corresponding to multiple preset objects, wherein each preset object corresponds to a group of first image feature information.
[0073] Optionally, the preset object can be any object that can be included in the image, such as a table, a rabbit, the sky, etc.
[0074] Specifically, for the determination method of the first image feature information corresponding to any one of the preset objects, the above method further includes:
[0075] Obtain an image of a preset object;
[0076] Extract features from the image of the preset object to obtain a set of feature vectors of the image corresponding to the preset object;
[0077] Use this set of feature vectors as the first image feature information corresponding to the preset object.
[0078] Optionally, the present application does not limit the manner of extracting features from the image of the preset object.
[0079] In some alternative embodiments of the present application, in the aforementioned S204, the step of annotating the target semantic categories of each object region in the image to be analyzed based on the at least one object region and the reference semantic information may include: for each object region, determining the image feature information that best matches the second image feature information according to the second image feature information corresponding to the object region and the multiple sets of first image feature information, and using the semantic category corresponding to the best-matching image feature information as the target semantic category of the object region.
[0080] Among them, determining the image feature information that best matches the second image feature information according to the second image feature information corresponding to the object region and the multiple sets of first image feature information means determining the image feature information that best matches the second image feature information among the multiple sets of first image feature information according to the second image feature information corresponding to the object region and the multiple sets of first image feature information.
[0081] Optionally, the image feature information that best matches the second image feature information among the multiple sets of first image feature information may refer to the image feature information with the highest similarity to the second image feature information among the multiple sets of first image feature information.
[0082] The similarity between the second image feature information and the first image feature information may refer to the cosine similarity between the second image feature information and the first image feature information. The determination method of the cosine similarity between the second image feature information and the first image feature information can be implemented by the following formula:
[0083]
[0084] Among them, f(s) refers to the second image feature information, f(r) refers to the first image feature information, cosine-similarity(f(s), f(r)) is the cosine similarity function, and m(s, r) is the cosine similarity between the second image feature information and the first image feature information.
[0085] In some alternative embodiments of the present application, the determination of the second image feature information corresponding to the object region includes:
[0086] Obtaining the image region information corresponding to the object region in the image to be analyzed;
[0087] Taking the image feature information of the image information in the image region information as the second image feature information corresponding to the object region.
[0088] In some alternative embodiments of the present application, determining the image feature information that best matches the second image feature information based on the second image feature information corresponding to the object region and the multiple sets of first image feature information may include the following S31 - S32:
[0089] S31. Determine a preset number of sets of candidate image feature information that match the second image feature information from the multiple sets of first image feature information;
[0090] S32. Based on the semantic categories corresponding to the preset number of sets of candidate image feature information, determine the image feature information that best matches the second image feature information among the preset number of sets of candidate image feature information.
[0091] In some alternative embodiments, the preset number of sets of candidate image feature information that match the second image feature information refers to: among the multiple sets of image feature information, the first image feature information corresponding to the preset number of highest similarities with the second image feature information. Specifically, the first image feature information corresponding to the preset number of highest similarities with the second image feature information in the image feature information set can be determined by KNN (k-Nearest Neighbor).
[0092] In some alternative embodiments, in the foregoing S31, determining a preset number of sets of candidate image feature information that match the second image feature information from the multiple sets of first image feature information may include:
[0093] For each set of first image feature information in the multiple sets of first image feature information, determine the similarity between the second image feature information and the first image feature information to obtain multiple similarities corresponding to the multiple sets of first image feature information;
[0094] Take the first image feature information corresponding to the preset number of highest similarities among the multiple similarities as the preset number of sets of candidate image feature information.
[0095] Optionally, each first image feature information and its corresponding semantic category may be stored in a database.
[0096] In some alternative embodiments, among the preset number of sets of candidate image feature information, the semantic categories corresponding to different sets of candidate image feature information may be the same or different.
[0097] Specifically, when among the preset number of sets of candidate image feature information, the semantic categories corresponding to some sets of candidate image feature information are the same, the number of sets of candidate image feature information is greater than the number of semantic categories.
[0098] In some alternative embodiments of the present application, in the foregoing S204, based on the at least one object region and the reference semantic information, marking the target semantic category of each object region in the image to be analyzed includes, for each object region in the at least one object region, performing the following S041-S042:
[0099] S41. Determine a preset number of groups of alternative image feature information that matches the second image feature information corresponding to the object region from the multiple groups of first image feature information;
[0100] S42. For each semantic category, count the number of groups of alternative image feature information corresponding to the semantic category in the preset number of groups of alternative image feature information to obtain the category number corresponding to the semantic category;
[0101] S43. Use the semantic category with the largest corresponding category number among the at least one semantic category as the target semantic category corresponding to the object region.
[0102] In some alternative embodiments, the foregoing S41-_S43 can be implemented by the following formula:
[0103]
[0104] Let M n be an object region of the nth object, D i3 be the semantic category of the i3th group of first image feature information, K be the preset number, and Class(M n ) represents the instance mask, that is, the object region, and M n 's target semantic category. I1(condition) is an indicator function that returns 1 if the condition is true and 0 otherwise.
[0105] Optionally, D can be a set of first image feature information stored in the database and / or the corresponding semantic categories, and K can also be the number of neighbors in the KNN algorithm.
[0106] In some alternative embodiments of the present application, in the foregoing S204, based on the at least one object region and the reference semantic information, marking the target semantic category of each object region in the image to be analyzed includes: for each object region in the at least one object region, based on the region type of the object region and the reference semantic information, marking the target semantic category of the object region in the image to be analyzed to obtain the target semantic category of each object region in the image to be analyzed.
[0107] In some alternative embodiments of the present application, the method further includes: determining the region type of the object region according to the instance segmentation model or annotation information.
[0108] Optionally, the region type may be a background type, such as sky, grassland, background information, etc.
[0109] Optionally, the region type may be a non-background type, such as people, animals, etc.
[0110] Optionally, the foregoing annotation information may be information manually annotated.
[0111] In some optional embodiments of the present application, when the region type is a background type, the foregoing reference semantic information includes the foregoing at least one semantic category.
[0112] In some optional embodiments of the present application, when the region type is a non-background type, the foregoing reference semantic information includes an image feature information set and a semantic category corresponding to the image feature information set.
[0113] After the target semantic categories of each object region in the image to be analyzed are annotated in the present application, the target semantic categories of each pixel in the image to be analyzed can be obtained. Specifically, the target semantic category of a pixel in the image to be analyzed is the target semantic category of the object region to which the pixel belongs.
[0114] Optionally, the semantic segmentation model proposed in the present application is replaced by any semantic segmentation model, such as: InternImage, Mask2Former, FCN (Fully Convolutional Networks), U-Net, SegNet, DeepLab, SPNet (Pyramid Scene Parsing Network), Mask R-CNN (Region-Convolutional Neural Networks), HRNet (High-Resolution Network), BiSeNet (Bilateral SegmentationNetwork), DANet (Dual Attention Network), DeeplabV3, etc.
[0115] The instance segmentation model proposed in this application can be any model of the same category, such as: segment-anything, FastSAM, MobieSAM, GrabCut, GraphCut, LazySnapping, Random Walks, GeodesicStar Convexity, ScribbleSup, SuperParsing, BSR (Bilateral Solver for Interactive Segmentation), SIOX (Simple Interactive Object Extraction), XDoG (Extended Difference of Gaussians).
[0116] The present disclosure provides a solution for obtaining an image to be analyzed; inputting the image to be analyzed into an instance segmentation model to obtain at least one object region in the image to be analyzed; determining reference semantic information; and annotating the target semantic categories of each object region in the image to be analyzed based on the at least one object region and the reference semantic information. The solution of this application annotates the target semantic categories of each object region in the image to be analyzed by analyzing at least one object region in the image to be analyzed and reference semantic information. By using an instance segmentation model to determine the mask of an object, that is, the object region, the accuracy of determining the edge of the object mask can be improved, thereby improving the accuracy of determining the target semantic categories of each object region in the image to be analyzed, and also improving the speed and efficiency of annotating the semantics of each object region in the image to be analyzed.
[0117] In addition, this application uses a semantic segmentation model to perform semantic segmentation on the entire picture to obtain a complete semantic map. The semantic segmentation model can assign each pixel in the image to a specific semantic category to achieve semantic understanding of the image to be analyzed.
[0118] Through the solution of this application, semantic understanding of the image to be analyzed and accurate segmentation of object instances can be achieved. By dividing into two branches, a semantic map and an instance mask of the entire picture can be obtained and associated with each other, so as to obtain accurate semantic information of each object. This method is expected to improve the accuracy of the object mask edge and provide a more reliable basis for subsequent image analysis and understanding tasks.
[0119] Figure 3 Schematic diagram of the structure of a data processing device provided for an exemplary embodiment of the present disclosure;
[0120] Among them, the device includes:
[0121] An acquisition unit 31 for acquiring an image to be analyzed;
[0122] An input unit 32 for inputting the image to be analyzed into an instance segmentation model to obtain at least one object region in the image to be analyzed;
[0123] A determination unit 33 for determining reference semantic information;
[0124] A labeling unit 34 for labeling the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information.
[0125] Optionally, the reference semantic information is obtained by inputting the image to be analyzed into a preset semantic segmentation model, and the reference semantic information includes the pixel semantic category of each pixel in the image to be analyzed.
[0126] Optionally, the reference semantic information includes at least one semantic category. When the device is used to label the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information, it is specifically used for:
[0127] For each object region in the at least one object region, count the number of pixels whose pixel semantic category in the image to be analyzed belongs to each of the semantic categories in the object region to obtain the number of pixels corresponding to each of the semantic categories;
[0128] Take the semantic category corresponding to the largest number of pixels among the number of pixels as the target semantic category of the object region.
[0129] Optionally, the device is further used for:
[0130] Obtain the ratio of the largest number of pixels to the total number of pixels in the object region;
[0131] If the ratio is greater than a preset threshold, then trigger the execution of taking the semantic category corresponding to the largest number of pixels among the number of pixels as the target semantic category of the object region.
[0132] Optionally, the reference semantic information includes an image feature information set and the semantic category corresponding to the image feature information set.
[0133] Optionally, the image feature information set includes multiple groups of first image feature information;
[0134] Optionally, when the device is used to label the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information, it is specifically used for:
[0135] For each object region, determine the image feature information that best matches the second image feature information corresponding to the object region based on the second image feature information corresponding to the object region and the multiple groups of first image feature information, and use the semantic category corresponding to the best-matching image feature information as the target semantic category of the object region.
[0136] Optionally, when the device is used to label the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information, it is specifically configured to:
[0137] For each object region in the at least one object region, label the target semantic category of the object region in the image to be analyzed based on the region type of the object region and the reference semantic information, so as to obtain the target semantic category of each object region in the image to be analyzed.
[0138] Optionally, the device is further configured to: determine the region type of the object region according to the instance segmentation model or the annotation information.
[0139] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they will not be elaborated here. Specifically, the device can execute the above method embodiments, and the foregoing and other operations and / or functions of each module in the device respectively correspond to the corresponding processes in each method in the above method embodiments. For the sake of brevity, they will not be elaborated here.
[0140] In the foregoing, the device of the embodiments of the present disclosure has been described from the perspective of functional modules. It should be understood that the functional module can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present disclosure can be completed by the integrated logic circuit in hardware in the processor and / or instructions in software. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.
[0141] Figure 4 is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. The electronic device may include:
[0142] A memory 401 and a processor 402, where the memory 401 is used to store a computer program and transmit the program code to the processor 402. In other words, the processor 402 can call and run the computer program from the memory 401 to implement the method in the embodiments of the present disclosure.
[0143] For example, the processor 402 can be used to execute the above method embodiments according to the instructions in the computer program.
[0144] In some embodiments of the present disclosure, the processor 402 may include but is not limited to:
[0145] A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.
[0146] In some embodiments of the present disclosure, the memory 401 includes but is not limited to:
[0147] A volatile memory and / or a non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0148] In some embodiments of the present disclosure, the computer program may be divided into one or more modules, which are stored in the memory 401 and executed by the processor 402 to complete the method provided by the present disclosure. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.
[0149] As Figure 4 shown, the electronic device may further include:
[0150] a transceiver 403, which may be connected to the processor 402 or the memory 401.
[0151] Among them, the processor 402 may control the transceiver 403 to communicate with other devices. Specifically, it may send information or data to other devices, or receive information or data sent by other devices. The transceiver 403 may include a transmitter and a receiver. The transceiver 403 may further include an antenna, and the number of antennas may be one or more.
[0152] It should be understood that the various components in the electronic device are connected through a bus system. Among them, the bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus.
[0153] The present disclosure also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by the computer, the computer can execute the method of the above method embodiment. Or, the embodiments of the present disclosure also provide a computer program product including instructions. When the instructions are executed by the computer, the computer executes the method of the above method embodiment.
[0154] When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present disclosure are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0155] According to one or more embodiments of the present disclosure, there is provided a labeling method, the method including:
[0156] Obtain an image to be analyzed;
[0157] Input the image to be analyzed into an instance segmentation model to obtain at least one object region in the image to be analyzed;
[0158] Determine reference semantic information;
[0159] Based on the at least one object region and the reference semantic information, label the target semantic category of each object region in the image to be analyzed.
[0160] According to one or more embodiments of the present disclosure, the reference semantic information is obtained by inputting the image to be analyzed into a preset semantic segmentation model, and the reference semantic information includes the pixel semantic category of each pixel in the image to be analyzed.
[0161] According to one or more embodiments of the present disclosure, the reference semantic information includes at least one semantic category, and the labeling the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information includes:
[0162] For each object region among the at least one object region, count the number of pixels in the image to be analyzed that belong to the object region and whose pixel semantic category is each of the semantic categories, to obtain the number of pixels corresponding to each of the semantic categories;
[0163] Take the semantic category corresponding to the largest number of pixels among the number of pixels as the target semantic category of the object region.
[0164] According to one or more embodiments of the present disclosure, the method further includes:
[0165] Obtain the ratio of the largest number of pixels to the total number of pixels in the object region;
[0166] If the ratio is greater than a preset threshold, then trigger the execution of taking the semantic category corresponding to the largest number of pixels among the number of pixels as the target semantic category of the object region.
[0167] According to one or more embodiments of the present disclosure, the reference semantic information includes an image feature information set and the semantic category corresponding to the image feature information set.
[0168] According to one or more embodiments of the present disclosure, the image feature information set includes multiple groups of first image feature information;
[0169] The step of labeling the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information includes:
[0170] For each object region, determine the image feature information that best matches the second image feature information corresponding to the object region according to the second image feature information corresponding to the object region and the multiple groups of first image feature information, and take the semantic category corresponding to the best-matched image feature information as the target semantic category of the object region.
[0171] According to one or more embodiments of the present disclosure, the step of labeling the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information includes:
[0172] For each object region among the at least one object region, label the target semantic category of the object region in the image to be analyzed based on the region type of the object region and the reference semantic information, to obtain the target semantic category of each object region in the image to be analyzed.
[0173] According to one or more embodiments of the present disclosure, the method further includes:
[0174] Determine the region type of the object region according to the instance segmentation model or the annotation information.
[0175] According to one or more embodiments of the present disclosure, a labeling device is provided, including:
[0176] An acquisition unit, configured to acquire an image to be analyzed;
[0177] An input unit, configured to input the image to be analyzed into an instance segmentation model to obtain at least one object region in the image to be analyzed;
[0178] A determination unit, configured to determine reference semantic information;
[0179] A labeling unit, configured to label the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information.
[0180] According to one or more embodiments of the present disclosure, the reference semantic information is obtained by inputting the image to be analyzed into a preset semantic segmentation model, and the reference semantic information includes the pixel semantic category of each pixel in the image to be analyzed.
[0181] According to one or more embodiments of the present disclosure, the reference semantic information includes at least one semantic category. When the device is used to label the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information, it is specifically configured to:
[0182] For each object region in the at least one object region, count the number of pixels whose pixel semantic category in the image to be analyzed belongs to each of the semantic categories of the object region to obtain the number of pixels corresponding to each of the semantic categories;
[0183] Take the semantic category corresponding to the largest number of pixels among the number of pixels as the target semantic category of the object region.
[0184] According to one or more embodiments of the present disclosure, the device is further configured to:
[0185] Obtain the ratio of the largest number of pixels to the total number of pixels in the object region;
[0186] If the ratio is greater than a preset threshold, then trigger the execution of taking the semantic category corresponding to the largest number of pixels among the number of pixels as the target semantic category of the object region.
[0187] According to one or more embodiments of the present disclosure, the reference semantic information includes an image feature information set and the semantic category corresponding to the image feature information set.
[0188] According to one or more embodiments of the present disclosure, the image feature information set includes multiple groups of first image feature information;
[0189] According to one or more embodiments of the present disclosure, when the device is used to label the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information, it is specifically configured to:
[0190] For each object region, determine the image feature information that best matches the second image feature information according to the second image feature information corresponding to the object region and the multiple sets of first image feature information, and use the semantic category corresponding to the best-matched image feature information as the target semantic category of the object region.
[0191] According to one or more embodiments of the present disclosure, when the device is used to label the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information, it is specifically configured to:
[0192] For each object region in the at least one object region, label the target semantic category of the object region in the image to be analyzed based on the region type of the object region and the reference semantic information, so as to obtain the target semantic category of each object region in the image to be analyzed.
[0193] According to one or more embodiments of the present disclosure, the device is further configured to: determine the region type of the object region according to the instance segmentation model or the annotation information.
[0194] According to one or more embodiments of the present disclosure, there is provided an electronic device, including:
[0195] A processor; and
[0196] A memory for storing executable instructions of the processor;
[0197] Wherein, the processor is configured to execute any one of the above methods by executing the executable instructions.
[0198] According to one or more embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, any one of the above methods is implemented.
[0199] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0200] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or modules can be in electrical, mechanical, or other forms.
[0201] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of the present disclosure, the various functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0202] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A marking method, characterized in that, The method includes: Obtaining an image to be analyzed; Inputting the image to be analyzed into an instance segmentation model to obtain at least one object region in the image to be analyzed; Determining reference semantic information; Based on the at least one object region and the reference semantic information, annotating the target semantic category of each object region in the image to be analyzed.
2. The method according to claim 1, wherein The reference semantic information is obtained by inputting the image to be analyzed into a preset semantic segmentation model, and the reference semantic information includes the pixel semantic category of each pixel in the image to be analyzed.
3. The method according to claim 2, characterized in that The reference semantic information includes at least one semantic category. The annotating the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information includes: For each object region in the at least one object region, counting the number of pixels whose pixel semantic category in the image to be analyzed belongs to each of the semantic categories in the object region to obtain the number of pixels corresponding to each of the semantic categories; Taking the semantic category corresponding to the largest number of pixels among the number of pixels as the target semantic category of the object region.
4. The method according to claim 3, characterized in that, The method further includes: Obtaining the ratio of the largest number of pixels to the total number of pixels in the object region; If the ratio is greater than a preset threshold, then triggering the execution of taking the semantic category corresponding to the largest number of pixels among the number of pixels as the target semantic category of the object region.
5. The method according to claim 1, characterized in that, The reference semantic information includes an image feature information set and the semantic category corresponding to the image feature information set.
6. The method according to claim 5, characterized in that The image feature information set includes multiple groups of first image feature information; The annotating the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information includes: For each object region, according to the second image feature information corresponding to the object region and the multiple groups of first image feature information, determining the image feature information that best matches the second image feature information, and taking the semantic category corresponding to the best-matched image feature information as the target semantic category of the object region.
7. The method according to claim 1, wherein The annotating the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information includes: For each object region in the at least one object region, based on the region type of the object region and the reference semantic information, annotating the target semantic category of the object region in the image to be analyzed to obtain the target semantic category of each object region in the image to be analyzed.
8. The method according to claim 7, wherein The method further includes: Determining the region type of the object region according to the instance segmentation model or the annotation information.
9. A labeling device, characterized in that, Includes: An obtaining unit for obtaining an image to be analyzed; An inputting unit for inputting the image to be analyzed into an instance segmentation model to obtain at least one object region in the image to be analyzed; A determining unit for determining reference semantic information; An annotating unit for annotating the target semantic category of each object region in the image to be analyzed based on the at least one object region and the reference semantic information.
10. An electronic device, characterized in that, Includes: A processor; And A memory for storing the executable instructions of the processor; Wherein, the processor is configured to execute the method according to any one of claims 1-8 by executing the executable instructions.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1-8 is implemented.