An automatic annotation method, device, and storage medium
Through the preset contour analysis model and feature point matching technology based on deep learning, the target objects in unlabeled image data are automatically marked, solving the problems of inefficiency and relying on manual accuracy in the existing technology, and achieving efficient product labeling.
Patent Information
- Application Number
- CN202111258064.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-27
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-10-27
AI Technical Summary
In the prior art, the automatic labeling of target objects without image data is inefficient, and the labeling results depend on the level of manual labeling, which cannot effectively ensure quality. Especially in the retail industry, manual labeling is laborious and has low accuracy when there are many types of products and different shapes.
The preset contour analysis model based on deep learning training is used to locate the unlabeled image data, determine the empty label of the target object, and automatically label it through feature point matching and single-item map, including positioning, cutting, clarity type determination and feature point matching steps to obtain the labeled data.
It has achieved significant savings in unsupervised and assisted manual labeling of shop products, improved the efficiency of inspection tasks, and ensured the quality of labeling.
Smart Images

Figure CN114092689B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information processing, and more particularly, to an automatic annotation method, apparatus, and storage medium. Background Art
[0002] In recent years, computer vision recognition technology based on deep learning has been widely applied in various industries. An excellent deep learning model requires a large amount of high-quality annotated data for support, and currently, almost all of these high-quality annotated data are obtained by manual annotation. However, the manual data annotation method is very inefficient, and whether the annotation result is accurate largely depends on the annotation level of the annotator. Therefore, the quality of data annotation by the manual annotation method cannot be effectively guaranteed.
[0003] In the retail industry, there are a wide variety of product categories and different shapes. The same category of products on the shelf may appear repeatedly, and the same product may be at different angles, making manual annotation rather laborious.
[0004] Therefore, an automatic annotation method is needed. Summary of the Invention
[0005] The problems to be solved by the present invention include how to obtain a cutout of a single target object from unannotated image data and annotate the target object. However, due to the diversity of unannotated image data and single target objects, there is no practical technical solution for automatically annotating target objects in the prior art.
[0006] To solve the above technical problems such as obtaining a cutout of a single target object from unannotated image data and annotating the target object, the present invention is proposed. Embodiments of the present invention provide an automatic annotation method, apparatus, and storage medium.
[0007] According to one aspect of an embodiment of the present invention, an automatic annotation method is provided, the method comprising:
[0008] Obtain unannotated image data, and use a preset contour analysis model to locate a first target object in the unannotated image data, and determine an empty label corresponding to each first target object;
[0009] Cut the unannotated image data based on the position information of the empty label to obtain the first target object;
[0010] Determine the clarity type of each first target object, and determine a second target object according to the clarity type;
[0011] Perform feature point matching between the second target object and the preset single product image to determine the number of feature point matches corresponding to the preset single product image and each second target object;
[0012] Select a preset number of second target objects based on the number of feature point matches, and label the selected preset number of second target objects according to the annotation information of the single product image to obtain annotation data.
[0013] Optionally, determining the clarity type of each first target object includes:
[0014] Use OpenCV to calculate the second derivative of each first target object to obtain the edge of each first target object, and calculate the variance of the edge;
[0015] Determine the clarity type of each first target object according to the variance corresponding to each first target object and the preset clarity threshold.
[0016] Optionally, the method further includes:
[0017] Before calculating the second derivative of the first target object, perform magnification processing on each first target object.
[0018] Optionally, the method further includes:
[0019] Perform mean blur processing and color enhancement processing on the second target object.
[0020] Optionally, the method further includes:
[0021] Based on the first target objects with the clarity type of the fuzzy type, train the initial fuzzy detection classification model to obtain an optimized fuzzy detection classification model.
[0022] Optionally, the method further includes:
[0023] Train the initial classification model according to the annotation data with different labels and the preset single product image corresponding to the annotation data to obtain an optimized classification model.
[0024] Optionally, the method further includes:
[0025] Obtain unlabeled third target objects, and use the optimized classification model to automatically label the third target objects to obtain annotation data.
[0026] According to another aspect of the embodiments of the present invention, there is also provided a storage medium, the storage medium includes a stored program, wherein, when the program runs, the above-mentioned method is executed by a processor.
[0027] According to another aspect of the embodiments of the present invention, an automatic annotation device is provided, and the device includes:
[0028] A positioning module, configured to obtain unannotated image data, and use a preset contour analysis model to locate a first target object in the unannotated image data, and determine an empty label corresponding to each first target object;
[0029] A first target object determination module, configured to cut the unannotated image data based on the position information of the empty label to obtain a first target object;
[0030] A second target object determination module, configured to determine the clarity type of each first target object, and determine a second target object according to the clarity type;
[0031] A matching module, configured to perform feature point matching between the second target object and a preset single-item image, and determine the number of feature point matches corresponding to the preset single-item image and each second target object;
[0032] An annotation data acquisition module, configured to select a preset number of second target objects based on the number of feature point matches, and perform annotation on the selected preset number of second target objects according to the annotation information of the single-item image to obtain annotation data.
[0033] According to another aspect of the embodiments of the present invention, an automatic annotation device is provided, and the device includes:
[0034] A processor; and
[0035] A memory, connected to the processor, and configured to provide instructions for the processor to perform the following processing steps:
[0036] Obtain unannotated image data, and use a preset contour analysis model to locate a first target object in the unannotated image data, and determine an empty label corresponding to each first target object;
[0037] Cut the unannotated image data based on the position information of the empty label to obtain a first target object;
[0038] Determine the clarity type of each first target object, and determine a second target object according to the clarity type;
[0039] Perform feature point matching between the second target object and a preset single-item image, and determine the number of feature point matches corresponding to the preset single-item image and each second target object;
[0040] Select a preset number of second target objects based on the number of feature point matches, and perform annotation on the selected preset number of second target objects according to the annotation information of the single-item image to obtain annotation data.
[0041] Embodiments of the present invention provide an automatic annotation method, apparatus, and storage medium. The first target object in the unannotated image data can be located according to a preset contour analysis model, an empty label corresponding to each first target object can be determined, and the target image in the unannotated image data can be quickly determined according to the position information of the empty label. Moreover, through feature point matching, the target object corresponding to the single-item map can be automatically annotated. The preset contour analysis model of the present invention is an identification model trained based on deep learning, and this identification model can well locate the first target object; based on the feature point matching between the single-item map and the target object, automatic annotation is realized, thereby obtaining annotation data, which can be used in the unsupervised assisted manual annotation of shop commodities, and can greatly save labor costs and time while improving the work efficiency of the detection task.
[0042] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] By referring to the following drawings, the exemplary embodiments of the present invention can be more fully understood:
[0044] Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing the method according to Embodiment 1 of the present invention;
[0045] Figure 2 is a flowchart of an automatic annotation method 200 according to the first aspect of Embodiment 1 of the present invention;
[0046] Figure 3 is a schematic structural diagram of an automatic annotation apparatus 300 according to Embodiment 2 of the present invention;
[0047] Figure 4 is a schematic structural diagram of an automatic annotation apparatus 400 according to Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited by the exemplary embodiments described herein.
[0049] It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present invention.
[0050] Those skilled in the art can understand that terms such as "first" and "second" in the embodiments of the present invention are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.
[0051] It should also be understood that in the embodiments of the present invention, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.
[0052] It should also be understood that for any component, data or structure mentioned in the embodiments of the present invention, without clear limitation or contrary revelation in the context, it can generally be understood as one or more.
[0053] In addition, the term "and / or" in the present invention is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the associated objects before and after.
[0054] It should also be understood that the present invention emphasizes the differences between the various embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be described in detail one by one.
[0055] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0056] The following description of at least one exemplary embodiment is actually only illustrative and in no way a limitation on the present invention and its application or use.
[0057] Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods and devices should be regarded as part of the specification.
[0058] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0059] Embodiments of the present invention can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate together with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.
[0060] Electronic devices such as terminal devices, computer systems, servers, etc. can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, target programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.
[0061] Embodiment 1
[0062] According to this embodiment, a method embodiment of an automatic annotation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0063] The method embodiment provided in this embodiment can be executed in a mobile terminal, a computer terminal, or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing the automatic annotation method is shown. As Figure 1 shown, the computer terminal 10 (or mobile device 10) can include one or more processors 102 (shown as 102a, 102b,..., 102n in the figure) (the processor 102 can include, but is not limited to, processing devices such as GPUs, microprocessor MCUs, or programmable logic devices FPGAs), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it can also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand,Figure 1 The structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than those shown in Figure 1 or have a configuration different from that shown in Figure 1 .
[0064] It should be noted that the above one or more processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of the present invention, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0065] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the automatic annotation method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizes the automatic annotation method of the above application program. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely provided with respect to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0066] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0067] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0068] It should be noted here that in some alternative embodiments, the aboveFigure 1 The computer device (or mobile device) shown may include hardware components (including circuits), software components (including computer code stored on a computer-readable medium), or a combination of both hardware and software components. It should be noted that Figure 1 is only an example of a specific concrete instance and is intended to illustrate the types of components that may exist in the above computer device (or mobile device).
[0069] Under the above operating environment, according to the first aspect of this embodiment, an automatic annotation method is provided. Figure 2 A flowchart of the method is shown. Referring to Figure 2 as shown, the method includes:
[0070] Step 201, obtain unannotated image data, and use a preset contour analysis model to locate the first target objects in the unannotated image data, and determine empty labels corresponding to each first target object.
[0071] In the embodiments of the present invention, the preset contour analysis model can be obtained in various ways. For example, through pre-trained models for commodity detection such as beverages, grain and oil, daily chemicals, and non-staple foods existing on Baidu PaddlePaddle, calculate the feature vectors of the convolutional layers of the pre-trained model for all training and test data, and then based on the feature vectors, a customized simplified version of the fully connected network is obtained, thereby obtaining the preset contour analysis model. Among them, the preset contour analysis model only distinguishes the rectangular contour frames of products and does not distinguish the categories of objects.
[0072] In the embodiments of the present invention, in the background management platform, upload the shelf images including commodities (i.e., unannotated image data) collected in advance in physical stores, and upload them in the form of a zip compression package. After the upload is completed, start the task, and use the preset contour analysis model to perform location processing on all unannotated commodities (i.e., the first target objects) in the unannotated image data, and generate corresponding empty labels. Through the preset contour analysis model set in advance in the present invention, unannotated commodities can be determined efficiently and quickly. The contour analysis model of the present invention can find more than 98% of the commodities on the shelf, and may find irrelevant results, but it has no impact on the subsequent results.
[0073] Among them, the images of the collected shelves can be obtained by various image acquisition devices. For example, the images of the shelves can be collected by an image acquisition device with a physical rotation function.
[0074] At the same time, when performing annotation, single-item images corresponding to different target objects are also required. At this time, the front and side views of the items from different perspectives can be taken in advance by a physical rotation device. Through the physical rotation device, the acquisition of single-item images can be realized with more standardized perspectives while keeping the camera fixed.
[0075] Step 202: Cut the unlabeled image data based on the position information of the empty label to obtain the first target object.
[0076] In an embodiment of the present invention, after obtaining a plurality of empty labels, the unlabeled image data can be cut based on the position information of the empty label, so as to obtain cut images of a plurality of unlabeled commodities, that is, the first target object.
[0077] For example, if the collected unlabeled image data includes 10 unlabeled commodities, first, through a preset contour analysis model, the positions of these 10 unlabeled commodities can be determined, and the position information of the empty label corresponding to each unlabeled commodity can be generated, and these 10 unlabeled commodities can be cut out from the unlabeled image data.
[0078] Step 203: Determine the clarity type of each first target object, and determine the second target object according to the clarity type.
[0079] Optionally, the determining the clarity type of each first target object includes:
[0080] Taking the second derivative of each first target object using OpenCV to obtain the edge of each first target object, and calculating the variance of the edge;
[0081] Determining the clarity type of each first target object according to the variance corresponding to each first target object and a preset clarity threshold.
[0082] Optionally, the method further includes:
[0083] Before taking the second derivative of the first target object, perform a magnification process on each first target object.
[0084] In an embodiment of the present invention, after obtaining the first target object, it is necessary to determine the clarity type of the first target object, and screen out the first target object that meets the requirements according to the clarity type as the second target object.
[0085] Specifically, use the dnn_superres module of OpenCV to load the LapSRN model, magnify the resolution of each first target object by 3 times, then take the second derivative of the magnified first target object to obtain the edge; then calculate the variance of the edge; finally, compare the variance with a preset clarity threshold to determine the clarity type of the first target object. Among them, the clarity type includes: clear and blurred. When the variance is less than or equal to the preset clarity threshold, the clarity type is determined to be clear; on the contrary, when the variance is greater than the preset clarity threshold, the clarity type is determined to be blurred.
[0086] When the clarity type is determined, the first target object can be classified, and the first target object with a clear clarity type is used as the second target object.
[0087] Optionally, the method further includes:
[0088] Based on the first target object with a fuzzy clarity type, the initial fuzzy detection and classification model is trained to obtain an optimized fuzzy detection and classification model.
[0089] In an embodiment of the present invention, the first target object with a fuzzy clarity type can also be screened out as a training set for training the initial fuzzy detection and classification model for determining the clarity type of the target object, so as to obtain an optimized fuzzy detection and classification model. Through the fuzzy detection and classification model, the fuzzy objects in the first target object can be directly determined, and thus the clear first target objects can be screened out as the second target objects.
[0090] Step 204: Perform feature point matching between the second target object and the preset single product image, and determine the number of feature point matches corresponding to the preset single product image and each second target object.
[0091] Optionally, the method further includes:
[0092] Perform mean blur processing and color enhancement processing on the second target image.
[0093] In an embodiment of the present invention, before performing feature point matching, mean blur processing can be performed on the obtained clear cut image (i.e., the second target object), which can make the points with relatively large noise in the cut image smoother and eliminate interference from other factors; then color enhancement processing is performed to adjust the brightness and enhance the color of the contour part to increase artificial recognition.
[0094] In an embodiment of the present invention, when performing feature point matching, based on OpenCV, each second target object that has undergone mean blur processing and color enhancement processing is respectively subjected to surf feature point matching with multiple single product images from different perspectives to obtain the number of feature point matches corresponding to each single product image and each second target object.
[0095] Step 205: Select a preset number of second target objects based on the number of feature point matches, and label the selected preset number of second target objects according to the annotation information of the single product image to obtain annotation data.
[0096] In an embodiment of the present invention, after obtaining the number of feature point matches, the second target objects are sorted according to the number of feature point matches; then, based on the sorting result, annotation is performed to obtain annotation data. Specifically, for any single product image, based on the feature point matching data corresponding to the second target objects and the single product image, the second target objects are sorted in descending order, and a preset number of the second target objects ranked at the front are selected. Then, the selected second target objects are annotated according to the classification label of the single product image, thereby obtaining annotation data.
[0097] In an embodiment of the present invention, manual annotation can also be performed. After obtaining the sorting of the number of feature point matches, an operator logs in to the background, enters the annotation list to select a single product image, and can view the sorting of the second target objects matching the single product image. Each cut image is accompanied by a checkbox. The operator only needs to check the checkboxes of a preset number of second target objects that conform to the product corresponding to the single product image and then click submit, and the second target objects can be annotated according to the classification label of the single product image, thereby obtaining annotation data. Optionally, the method further includes:
[0098] Training the initial classification model according to the annotation data with different labels and the preset single product images corresponding to the annotation data to obtain an optimized classification model.
[0099] Optionally, the method further includes:
[0100] Obtaining unannotated third target objects, and automatically annotating the third target objects by using the optimized classification model to obtain annotation data.
[0101] In an embodiment of the present invention, the obtained annotation data and the corresponding single product Figure 1 are added to ImageNet to train the initial classification model, and an image classification label is trained. When the training of all the labels to be recognized for a certain brand is completed, an optimized classification model can be obtained. After obtaining the optimized classification model for annotating cut images, when an unannotated data set (i.e., third target objects) is input and the target annotation number is set, the model can automatically annotate the data set. The target annotation number can be set between 200 and 250 to control the overall number of annotations to be more than 200 on average. When the annotated cut images are obtained, an operator can also check based on multiple annotated cut images of each product, make appropriate increases or decreases, and finally export the data set for model training. Whether it is classification based on the classification model or manual review, it is based on the same cut images of the product, and there is no need to care about the specific position of the annotation box on the data set.
[0102] The method of the present invention can save the cost of manually searching and annotating thousands of small cut images. After the number of cut images reaches a certain amount, the system trains a classification model, which can assist manual labor in more quickly sorting unannotated dataset images. Uniformly control the number of target objects in the annotation box to make the distribution of the training set for each product in the subsequent product detection model training relatively average. Finally, the annotated data is exported in the json format and can be uploaded to Baidu PaddlePaddle or used for offline training.
[0103] As described in the background art above, the manual data annotation method is very inefficient, and whether the annotation result is accurate largely depends on the annotation level of the annotator. Therefore, the quality of data annotation by the manual annotation method cannot be effectively guaranteed. For the retail industry, its commodity categories are numerous and the shapes are diverse. The same category of commodities on the shelves may appear repeatedly, and the same commodity may be at different angles, making it somewhat laborious and time-consuming based on manual annotation.
[0104] In view of the problems existing in the above background art, in this embodiment, unannotated image data is obtained, and a preset contour analysis model is used to locate a first target object in the unannotated image data to determine an empty label corresponding to each first target object; the unannotated image data is cut based on the position information of the empty label to obtain the first target object; the clarity type of each first target object is determined, and a second target object is determined according to the clarity type; the second target object and a preset single product image are subjected to feature point matching to determine the number of feature point matches corresponding to the preset single product image and each second target object; a preset number of second target objects are selected based on the number of feature point matches, and the selected preset number of second target objects are annotated according to the annotation information of the single product image, thereby obtaining annotated data.
[0105] Thus, in this way, the first target object in the unannotated image data can be located according to the preset contour analysis model, the empty label corresponding to each first target object can be determined, and the target image in the unannotated image data can be quickly determined according to the position information of the empty label. And through feature point matching, the target object corresponding to the single product image can be automatically annotated. The preset contour analysis model of the present invention is an identification model trained based on deep learning, and this identification model can well locate the first target object; based on the feature point matching between the single product image and the target object, automatic annotation is realized, thereby obtaining annotated data, which can be used for unsupervised auxiliary manual annotation of shop commodities, can greatly save labor costs and time, and improve the working efficiency of the detection task, solving the technical problems of the existing manual annotation being cumbersome, consuming manpower and material resources and having low accuracy. Specifically, in the embodiment of the present invention, the process of realizing automatic annotation includes:
[0106] Step A: Pretrain a contour analysis model to obtain target objects in the picture that may be products and assign empty labels to them.
[0107] Step B: On the background management platform, upload the shelf images collected in advance in the physical store and the single-product images taken from different front and side perspectives in advance through a physical rotation device. Among them, the end-cap data set is uploaded in the form of a zip compressed package. After the upload is completed, click to start the task. The system will use the contour analysis model to locate all unlabeled product objects in the product images and generate corresponding empty labels.
[0108] Step C: Cut out each target object based on the positioning information box of the empty label, determine the clarity type, and determine the second target object according to the clarity type;
[0109] Step D: Enlarge the resolution of each second target object, perform blur filtering, and adjust the brightness and enhance the color.
[0110] Step E: Match the enhanced cut-out objects with multiple single-product images from different perspectives using SURF feature points, and label the second target objects according to the number of feature point matches to obtain labeled data.
[0111] In addition, referring to Figure 1 As shown, according to the second aspect of this embodiment, a storage medium 104 is provided. The storage medium 104 includes a stored program, wherein the method described in any one of the above is executed by a processor when the program runs.
[0112] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0114] Example 2
[0115] Figure 3 The automatic annotation device 300 according to the present embodiment is shown. The device 300 corresponds to the method described in the first aspect of Embodiment 1. Refer to Figure 3 As shown, the device 300 includes:
[0116] A positioning module 301, configured to obtain unannotated image data, and use a preset contour analysis model to locate a first target object in the unannotated image data, and determine an empty label corresponding to each first target object;
[0117] A first target object determination module 302, configured to cut the unannotated image data based on the position information of the empty label to obtain a first target object;
[0118] A second target object determination module 303, configured to determine the clarity type of each first target object, and determine a second target object according to the clarity type;
[0119] A matching module 304, configured to perform feature point matching between the second target object and a preset single product image, and determine the number of feature point matches corresponding to the preset single product image and each second target object;
[0120] An annotation data acquisition module 305, configured to select a preset number of second target objects based on the number of feature point matches, and perform annotation on the selected preset number of second target objects according to the annotation information of the single product image to obtain annotation data.
[0121] Optionally, the second target object determination module 303 determines the clarity type of each first target object, including:
[0122] Using OpenCV to find the second derivative of each first target object to obtain the edge of each first target object, and calculating the variance of the edge;
[0123] Determining the clarity type of each first target object according to the variance corresponding to each first target object and a preset clarity threshold.
[0124] Optionally, the second target object determination module 303 further includes:
[0125] Before finding the second derivative of the first target object, perform a magnification process on each first target object.
[0126] Optionally, the second further includes:
[0127] Performing mean blur processing and color enhancement processing on the second target image.
[0128] Optionally, the device further includes:
[0129] A fuzzy detection and classification optimization model determination module, configured to train an initial fuzzy detection and classification model based on a first target object with a fuzzy clarity type to obtain a fuzzy detection and classification optimization model.
[0130] Optionally, the device further includes:
[0131] A classification optimization model determination module, configured to train an initial classification model based on labeled data with different labels and corresponding preset single-item images to obtain a classification optimization model.
[0132] Optionally, the device further includes:
[0133] An automatic annotation module, configured to obtain an unlabeled third target object and automatically annotate the third target object by using the classification optimization model to obtain labeled data.
[0134] Therefore, according to this embodiment, the first target object in the unlabeled image data can be located according to the preset contour analysis model, an empty label corresponding to each first target object can be determined, and the target image in the unlabeled image data can be quickly determined according to the position information of the empty label. Moreover, through feature point matching, the target object corresponding to the single-item image can be automatically annotated. The preset contour analysis model of the present invention is an identification model trained based on deep learning, and this identification model can well locate the first target object; based on the feature point matching between the single-item image and the target object, automatic annotation is realized to obtain labeled data, which can be used in the unsupervised assisted manual annotation of shop commodities, and can greatly save labor costs and time while improving the working efficiency of the detection task.
[0135] Embodiment 3
[0136] Figure 4 Fig. shows an automatic annotation device 400 according to this embodiment. The device 400 corresponds to the method described in the first aspect of Embodiment 1. Refer to Figure 4As shown, the device 400 includes: a processor 410; and a memory 420, connected to the processor 410, for providing instructions for the processor 410 to process the following processing steps: cutting the unlabeled image data based on the position information of the empty label to obtain a first target object; determining the clarity type of each first target object, and determining a second target object according to the clarity type; performing feature point matching between the second target object and a preset single-item image, and determining the number of feature point matches corresponding to the preset single-item image and each second target object; selecting a preset number of second target objects based on the number of feature point matches, and performing labeling on the selected preset number of second target objects according to the labeling information of the single-item image to obtain labeled data.
[0137] Optionally, the determining the clarity type of each first target object includes:
[0138] Taking the second derivative of each first target object using OpenCV to obtain the edge of each first target object, and calculating the variance of the edge;
[0139] Determining the clarity type of each first target object according to the variance corresponding to each first target object and a preset clarity threshold.
[0140] Optionally, the method further includes:
[0141] Before taking the second derivative of the first target object, performing magnification processing on each first target object.
[0142] Optionally, the method further includes:
[0143] Performing mean blur processing and color enhancement processing on the second target image.
[0144] Optionally, the method further includes:
[0145] Training an initial fuzzy detection classification model based on the first target objects with a fuzzy clarity type to obtain an optimized fuzzy detection classification model.
[0146] Optionally, the method further includes:
[0147] Training an initial classification model according to the labeled data with different labels and the preset single-item images corresponding to the labeled data to obtain an optimized classification model.
[0148] Optionally, the method further includes:
[0149] Obtaining unlabeled third target objects, and automatically labeling the third target objects using the optimized classification model to obtain labeled data.
[0150] Thus, according to this embodiment, the first target objects in the unlabeled image data can be located according to the preset contour analysis model, the empty labels corresponding to each first target object can be determined, and the target images in the unlabeled image data can be quickly determined according to the position information of the empty labels. Moreover, through feature point matching, the target objects corresponding to the single-item images can be automatically labeled. The preset contour analysis model of the present invention is an identification model trained based on deep learning, and this identification model can well locate the first target objects. Based on the feature point matching between the single-item images and the target objects, automatic labeling is achieved, thereby obtaining labeled data, which can be used for unsupervised assisted manual labeling of shop commodities, and can greatly save labor costs and time while improving the working efficiency of the detection task.
[0151] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0152] In the above embodiments of the present invention, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0153] In several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the units or modules can be in electrical or other forms.
[0154] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0155] In addition, each functional unit in the various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0156] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0157] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An automatic annotation method, characterized in that, The method includes: Obtain unlabeled image data, and use a preset contour analysis model to locate the first target object in the unlabeled image data, and determine an empty label corresponding to each first target object; Cut the unlabeled image data based on the position information of the empty label to obtain the first target object; Determine the clarity type of each first target object, and determine the second target object according to the clarity type; Perform feature point matching between the second target object and a preset single-item image, and determine the number of feature point matches between the preset single-item image and each second target object; Select a preset number of second target objects based on the number of feature point matches, and label the selected preset number of second target objects according to the annotation information of the single-item image to obtain labeled data; The determining the clarity type of each first target object includes: Use OpenCV to calculate the second derivative of each first target object to obtain the edge of each first target object, and calculate the variance of the edge; Determine the clarity type of each first target object according to the variance corresponding to each first target object and a preset clarity threshold; The method further includes: Before calculating the second derivative of the first target object, perform a magnification process on each first target object.
2. The method according to claim 1, wherein The method further includes: Perform mean blur processing and color enhancement processing on the second target object.
3. The method according to claim 1, wherein The method further includes: Based on the first target objects with the clarity type being the fuzzy type, train an initial fuzzy detection classification model to obtain an optimized fuzzy detection classification model.
4. The method according to claim 1, wherein The method further includes: Train an initial classification model according to the labeled data with different labels and the preset single-item images corresponding to the labeled data to obtain an optimized classification model.
5. The method according to claim 4, wherein The method further includes: Obtain unlabeled third target objects, and use the optimized classification model to automatically label the third target objects to obtain labeled data.
6. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program runs, the method according to any one of claims 1 to 5 is executed by a processor.
7. An automatic annotation device, characterized in that, The apparatus includes: A positioning module, configured to obtain unlabeled image data, and use a preset contour analysis model to locate the first target object in the unlabeled image data, and determine an empty label corresponding to each first target object; A first target object determination module, configured to cut the unlabeled image data based on the position information of the empty label to obtain the first target object; A second target object determination module, configured to determine the clarity type of each first target object, and determine the second target object according to the clarity type; A matching module, configured to perform feature point matching between the second target object and a preset single-item image, and determine the number of feature point matches between the preset single-item image and each second target object; A labeled data acquisition module, configured to select a preset number of second target objects based on the number of feature point matches, and label the selected preset number of second target objects according to the annotation information of the single-item image to obtain labeled data; The determining the clarity type of each first target object includes: Use OpenCV to calculate the second derivative of each first target object to obtain the edge of each first target object and calculate the variance of the edge; Determine the clarity type of each first target object according to the variance corresponding to each first target object and a preset clarity threshold; The device further includes: Before calculating the second derivative of the first target object, perform magnification processing on each first target object.
8. An automatic labeling device, characterized in that, The device includes: A processor; and A memory, connected to the processor, for providing instructions for the processor to perform the following processing steps: Obtain unlabeled image data, and use a preset contour analysis model to locate the first target object in the unlabeled image data to determine an empty label corresponding to each first target object; Cut the unlabeled image data based on the position information of the empty label to obtain the first target object; Determine the clarity type of each first target object, and determine the second target object according to the clarity type; Perform feature point matching between the second target object and a preset single-item image to determine the number of feature point matches between the preset single-item image and each second target object; Select a preset number of second target objects based on the number of feature point matches, and label the selected preset number of second target objects according to the annotation information of the single-item image to obtain annotation data; The determining the clarity type of each first target object includes: Use OpenCV to calculate the second derivative of each first target object to obtain the edge of each first target object and calculate the variance of the edge; Determine the clarity type of each first target object according to the variance corresponding to each first target object and a preset clarity threshold; The device further includes: Before calculating the second derivative of the first target object, perform magnification processing on each first target object.
Citation Information
Patent Citations
Training sample obtaining method and device, electronic device and storage medium
CN109753975A
Data labeling method, computer device and readable storage medium
CN112148685A