Annotation apparatus and learning apparatus
By combining the image signal acquisition unit, image recognition unit, and training dataset generation unit, the training dataset is generated automatically or semi-automatically, solving the problem of high annotation workload and realizing an efficient and low-burden annotation process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-03-04
- Publication Date
- 2026-03-03
AI Technical Summary
In existing technologies, annotation tasks place a heavy burden on the person in charge, especially in annotation tasks involving large amounts of learning data, leading to excessive human workload.
It employs an image signal acquisition unit, an image recognition unit, and a learning dataset generation unit to generate the learning dataset through automated or semi-automated machine learning, reducing the need for manual annotation.
It has enabled the automation or semi-automation of annotation work, reducing the workload of the person in charge of annotation and improving annotation efficiency and accuracy.
Smart Images

Figure CN115176277B_ABST
Abstract
Description
Technical Field
[0001] This application relates to annotation devices and learning devices. Background Technology
[0002] Various techniques related to supervised learning have been developed in the past. In supervised learning, the learning data is pre-labeled. Patent Document 1 discloses a technique for predicting learning performance based on the state of the labels on the learning data.
[0003] Existing technical documents
[0004] Patent documents
[0005] Patent Document 1: International Publication No. 2018 / 079020 Summary of the Invention
[0006] The problem that the invention aims to solve
[0007] Typically, annotation of learning data is based on manual human work. Hereinafter, the person who annotates the learning data will sometimes be referred to as the "annotation lead." Additionally, the task of annotating the learning data will sometimes be called an "annotation task."
[0008] Previously, various techniques related to object detection have been developed in object recognition within computer vision. Additionally, various techniques related to scene segmentation have been developed. In object detection, tools such as "labellmg" are used for annotation tasks. In scene segmentation, tools such as "Labelbox" are used for annotation tasks.
[0009] Even with these tools, the workload for those responsible for annotation tasks will still increase due to the annotation work. This is especially true when annotation tasks are required for large amounts of training data, where the workload for those responsible for annotation can be quite heavy.
[0010] This application was completed to address the aforementioned issues, with the aim of reducing the workload of those responsible for annotation work.
[0011] Methods for solving problems
[0012] The annotation apparatus of this application includes: an image signal acquisition unit that acquires an image signal representing an image captured by a camera; an image recognition unit that is a machine learning-based learned image recognition unit that performs image recognition on the captured image; and a learning dataset generation unit that generates a learning dataset containing image data corresponding to each object and annotation data corresponding to each object by performing annotation on each object contained in the captured image based on the image recognition result.
[0013] Invention Effects
[0014] According to this application, due to the configuration described above, annotation work can be automated or semi-automated. As a result, the workload for those responsible for annotation can be reduced. Attached Figure Description
[0015] Figure 1 This is a block diagram showing the main parts of the annotation system of Embodiment 1.
[0016] Figure 2 This is a block diagram showing the main parts of the image recognition unit in the annotation device of Embodiment 1.
[0017] Figure 3 This is a block diagram showing the main parts of the learning database update unit in the learning device of Embodiment 1.
[0018] Figure 4 This is an illustrative diagram showing an example of a captured image.
[0019] Figure 5 It is shown that... Figure 4 An explanatory diagram of an example of the first feature map corresponding to the captured image shown.
[0020] Figure 6 This is an illustrative diagram showing examples of other captured images.
[0021] Figure 7 It is shown that... Figure 6 An explanatory diagram of an example of the first feature map corresponding to the captured image shown.
[0022] Figure 8 It is shown that... Figure 4 An explanatory diagram of an example of the second feature map corresponding to the captured image.
[0023] Figure 9 This is an illustrative diagram showing the structure of the neural network in "Mask R-CNN+GSoC".
[0024] Figure 10 It is shown that... Figure 4 An illustrative diagram of an example of the third feature map corresponding to the captured image.
[0025] Figure 11 This is an illustrative diagram showing the structure of the neural network in the first convolutional block of "Mask R-CNN+GSoC".
[0026] Figure 12 This is an explanatory diagram showing an example of recognition results based on object recognition using comparison.
[0027] Figure 13 This is an explanatory diagram showing an example of the recognition result based on the object recognition method of Implementation 1.
[0028] Figure 14 These are explanatory diagrams illustrating examples of recognition accuracy based on object recognition using comparison and examples of recognition accuracy based on object recognition according to Embodiment 1.
[0029] Figure 15 This is an illustrative diagram showing an example of a reliability map.
[0030] Figure 16 This is a block diagram showing the hardware structure of the main parts of the labeling device according to Embodiment 1.
[0031] Figure 17 This is a block diagram showing other hardware structures of the main parts of the labeling device in Embodiment 1.
[0032] Figure 18 This is a block diagram showing other hardware structures of the main parts of the labeling device in Embodiment 1.
[0033] Figure 19 This is a block diagram showing the hardware structure of the main parts of the learning device according to Embodiment 1.
[0034] Figure 20 This is a block diagram showing other hardware structures of the main parts of the learning device according to Embodiment 1.
[0035] Figure 21 This is a block diagram showing other hardware structures of the main parts of the learning device according to Embodiment 1.
[0036] Figure 22 This is a flowchart illustrating the operation of the labeling device in Embodiment 1.
[0037] Figure 23 This is a flowchart illustrating the operation of the learning device according to Embodiment 1.
[0038] Figure 24 This is a block diagram showing the main parts of another annotation system in Embodiment 1.
[0039] Figure 25 This is a block diagram showing the main parts of another annotation system in Embodiment 1.
[0040] Figure 26 This is a block diagram showing the main parts of the annotation system of Embodiment 2.
[0041] Figure 27 This is a flowchart illustrating the operation of the labeling device in Embodiment 2.
[0042] Figure 28 This is a block diagram showing the main parts of another annotation system in Embodiment 2.
[0043] Figure 29 This is a block diagram showing the main parts of another annotation system in Embodiment 2. Detailed Implementation
[0044] Hereinafter, in order to illustrate this application in more detail, the manner in which this application is implemented is described with reference to the accompanying drawings.
[0045] Implementation Method 1
[0046] Figure 1 This is a block diagram showing the main parts of the annotation system of Embodiment 1. Figure 2 This is a block diagram showing the main parts of the image recognition unit in the annotation device of Embodiment 1. Figure 3 This is a block diagram showing the main parts of the learning database update unit in the learning device of Embodiment 1. (Refer to...) Figures 1-3 The annotation system of Implementation Method 1 will be described.
[0047] like Figure 1 As shown, the annotation system 1 includes a camera 2, a storage device 3, a storage device 4, an annotation device 100, and a learning device 200. The storage device 3 has a learning dataset storage unit 11. The storage device 4 has a learning database storage unit 12. The annotation device 100 has an image signal acquisition unit 21, an image recognition unit 22, and a learning dataset generation unit 23. The learning device 200 has a learning database update unit 31 and a learning unit 32.
[0048] Camera 2 is a surveillance camera. Specifically, for example, camera 2 is a surveillance camera, a security camera, or a camera used for electronic mirrors. Camera 2 is composed of a visible light camera or an infrared camera, and is also composed of a camera for capturing moving images. Hereinafter, the individual still images that constitute the moving images captured by camera 2 will sometimes be referred to as "captured images".
[0049] The image signal acquisition unit 21 acquires an image signal representing a captured image. The image recognition unit 22 uses the acquired image signal to perform image recognition. Here, the image recognition performed by the image recognition unit 22 includes object recognition and tracking. Furthermore, the object recognition performed by the image recognition unit 22 includes at least one of object detection and region segmentation.
[0050] That is, such as Figure 2 As shown, the image recognition unit 22 includes a feature extraction unit 41, an object recognition unit 42, and an object tracking unit 43. The feature extraction unit 41 includes a first feature extraction unit 41_1 and a second feature extraction unit 41_2.
[0051] The first feature extraction unit 41_1 uses the image signal obtained above to generate a feature map (hereinafter, sometimes referred to as "first feature map") corresponding to each captured image. The first feature map is composed of a plurality of feature quantities (hereinafter, sometimes referred to as "first feature quantities") arranged in two mutually perpendicular directions.
[0052] Here, the first feature map corresponds to the foreground mask for each attribute. In this case, the first feature extraction unit 41_1 generates the first feature map, for example, by using the background subtraction method developed in GSoC (Google Summer of Code) 2017 to generate the foreground mask for each attribute. Figure 4 Examples of captured images are shown. Figure 5 This is the first feature map corresponding to the captured image, and an example of a first feature map based on background subtraction is shown. More specifically, Figure 5 This shows an example of a foreground mask corresponding to the attribute "person".
[0053] Alternatively, the first feature map corresponds to a mid-level feature, which corresponds to objectness. That is, each first feature quantity in the first feature map uses this mid-level feature. Furthermore, "mid-level" is the same level as the hierarchy in human-based visual models. In other words, "mid-level" is a lower level than the hierarchy of features used for existing object recognition.
[0054] For example, attention is used for intermediate-order features. In this case, the first feature extraction unit 41_1 generates an attention map, for example, through an attention mechanism, thus generating the first feature map. Figure 6 Examples of captured images are shown. Figure 7 This is the first feature map corresponding to the captured image, and an example of an attention-based first feature map is shown.
[0055] Alternatively, intermediate-order features may use saliency, for example. In this case, the first feature extraction unit 41_1 generates the first feature map, for example, by generating a saliency map using the same method as described in Reference 1 below. That is, the first feature extraction unit 41_1 generates the first feature map by performing saliency estimation.
[0056] [Reference 1]
[0057] International Publication No. 2018 / 051459
[0058] Furthermore, the intermediate-order features only need to correspond to objectivity, and are not limited to attention or saliency. Moreover, the method for generating the first feature map of the first feature extraction unit 41_1 is not limited to the specific example described above. For example, the first feature extraction unit 41_1 may also use at least one of image gradient detection, saliency estimation, background subtraction, objectivity estimation, attention, and region segmentation to generate the first feature map.
[0059] The following explanation will focus on the example where the first feature extraction unit 41_1 generates a foreground mask for each attribute using the background subtraction method.
[0060] The second feature extraction unit 41_2 uses the image signals obtained above to generate one or more feature maps (hereinafter sometimes referred to as "second feature maps") corresponding to each captured image. The second feature maps are formed sequentially, for example, using a convolutional neural network (hereinafter sometimes referred to as "CNN"). Each second feature map is composed of multiple feature quantities (hereinafter sometimes referred to as "second feature quantities") arranged in two mutually perpendicular directions.
[0061] Here, the second feature map corresponds to the high-level feature. That is, each second feature uses this high-level feature. Furthermore, "high-level" is the same level as the features used for existing object recognition. In other words, "high-level" is a level higher than the level of human-based visual models. Figure 8 Showing with Figure 4 The example shown is the second feature map corresponding to the captured image.
[0062] The object recognition unit 42 performs object recognition using the first feature map and the second feature map generated above. As described above, the object recognition performed by the object recognition unit 42 includes at least one of object detection and region segmentation.
[0063] In object detection, for each object contained in an image, its position is estimated through regression, and its attributes are estimated through classification. Object detection outputs information such as bounding boxes corresponding to coordinates (x, y, w, h), annotations corresponding to attributes, and reliability information for each bounding box, where the coordinates (x, y, w, h) correspond to position and size.
[0064] Region segmentation is the process of dividing a captured image into regions corresponding to various attributes. Through region segmentation, the captured image is divided into multiple regions, with each region measured in pixels. The output of region segmentation includes information representing the area of each region and information representing the attributes corresponding to each region.
[0065] Specifically, for example, the object recognition unit 42 performs both object detection and region segmentation using a masked R-CNN (Region-based CNN). The following explanation focuses on an example where a masked R-CNN is used in the object recognition unit 42. Masked R-CNN is described in reference 2 below.
[0066] [Reference 2]
[0067] Kaiming He, Georgia Gkioxari, Ross Girshick, et al. “Mask R-CNN,” v3, 24Jan 2018, https: / / arxiv.org / pdf / 1703.06870v3.pdf
[0068] Figure 9 This is a diagram showing the foreground mask for each attribute generated by the first feature extraction unit 41_1 using the background subtraction method, and an example of the structure of the neural network corresponding to the feature extraction unit 41 and the object recognition unit 42 when the object recognition unit 42 performs both object detection and region segmentation using Mask R-CNN. Hereinafter, this neural network is sometimes referred to as "Mask R-CNN+GSoC".
[0069] In the accompanying drawings, "GSoC background removal" corresponds to the first feature extraction unit 41_1. Furthermore, the CNN in "Fast R-CNN" within "Mask R-CNN" corresponds to the second feature extraction unit 41_2. Additionally, the block group located after the CNN in "Mask R-CNN" corresponds to the object recognition unit 42.
[0070] The CNN in "Fast R-CNN" within "Mask R-CNN" is, for example, a network composed of FPN (Feature Pyramid Networks) combined with ResNet (Residual Network)-101. Furthermore, as... Figure 9 As shown, the "mask" in "Mask R-CNN" has multiple convolutional blocks (in the attached figure, "conv.").
[0071] Figure 9 The neural network shown is pre-learned using an existing large-scale database. Specifically, for example, Figure 9The neural network shown is pre-learned using Microsoft COCO (Common Objects in Context). In other words, the image recognition unit 22 is pre-learned using this large-scale database.
[0072] Furthermore, the database used for learning by the image recognition unit 22 is not limited to Microsoft COCO. The image recognition unit 22 may also use, for example, a publicly available database based on "OpenAI" for prior learning. The following explanation will focus on the example of the image recognition unit 22 using Microsoft COCO for prior learning.
[0073] By using prior learning with this large-scale database, objects with learned shapes within captured images can be identified with high accuracy. Furthermore, object recognition with a certain level of accuracy can also be achieved for objects with unlearned shapes within captured images.
[0074] Here, in the object recognition of the object recognition unit 42, a feature map (hereinafter sometimes referred to as "the third feature map") formed by combining the first feature map and the second feature map is used, as described below. Furthermore, in the object recognition of the object recognition unit 42, this threshold is set to a lower value compared to existing object recognition (i.e., object recognition that uses the second feature map instead of the third feature map). Specific examples of the third feature map will be described below.
[0075] <The first specific example of the third feature map>
[0076] When a foreground mask is used in the first feature map, the object recognition unit 42 uses each first feature quantity in the first feature map to perform weighting on the corresponding second feature quantity in each second feature map. At this time, the object recognition unit 42 sets the value (hereinafter referred to as "importance") W representing the weight in the weighting as follows.
[0077] That is, the object recognition unit 42 calculates the similarity S between each first feature quantity in the first feature map and the corresponding second feature quantity in each second feature map. The similarity S is, for example, based on the value of at least one of EMD (Earth Mover's Distance), cosine similarity, KLD (Kullback-Leibler Divergence), L2 norm, L1 norm, and Manhattan distance.
[0078] Next, the object recognition unit 42 uses the calculated similarity S to set the importance W corresponding to each of the second features. At this time, for each of the second features, the larger the similarity S (i.e., the smaller the corresponding distance), the larger the importance W is set to. In other words, for each of the second features, the smaller the similarity S (i.e., the larger the corresponding distance), the smaller the importance W is set to.
[0079] By applying this weighting, the second feature quantity related to the region corresponding to the foreground object in the captured image is relatively enhanced compared to the second feature quantity related to the region corresponding to the background in the captured image. In other words, the second feature quantity related to the region corresponding to the background in the captured image is relatively weakened compared to the second feature quantity related to the region corresponding to the foreground object in the captured image. Multiple third feature maps corresponding to multiple first feature maps are generated in this manner.
[0080] Figure 10 An example of a third feature map generated in this manner is shown. Figure 10 The third feature map shown is... Figure 4 The images shown correspond to the captured images. That is, Figure 10 The third feature map shown is obtained by using... Figure 5 The first feature map shown is targeted Figure 8 It is generated by weighting the second feature map shown.
[0081] This weighting is performed, for example, by the first convolutional block in the "mask" of "mask R-CNN+GSoC". Figure 11 This example illustrates the structure of the neural network in the first convolutional block in this case. For example... Figure 11 As shown, the neural network has a weight calculation layer (in the attached figure, "Weight Calc."). Using this weight calculation layer, the importance W is set as described above.
[0082] <The second specific example of the third feature map>
[0083] When a foreground mask is used in the first feature map, the object recognition unit 42 performs element-wise multiplication operations on each first feature quantity in the first feature map and the corresponding second feature quantity in each second feature map to operate on the inner product.
[0084] By performing this operation, the second feature quantity related to the region corresponding to the foreground object in the captured image is relatively enhanced compared to the second feature quantity related to the region corresponding to the background object in the captured image. In other words, the second feature quantity related to the region corresponding to the background object in the captured image is relatively weakened compared to the second feature quantity related to the region corresponding to the foreground object in the captured image. Multiple third feature maps corresponding to multiple first feature maps are generated in this manner.
[0085] This operation is performed, for example, by the first convolutional block in the "Mask" of "Mask R-CNN+GSoC".
[0086] <The third specific example of the third feature map>
[0087] When using attention in the first feature map, the object recognition unit 42 uses each first feature quantity in the first feature map to perform weighting on the corresponding second feature quantity in each second feature map. At this time, the object recognition unit 42 sets the importance W as follows.
[0088] That is, the object recognition unit 42 uses GAP (Global Average Pooling) to select representative values from each of the second feature maps. The object recognition unit 42 then sets an importance W based on these selected representative values. In other words, the object recognition unit 42 sets the importance W to a value corresponding to the selected representative value.
[0089] By performing this weighting, multiple third feature maps corresponding to multiple second feature maps are generated. Alternatively, one third feature map corresponding to one second feature map is generated.
[0090] This weighting is performed, for example, by the first convolutional block in the "mask" of "Mask R-CNN+GSoC". Figure 11 Compared to the structure shown, the neural network in the first convolutional block in this case has a GAP layer instead of a weight calculation layer. Therefore, the importance W is set as described above.
[0091] By using the third feature map generated from the first, second, or third specific example for object recognition, compared to using the second feature map for object recognition, it is possible to avoid misidentification where a portion of the background is identified as an object. Furthermore, by using a lower threshold and suppressing misidentification as described above, object recognition can be achieved with high accuracy. In particular, the recognition accuracy for objects with unlearned shapes can be improved.
[0092] Additionally, typically, the first convolutional block in the "mask" of "Mask R-CNN+GSoC" includes a process for performing convolution (hereinafter, sometimes referred to as "process 1"), a process for performing deconvolution (hereinafter, sometimes referred to as "process 2"), and a process for performing point-wise convolution (hereinafter, sometimes referred to as "process 3"). The weighting of the first instance can be performed in process 1 or in process 3. The computation of the second instance can be performed in process 1 or in process 3. The weighting of the third instance can be performed in process 1 or in process 3.
[0093] That is, the weighting of the first specific example, the operation of the second specific example, or the weighting of the third specific example is sometimes preferably performed in the first step, or sometimes preferably performed in the third step, depending on the number of layers in the neural network. It is acceptable to select the more preferred step among these steps.
[0094] Hereinafter, object recognition that uses the third feature map to perform object detection and object recognition is sometimes referred to as "object recognition of embodiment 1". That is, object recognition of embodiment 1 uses "masked R-CNN+GSoC". In contrast, object recognition that uses the second feature map to perform object detection and region segmentation is sometimes referred to as "object recognition for comparison". That is, object recognition for comparison uses "masked R-CNN".
[0095] Figure 12 An example of recognition results based on object identification through comparison is shown. In contrast, Figure 13 An example of the recognition result based on object recognition according to Implementation Method 1 is shown. More specifically, Figure 13 Examples of recognition results related to the attribute "person" are shown. These recognition results are... Figure 4 The image shown corresponds to the captured image.
[0096] Here, refer to Figure 14 This paper explains the effects of using "masked R-CNN + GSoC". Specifically, it explains the improvement in object recognition accuracy compared to using "masked R-CNN".
[0097] Figure 14 The left half of the table shows experimental results based on the recognition accuracy of object recognition. In contrast, Figure 14 The right half of the table shows experimental results for the object recognition accuracy based on Implementation 1. These experiments used evaluation data from 5317 frames in the MOT16 benchmark.
[0098] The values in each column of the table represent mAP (mean Average Precision), expressed as a percentage. Furthermore, "Visibility > 0.X" in the table indicates that only objects whose entire body contains more than X% of the image portion are considered for identification. In other words, it means excluding objects whose entire body contains less than X% of the image portion from the identification pool.
[0099] like Figure 14 As shown, by using "Mask R-CNN+GSoC", the mAP value increases significantly compared to using "Mask R-CNN". That is, the accuracy of object recognition is greatly improved.
[0100] The object tracking unit 43 performs tracking of each object contained in the captured image by using the object recognition results of the object recognition unit 42 in a time sequence. This allows for the suppression of the decrease in recognition accuracy caused by changes in the appearance and shape of each object in the moving image captured by the camera 2.
[0101] That is, for example, if an object moves, its shape may change over time in the moving images captured by camera 2. In this case, the object's shape may be learned in one captured image at a certain moment, but not in another. Furthermore, the object may not be recognized using object recognition at the latter moment, thus making object recognition unstable over time.
[0102] In contrast, by performing tracking on the object, it can also be identified at the latter time. This allows the object's recognition to stabilize over time. Consequently, the accuracy of object recognition can be further improved.
[0103] The tracking by the object tracking unit 43 is as follows: Based on the object detection results for the captured image corresponding to the Nth frame (N is any integer), the object tracking unit 43 performs regression based on the KLD-based tracking loss, considering the attributes corresponding to each object, the coordinates corresponding to each object, and the foreground-to-background ratio in each small region. Thus, the object tracking unit 43 predicts the position and size of each object in the captured image corresponding to the (N+1)th frame.
[0104] Next, the object tracking unit 43 compares the prediction result with the object detection result for the captured image corresponding to the (N+1)th frame, and detects objects that were detected by object detection for the captured image corresponding to the Nth frame but not by object detection for the captured image corresponding to the (N+1)th frame. Thus, objects that are included in the captured image corresponding to the (N+1)th frame but were not detected by object detection can be continuously detected.
[0105] Furthermore, various known techniques can be used in the tracking process of the object tracking unit 43. Detailed descriptions of these techniques are omitted.
[0106] The learning dataset generation unit 23 generates a learning dataset corresponding to each object contained in the captured image based on the object recognition results of the object recognition unit 42 and the tracking results of the object tracking unit 43.
[0107] Here, the learning dataset includes data representing the images within the bounding boxes corresponding to each object (hereinafter referred to as "image data"), data representing the annotations corresponding to the attributes of each object (hereinafter referred to as "annotation data"), and data representing the masks corresponding to the regions corresponding to each object (hereinafter referred to as "mask data"). Generating this learning dataset can also be described as annotating each object contained in the captured image.
[0108] In addition, the learning dataset includes data (hereinafter referred to as "priority assignment data") for assigning priority P by the priority assignment unit 53, as described later. The priority assignment data includes, for example, data representing the object recognition reliability associated with each object (hereinafter referred to as "reliability data").
[0109] Furthermore, the data used for prioritization is not limited to reliability data. For example, the data used for prioritization may replace reliability data or, in addition to reliability data, include at least one of the following: data representing the size associated with each object, data representing the high-dimensional image features associated with each object, data representing the low-dimensional image features associated with each object, data representing the objectivity associated with each object, data representing the saliency estimation results associated with each object, and data representing the attention associated with each object.
[0110] The following explanation focuses on an example where the learning dataset contains image data, labeled data, masked data, and reliability data.
[0111] Here, as described above, the object recognition unit 42 uses the third feature map for object recognition. This avoids misidentification where a portion of the background is identified as an object. As a result, in the learning dataset generation unit 23, it avoids generating a learning dataset equivalent to an Easy Example in Focal Loss. That is, it avoids generating a learning dataset containing image data corresponding to the background. Therefore, in the relearning or additional learning of the image recognition unit 22 by the learning unit 32 (described later), the convergence of this learning can be accelerated.
[0112] The learning dataset storage unit 11 stores the learning dataset generated by the learning dataset generation unit 23. The learning database update unit 31 uses the learning dataset stored in the learning dataset storage unit 11 to update the learning database stored in the learning database storage unit 12.
[0113] That is, such as Figure 3 As shown, the learning database update unit 31 includes a learning dataset acquisition unit 51, a learning dataset acquisition unit 52, a priority assignment unit 53, and a learning dataset appending unit 54.
[0114] The learning dataset acquisition unit 51 acquires the learning dataset (hereinafter, sometimes referred to as "the first learning dataset") stored in the learning dataset storage unit 11. The learning dataset acquisition unit 52 acquires multiple learning datasets (hereinafter, sometimes referred to as "the second learning dataset") contained in the learning database stored in the learning database storage unit 12.
[0115] The priority assignment unit 53 assigns the priority P to the first learning dataset obtained above in the relearning or supplementary learning process of the learning unit 32, as described later. At this time, the priority assignment unit 53 assigns priority P based on the deviation of the distribution D in the multiple second learning datasets obtained above, in a manner that the dataset with higher learning value has a higher learning priority (i.e., the dataset with lower learning value has a lower learning priority).
[0116] Here, distribution D is a distribution based on priority assigned to the data. Specifically, for example, distribution D is a distribution in the reliability graph based on reliability data. Figure 15 An example of a reliability diagram is shown. In this case, the priority assignment unit 53 sets the priority P as follows, for example.
[0117] That is, the priority assignment unit 53 classifies the multiple second learning datasets obtained above into a dataset group that has accumulated a sufficient number of datasets with high reliability (hereinafter referred to as the "first dataset group"), a dataset group that has accumulated a certain number of datasets with high reliability (hereinafter referred to as the "second dataset group"), and a dataset group that has insufficient datasets with high reliability (hereinafter referred to as the "third dataset group") based on the deviation in the reliability map. This classification is based, for example, on the shape of the annotations represented by the labeled data (i.e., the attributes of the corresponding objects) or the shape of the mask represented by the mask data (i.e., the appearance shape of the corresponding objects).
[0118] Next, the priority assignment unit 53 determines which of the three datasets—the first dataset group, the second dataset group, and the third dataset group—the acquired first learning dataset should be classified into. This determination is based, for example, on the shape of the annotations represented by the labeled data (i.e., the attributes of the corresponding objects) or the shape of the mask represented by the mask data (i.e., the appearance shape of the corresponding objects).
[0119] If the first learning dataset obtained above should be classified into the first dataset group, it is considered to have low learning value. Therefore, the priority assignment unit 53 sets the priority P of the first learning dataset to a low value. Furthermore, if the first learning dataset obtained above should be classified into the second dataset group, it is considered to have moderate learning value. Therefore, the priority assignment unit 53 sets the priority P of the first learning dataset to a moderate value. Furthermore, if the first learning dataset obtained above should be classified into the third dataset group, it is considered to have high learning value. Therefore, the priority assignment unit 53 sets the priority P of the first learning dataset to a high value.
[0120] Furthermore, distribution D can be any distribution based on priority-assigned data, and is not limited to distribution based on reliability data. For example, distribution D can also be a distribution based on at least one of reliability, size, high-dimensional image features, low-dimensional image features, objectivity, saliency estimation, and attention.
[0121] Furthermore, the method by which the priority assignment unit 53 assigns priority P is not limited to the specific example described above. The priority assignment unit 53 may assign priority P in a manner that prioritizes datasets with higher learning value (i.e., in a manner that prioritizes datasets with lower learning value).
[0122] The learning dataset appending unit 54 generates a dataset (hereinafter sometimes referred to as "the third learning dataset") by appending data representing the priority P assigned above to the first learning dataset obtained above (hereinafter referred to as "priority data"). The learning dataset appending unit 54 updates the learning database by appending the generated third learning dataset to the learning database stored in the learning database storage unit 12.
[0123] Furthermore, the learning dataset addition unit 54 can also exclude the third learning dataset, which corresponds to a priority P less than a predetermined value, from the list of additions to the learning database. This avoids adding datasets with low learning value to the learning database.
[0124] Furthermore, the learning dataset addition unit 54 can reassign priority P to each of the second learning datasets in the same way that priority P is assigned to the first learning dataset. Therefore, the learning dataset addition unit 54 can also adjust the priority P in the learning database as a whole.
[0125] Furthermore, shortly after the system containing camera 2 (e.g., surveillance system, security system, or electronic mirror) begins operation, the learning database may not contain any learning data. In such cases, the learning database update unit 31 can also generate a new learning database by having the learning database storage unit 12 newly store the third learning dataset generated in the manner described above. Then, the learning database update unit 31 can also update the learning database by appending the newly generated third learning dataset to the learning database at any time. That is, the learning database update unit 31 can also generate and update the learning database.
[0126] The learning unit 32 uses the learning database stored in the learning database storage unit 12 (that is, the learning database updated by the learning database update unit 31) to perform relearning or additional learning of the image recognition unit 22. Hereinafter, relearning or additional learning will sometimes be referred to as "relearning, etc."
[0127] That is, as described above, the image recognition unit 22 performs pre-learning using an existing large-scale database. In addition, the image recognition unit 22 learns freely using the updated learning database described above. Therefore, the learning unit 32 performs relearning, etc., using the updated learning database described above, for the image recognition unit 22.
[0128] The relearning of the first feature extraction unit 41_1, for example, is based on supervised learning or unsupervised learning. Therefore, various well-known techniques related to supervised or unsupervised learning can be used in the relearning of the first feature extraction unit 41_1. Detailed descriptions of these techniques are omitted.
[0129] The relearning of the second feature extraction unit 41_2 is, for example, based on supervised learning. Therefore, various well-known techniques related to supervised learning can be used in the relearning of the second feature extraction unit 41_2. Furthermore, as mentioned above, the second feature extraction unit 41_2 uses a CNN. Therefore, the relearning of the second feature extraction unit 41_2 can also be based on deep learning. Therefore, various well-known techniques related to deep learning can be used in the relearning of the second feature extraction unit 41_2. Detailed descriptions of these techniques are omitted.
[0130] The relearning of the object recognition unit 42 is based, for example, on supervised learning. Therefore, various well-known techniques related to supervised learning can be used in the relearning of the object recognition unit 42. Detailed descriptions of these techniques are omitted.
[0131] Here, as described above, each learning dataset contained in the learning database is assigned a priority P. Therefore, the learning unit 32 can also vary the learning rate η in relearning, etc., according to the assigned priority P, for each learning dataset or each label. For example, the learning unit 32 can also increase the learning rate η as the assigned priority P increases (i.e., decrease the learning rate η as the assigned priority P decreases).
[0132] Alternatively, the learning unit 32 may perform data augmentation on a subset of the multiple learning datasets contained in the learning database, based on the assigned priority P. For example, the learning unit 32 may perform data augmentation on the learning dataset with the higher assigned priority P. Various known techniques can be used in data augmentation. Detailed descriptions of these techniques are omitted.
[0133] By setting the learning rate η or expanding the data, efficient relearning can be achieved using the learning database stored in the learning database storage unit 12 (i.e., a database smaller than a known large-scale database).
[0134] Furthermore, the updated learning database described above is smaller in size than the aforementioned known large-scale databases. Additionally, the updated learning database is based on images that differ from those contained in the aforementioned known large-scale databases (i.e., images captured by camera 2). Moreover, the updated learning database may contain annotations that differ from those contained in the aforementioned known large-scale databases.
[0135] Therefore, the relearning of the image recognition unit 22 by the learning unit 32 can also be based on transfer learning. In other words, various well-known techniques related to transfer learning can be used in the relearning of the image recognition unit 22 by the learning unit 32. Detailed descriptions of these techniques are omitted.
[0136] Furthermore, the relearning of the image recognition unit 22 by the learning unit 32 can also be based on fine-tuning. In other words, various known techniques related to fine-tuning can be used in the relearning of the image recognition unit 22 by the learning unit 32. Detailed descriptions of these techniques are omitted.
[0137] Furthermore, the relearning of the image recognition unit 22 by the learning unit 32 can also be based on Few-shot Learning. In other words, various well-known techniques related to Few-shot Learning can be used in the relearning of the image recognition unit 22 by the learning unit 32. Detailed descriptions of these techniques are omitted.
[0138] Furthermore, the relearning of the image recognition unit 22 by the learning unit 32 can also be based on meta-learning. In other words, various well-known techniques related to meta-learning can be used in the relearning of the image recognition unit 22 by the learning unit 32. Detailed descriptions of these techniques are omitted.
[0139] Furthermore, the relearning of the image recognition unit 22 by the learning unit 32 can also be based on distillation. In other words, various known techniques related to distillation can be used in the relearning of the image recognition unit 22 by the learning unit 32. Detailed descriptions of these techniques are omitted.
[0140] When the system containing camera 2 (such as a surveillance system, security system, or electronic mirror) is running, the learning unit 32 repeatedly performs relearning, thereby gradually adapting the image recognition of the image recognition unit 22 to the environment where camera 2 is located. As a result, the annotation accuracy of the learning dataset generation unit 23 gradually improves.
[0141] Hereinafter, the functions of the image signal acquisition unit 21 will be collectively referred to as "image signal acquisition function". In addition, the symbol "F1" will sometimes be used in this image signal acquisition function. Furthermore, the processing performed by the image signal acquisition unit 21 will sometimes be collectively referred to as "image signal acquisition processing".
[0142] Hereinafter, the functions of the image recognition unit 22 will be collectively referred to as "image recognition function". In addition, the symbol "F2" will sometimes be used in this image recognition function. Furthermore, the processing performed by the image recognition unit 22 will sometimes be collectively referred to as "image recognition processing".
[0143] Hereinafter, the functions of the learning dataset generation unit 23 will sometimes be referred to as the "learning dataset generation function". Furthermore, the symbol "F3" will sometimes be used in this learning dataset generation function. Additionally, the processing performed by the learning dataset generation unit 23 will sometimes be referred to as the "learning dataset generation processing".
[0144] Hereinafter, the functions of the learning database update unit 31 will sometimes be referred to as "learning database update function". Furthermore, the symbol "F11" will sometimes be used in this learning database update function. Additionally, the processing performed by the learning database update unit 31 will sometimes be referred to as "learning database update processing".
[0145] Hereinafter, the functions of the learning unit 32 will sometimes be referred to as "learning functions". In addition, the symbol "F12" will sometimes be used in this learning function. Furthermore, the processing performed by the learning unit 32 will sometimes be referred to as "learning processing".
[0146] Next, refer to Figures 16-18 The hardware structure of the main parts of the labeling device 100 will be described.
[0147] like Figure 16 As shown, the labeling device 100 includes a processor 61 and a memory 62. The memory 62 stores programs corresponding to multiple functions F1 to F3. The processor 61 reads and executes the programs stored in the memory 62. Thus, the multiple functions F1 to F3 are implemented.
[0148] Or, such as Figure 17 As shown, the labeling device 100 includes a processing circuit 63. The processing circuit 63 performs processing corresponding to multiple functions F1 to F3. Thus, the multiple functions F1 to F3 are realized.
[0149] Or, such as Figure 18 As shown, the labeling device 100 includes a processor 61, a memory 62, and a processing circuit 63. The memory 62 stores programs corresponding to a portion of the multiple functions F1 to F3. The processor 61 reads and executes the programs stored in the memory 62. This enables the implementation of that portion of the functions. Furthermore, the processing circuit 63 performs processing corresponding to the remaining functions of the multiple functions F1 to F3. This enables the implementation of the remaining functions.
[0150] Processor 61 consists of one or more processors. Each processor may be a CPU (Central Processing Unit), GPU (Graphics Processing Unit), microprocessor, microcontroller, or DSP (Digital Signal Processor).
[0151] The memory 62 is composed of one or more non-volatile memories. Alternatively, the memory 62 is composed of one or more non-volatile memories and one or more volatile memories. That is, the memory 62 is composed of one or more memories. Each memory may be, for example, a semiconductor memory, a magnetic disk, an optical disk, an optical disc, a magnetic tape, or a magnetic drum. More specifically, each volatile memory may be, for example, RAM (Random Access Memory). In addition, each non-volatile memory may be, for example, ROM (Read Only Memory), flash memory, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), a solid-state drive, a hard disk, a floppy disk, a high-density disk, a DVD (Digital Versatile Disc), a Blu-ray disc, or a mini-disc.
[0152] The processing circuit 63 is composed of one or more digital circuits. Alternatively, the processing circuit 63 is composed of one or more digital circuits and one or more analog circuits. That is, the processing circuit 63 is composed of one or more processing circuits. Each processing circuit may use, for example, an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), an FPGA (Field Programmable Gate Array), a SoC (System on a Chip), or a system LSI (Large Scale Integration).
[0153] Here, when processor 61 is composed of multiple processors, the correspondence between the multiple functions F1 to F3 and the multiple processors is arbitrary. That is, the multiple processors can also read and execute programs corresponding to one or more functions among the multiple functions F1 to F3. Processor 61 may also include dedicated processors corresponding to each function F1 to F3.
[0154] Furthermore, when memory 62 is composed of multiple memories, the correspondence between the multiple functions F1 to F3 and the multiple memories is arbitrary. That is, the multiple memories can also store programs corresponding to one or more of the corresponding functions F1 to F3. Memory 62 can also include dedicated memories corresponding to each function F1 to F3.
[0155] Furthermore, when the processing circuit 63 is composed of multiple processing circuits, the correspondence between the multiple functions F1 to F3 and the multiple processing circuits is arbitrary. That is, the multiple processing circuits may each execute the processing corresponding to one or more of the multiple functions F1 to F3. The processing circuit 63 may also include dedicated processing circuits corresponding to each function F1 to F3.
[0156] Next, refer to Figures 19-21 The hardware structure of the main parts of the learning device 200 will be described.
[0157] like Figure 19 As shown, the learning device 200 includes a processor 71 and a memory 72. The memory 72 stores programs corresponding to multiple functions F11 and F12. The processor 71 reads and executes the programs stored in the memory 72. Thus, the multiple functions F11 and F12 can be implemented.
[0158] Or, such as Figure 20 As shown, the learning device 200 has a processing circuit 73. The processing circuit 73 performs processing corresponding to multiple functions F11 and F12. Thus, multiple functions F11 and F12 can be implemented.
[0159] Or, such as Figure 21 As shown, the learning device 200 includes a processor 71, a memory 72, and a processing circuit 73. The memory 72 stores programs corresponding to a portion of the multiple functions F11 and F12. The processor 71 reads and executes the programs stored in the memory 72, thereby implementing that portion of the function. Furthermore, the processing circuit 73 performs processing corresponding to the remaining functions of the multiple functions F11 and F12, thereby implementing the remaining functions.
[0160] The specific example of processor 71 is the same as that of processor 61. The specific example of memory 72 is the same as that of memory 62. The specific example of processing circuit 73 is the same as that of processing circuit 63. Therefore, detailed description is omitted.
[0161] Here, when processor 71 is composed of multiple processors, the correspondence between the multiple functions F11, F12 and the multiple processors is arbitrary. That is, the multiple processors can also read and execute programs corresponding to one or more of the functions F11, F12. Processor 71 can also include dedicated processors corresponding to each function F11, F12.
[0162] Furthermore, when memory 72 is composed of multiple memories, the correspondence between the multiple functions F11, F12 and the multiple memories is arbitrary. That is, the multiple memories can also store programs corresponding to one or more of the corresponding functions F11, F12. Memory 72 can also include dedicated memories corresponding to each function F11, F12.
[0163] Furthermore, when the processing circuit 73 is composed of multiple processing circuits, the correspondence between the multiple functions F11, F12 and the multiple processing circuits is arbitrary. That is, the multiple processing circuits may each execute the processing corresponding to one or more of the corresponding functions F11, F12. The processing circuit 73 may also include dedicated processing circuits corresponding to each function F11, F12.
[0164] Next, refer to Figure 22 The flowchart illustrates the operation of the labeling device 100.
[0165] First, the image signal acquisition unit 21 performs image signal acquisition processing (step ST1). Next, the image recognition unit 22 performs image recognition processing (step ST2). Next, the learning dataset generation unit 23 performs learning dataset generation processing (step ST3).
[0166] Next, refer to Figure 23 The flowchart illustrates the operation of the learning device 200.
[0167] First, the learning database update unit 31 performs learning database update processing (step ST11). Next, the learning unit 32 performs learning processing (step ST12).
[0168] Next, refer to Figure 24 A variation of annotation system 1 will be explained.
[0169] like Figure 24As shown, the learning device 200 may also include the annotation device 100. That is, the learning device 200 may also have an image signal acquisition unit 21, an image recognition unit 22, a learning dataset generation unit 23, a learning database update unit 31, and a learning unit 32.
[0170] Next, refer to Figure 25 Other variations of annotation system 1 will be explained.
[0171] like Figure 25 As shown, the annotation device 100 may also include a learning device 200. That is, the annotation device 100 may also have an image signal acquisition unit 21, an image recognition unit 22, a learning dataset generation unit 23, a learning database update unit 31, and a learning unit 32.
[0172] Next, other variations of annotation system 1 will be described.
[0173] The annotation device 100 can also be integrated with the camera 2. Furthermore, the learning device 200 can also be integrated with the camera 2. Thus, an AI (Artificial Intelligence) camera can be realized.
[0174] The annotation device 100 can also be configured as a server that can communicate freely with the camera 2. Similarly, the learning device 200 can also be configured as a server that can communicate freely with the camera 2. This server can also be an edge server. Thus, an edge AI camera can be realized.
[0175] As described above, the annotation apparatus 100 of Embodiment 1 includes: an image signal acquisition unit 21 that acquires an image signal representing an image captured by the camera 2; an image recognition unit 22, which is a machine learning-based, learned image recognition unit 22 that performs image recognition on the captured image; and a learning dataset generation unit 23 that generates a learning dataset containing image data corresponding to each object and annotation data corresponding to each object by performing annotation on each object contained in the captured image based on the image recognition results. Therefore, when generating the learning dataset using the image captured by the camera 2, the annotation work can be automated. As a result, the workload for the annotation manager can be reduced.
[0176] Furthermore, the image recognition unit 22 uses a known large-scale database to complete the learning process. Therefore, it can achieve high-precision object recognition for objects that have already been learned, and also achieve a certain level of accuracy for objects that have not yet been learned.
[0177] Furthermore, the image recognition unit 22 includes: a first feature extraction unit 41_1 that generates a first feature map corresponding to the captured image; a second feature extraction unit 41_2 that generates a second feature map corresponding to the captured image; and an object recognition unit 42 that performs object recognition using the first and second feature maps. The first feature map corresponds to a foreground mask or to an intermediate-order feature, which corresponds to objectivity, and the second feature map corresponds to a high-order feature. By using the first feature map in addition to the second feature map, the accuracy of object recognition can be improved. In particular, the accuracy of object recognition for objects that have not been learned can be improved.
[0178] Furthermore, the image recognition unit 22 includes an object tracking unit 43, which performs tracking of each object by using the results of object recognition in a time sequence. As a result, each object can be identified with higher accuracy.
[0179] Furthermore, the learning device 200 of Embodiment 1 is a learning device 200 for the annotation device 100, comprising: a learning database update unit 31, which updates the learning database by appending the learning dataset generated by the learning dataset generation unit 23 to the learning database; and a learning unit 32, which uses the learning database to perform relearning or append learning of the image recognition unit 22. Thus, relearning based on transfer learning, fine-tuning, few-shot learning, meta-learning, or distillation can be implemented for the image recognition unit 22. As a result, image recognition accuracy can be gradually improved, and annotation accuracy can be gradually improved. Furthermore, when automating the annotation job, a human equivalent to an Oracle in Active Learning is not required.
[0180] Furthermore, the learning database update unit 31 assigns a priority P to the learning dataset generated by the learning dataset generation unit 23 based on the deviation of the distribution D among the multiple learning datasets contained in the learning database. By using this priority P, efficient relearning can be achieved using a learning database that is smaller than a known large-scale database.
[0181] Furthermore, the learning unit 32 sets the learning rate η for relearning or supplementary learning based on the priority P. This enables efficient relearning, etc.
[0182] Furthermore, the learning unit 32 performs data expansion in the learning database according to priority P. This enables efficient relearning, etc.
[0183] Implementation Method 2
[0184] Figure 26 This is a block diagram showing the main parts of the annotation system in Embodiment 2. (Refer to...) Figure 26The annotation system of Implementation Method 2 will be described. Furthermore, in Figure 26 In the middle, to and Figure 1 Blocks that are identical to those shown are labeled with the same marker and their descriptions are omitted.
[0185] like Figure 26 As shown, the annotation system 1a includes a camera 2, a storage device 3, a storage device 4, an output device 5, an input device 6, an annotation device 100a, and a learning device 200. The annotation device 100a has an image signal acquisition unit 21, an image recognition unit 22, a learning dataset generation unit 23a, and a user interface control unit (hereinafter referred to as "UI control unit") 24.
[0186] Output device 5 may be, for example, a display or a speaker. Input device 6 may be a device corresponding to output device 5. For example, if output device 5 is a display, input device 6 may be a touch panel and a stylus. Alternatively, if output device 5 is a speaker, input device 6 may be a microphone.
[0187] The UI control unit 24 uses the output device 5 to control the output of the image recognition result of the image recognition unit 22. Furthermore, the UI control unit 24 processes inputs that accept operations using the input device 6, and processes inputs that accept operations to correct the image recognition result (hereinafter, sometimes referred to as "correction operations").
[0188] Specifically, for example, the UI control unit 24 uses a display to control the display of a screen (hereinafter sometimes referred to as a "correction screen") that includes an image representing the image recognition result of the image recognition unit 22. Furthermore, the UI control unit 24 processes input that accepts correction operations using a touch panel and a stylus. That is, the UI control unit 24 processes input that accepts correction operations based on handwriting input to the correction screen.
[0189] Alternatively, for example, the UI control unit 24 uses a speaker to perform voice control that outputs a voice message representing the image recognition result of the image recognition unit 22. Furthermore, the UI control unit 24 performs processing for accepting input for correction operations using a microphone. That is, the UI control unit 24 performs processing for accepting input for correction operations based on voice input. In this case, various known techniques related to voice recognition can be used in the processing for accepting input for correction operations.
[0190] Here, the UI related to the input of the correction operation can also use an interactive UI. As a result, the person in charge of annotation can easily correct the image recognition results of the image recognition unit 22.
[0191] The learning dataset generation unit 23a generates the same learning dataset as the learning dataset generated by the learning dataset generation unit 23. That is, based on the image recognition results from the image recognition unit 22, the learning dataset generation unit 23a generates a first learning dataset containing image data, annotation data, mask data, and reliability data. The learning dataset generation unit 23a generates a third learning dataset by appending priority data to the generated first learning dataset. The learning dataset generation unit 23a then causes the learning dataset storage unit 11 to store the generated third learning dataset.
[0192] However, if the image recognition result of the image recognition unit 22 is corrected by the correction operation, the learning dataset generation unit 23a generates the first learning data based on the correction result.
[0193] Hereinafter, the functions of the learning dataset generation unit 23a will sometimes be referred to as the "learning dataset generation function". Furthermore, the symbol "F3a" will sometimes be used in this learning dataset generation function. Additionally, the processing performed by the learning dataset generation unit 23a will sometimes be referred to as the "learning dataset generation processing".
[0194] Hereinafter, the functions of the UI control unit 24 will sometimes be collectively referred to as "UI control functions". Furthermore, the symbol "F4" will sometimes be used for these UI control functions. Additionally, the control and processing performed by the UI control unit 24 will sometimes be collectively referred to as "output control and operation input processing".
[0195] The hardware structure of the main parts of the labeling device 100a is the same as that described in Embodiment 1. Figures 16-18 The structure is the same as described above. Therefore, detailed descriptions are omitted. That is, the labeling device 100a has multiple functions F1, F2, F3a, and F4. The multiple functions F1, F2, F3a, and F4 can be implemented by the processor 61 and the memory 62 respectively, or they can also be implemented by the processing circuit 63.
[0196] Next, refer to Figure 27 The flowchart illustrates the operation of the labeling device 100a. Furthermore, in Figure 27 In the middle, to and Figure 22 Steps that are identical in the steps shown are marked with the same symbols and their descriptions are omitted.
[0197] First, the processing in step ST1 is executed. Next, the processing in step ST2 is executed. Then, the UI control unit 24 performs output control and operation input processing (step ST4). Next, the learning dataset generation unit 23a performs learning dataset generation processing (step ST3a).
[0198] Next, refer to Figure 28A variation of annotation system 1a will be explained.
[0199] like Figure 28 As shown, the learning device 200 may also include a labeling device 100a. That is, the learning device 200 may also have an image signal acquisition unit 21, an image recognition unit 22, a learning dataset generation unit 23a, a UI control unit 24, a learning database update unit 31, and a learning unit 32.
[0200] Next, refer to Figure 29 Other variations of annotation system 1a will be explained.
[0201] like Figure 29 As shown, the annotation device 100a may also include the learning device 200. That is, the annotation device 100a may also have an image signal acquisition unit 21, an image recognition unit 22, a learning dataset generation unit 23a, a UI control unit 24, a learning database update unit 31, and a learning unit 32.
[0202] Next, other variations of annotation system 1a will be explained.
[0203] The annotation device 100a can also be integrated with the camera 2. Furthermore, the learning device 200 can also be integrated with the camera 2. Thus, an AI camera can be realized.
[0204] The annotation device 100a can also be configured as a server that can communicate freely with the camera 2. Similarly, the learning device 200 can also be configured as a server that can communicate freely with the camera 2. This server can, for example, be an edge server. Thus, an edge AI camera can be realized.
[0205] As described above, the annotation apparatus 100a of Embodiment 2 includes a UI control unit 24, which controls the output of the image recognition result and processes inputs for operations that correct the image recognition result. The learning dataset generation unit 23a generates a learning dataset based on the operation-based correction result. Therefore, when generating the learning dataset using images captured by the camera 2, the annotation work can be semi-automated. In other words, it can support the annotation work of the annotation manager. As a result, the workload for the annotation manager can be reduced.
[0206] Furthermore, the UI control unit 24 controls the display of the screen and processes input based on handwriting input to the screen, where the screen contains an image representing the result of image recognition. Using this UI, the image recognition result can be easily corrected.
[0207] Furthermore, the UI control unit 24 performs voice control to output the image recognition result and processes input that accepts voice-based input. By using this UI, the image recognition result can be easily corrected.
[0208] Furthermore, within the scope of its disclosure, this application enables free combination of various embodiments, modification of any structural elements in various embodiments, or omission of any structural elements in various embodiments.
[0209] Industrial availability
[0210] The annotation and learning devices of this application can be used, for example, in surveillance systems, security systems, or electronic mirrors.
[0211] Label Explanation
[0212] 1, 1a: Annotation system; 2: Camera; 3: Storage device; 4: Storage device; 5: Output device; 6: Input device; 11: Learning dataset storage unit; 12: Learning database storage unit; 21: Image signal acquisition unit; 22: Image recognition unit; 23, 23a: Learning dataset generation unit; 24: UI control unit; 31: Learning database update unit; 32: Learning unit; 41: Feature extraction unit; 41_1: First feature extraction unit; 41_2: Second feature extraction unit; 42: Object recognition unit; 43: Object tracking unit; 51: Learning dataset acquisition unit; 52: Learning dataset acquisition unit; 53: Priority assignment unit; 54: Learning dataset appending unit; 61: Processor; 62: Memory; 63: Processing circuit; 71: Processor; 72: Memory; 73: Processing circuit; 100, 100a: Annotation device; 200: Learning device.
Claims
1. A labeling device, wherein, The labeling device has the following features: The image signal acquisition unit acquires an image signal representing an image captured by the camera; An image recognition unit, which is a machine learning-based learned image recognition unit, performs image recognition for the captured image; as well as The learning dataset generation unit generates a learning dataset by performing annotation on each object contained in the captured image based on the image recognition result, and generating image data corresponding to each object and annotation data corresponding to each object. The image recognition unit includes: a first feature extraction unit that generates a first feature map corresponding to the captured image; a second feature extraction unit that generates a second feature map corresponding to the captured image; and an object recognition unit that performs object recognition using the first feature map and the second feature map. The first feature map corresponds to either the foreground mask or a mid-level feature, which corresponds to objectivity. The second feature map corresponds to the higher-order features.
2. The labeling device according to claim 1, characterized in that, The image recognition unit uses an existing large-scale database to perform the learning process.
3. The labeling device according to claim 1, characterized in that, The first feature extraction unit generates the first feature map using at least one of image gradient detection, saliency estimation, background subtraction, objectivity estimation, attention, and region segmentation.
4. The labeling device according to claim 1, characterized in that, The object recognition unit uses each of the first feature quantities in the first feature map to perform weighting on the corresponding second feature quantities in the second feature map.
5. The labeling device according to claim 4, characterized in that, The object recognition unit sets the importance in the weighting based on the similarity between each first feature and the corresponding second feature.
6. The labeling device according to claim 5, characterized in that, The similarity is based on the value of at least one of EMD, cosine similarity, KLD, L2 norm, L1 norm, and Manhattan distance.
7. The labeling device according to claim 4, characterized in that, When attention is used in the first feature map, the object recognition unit selects a representative value in the first feature map and sets the importance in the weighting based on the representative value.
8. The labeling device according to claim 1, characterized in that, The object recognition includes at least one of object detection and region segmentation. In the object detection, the positions of each object are estimated through regression, and the attributes of each object are estimated through classification. In the region segmentation, the captured image is divided into regions corresponding to each attribute.
9. The labeling device according to claim 1, characterized in that, The image recognition unit has an object tracking unit that performs tracking of each object by using the results of the object recognition in a time sequence.
10. The labeling device according to claim 1, characterized in that, The first feature extraction unit can learn freely through supervised or unsupervised learning.
11. The labeling device according to claim 1, characterized in that, The second feature extraction unit learns freely through supervised learning.
12. The labeling device according to claim 1, characterized in that, The second feature extraction unit learns freely through deep learning.
13. The labeling device according to claim 1, characterized in that, The second feature extraction unit uses a convolutional neural network.
14. The labeling device according to claim 1, characterized in that, The object recognition unit learns freely through supervised learning.
15. The labeling device according to claim 1, characterized in that, The annotation device includes a UI control unit that controls the output of the image recognition result and processes inputs for accepting and correcting the image recognition result. The learning dataset generation unit generates the learning dataset based on the correction result of the operation.
16. The labeling device according to claim 15, characterized in that, The UI control unit performs control over the display of the screen and performs processing for accepting input based on the operation of handwriting input to the screen, wherein the screen includes an image representing the result of the image recognition.
17. The labeling device according to claim 15, characterized in that, The UI control unit performs control by outputting voice representing the result of the image recognition, and performs processing of input that accepts the operation based on voice input.
18. The labeling device according to claim 1, characterized in that, The camera in question is a surveillance camera.
19. The labeling device according to claim 18, characterized in that, The camera is a surveillance camera, a security camera, or a camera used for electronic mirrors.
20. A learning device, wherein the learning device is used with the annotation device according to claim 1, characterized in that, This learning device has the following features: The learning database update unit updates the learning database by appending the learning dataset generated by the learning dataset generation unit to the learning database; and The learning unit uses the learning database to perform relearning or additional learning by the image recognition unit.
21. The learning device according to claim 20, characterized in that, The learning database update unit assigns priority to the learning datasets generated by the learning dataset generation unit based on the distribution deviation among the multiple learning datasets contained in the learning database.
22. The learning device according to claim 21, characterized in that, The priority is set to a value corresponding to the learning value of the learning dataset generated by the learning dataset generation unit.
23. The learning device according to claim 21, characterized in that, The distribution is based on at least one of reliability, size, high-dimensional image features, low-dimensional image features, objectivity, saliency estimation, and attention.
24. The learning device according to claim 21, characterized in that, The learning unit sets the learning rate for the relearning or the additional learning based on the priority.
25. The learning device according to claim 21, characterized in that, The learning unit performs data expansion in the learning database according to the priority.
Citation Information
Patent Citations
Information processor and information-processing method
WO2018079020A1
An image automatic labeling method and system based on deep learning
CN108985293A