Object detection device, object detection method, and program

The object detection device streamlines the creation of teacher data by using a graphical user interface to visualize detection results and facilitate easy image selection and classification, addressing the inefficiencies of existing methods and improving training accuracy.

JP2025081496AActive Publication Date: 2025-05-27NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025024657
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2015-03-19
Filing Date
2025-02-19
Publication Date
2025-05-27
Estimated Expiration
2036-03-15

AI Technical Summary

Technical Problem

Existing methods for creating teacher data for object detection devices are time-consuming and labor-intensive, requiring manual selection and classification of images based on detected objects.

Method used

An object detection device that includes detection, receiving, display, and generating means to efficiently select and process images for teacher data creation based on the number of detected objects, using a graphical user interface to visualize detection results and allow for easy selection and classification.

Benefits of technology

Enables efficient and easy selection of images for creating high-quality teacher data, reducing the time and effort required for manual processing and improving the accuracy of object detection training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025081496000001_ABST
    Figure 2025081496000001_ABST
Patent Text Reader

Abstract

To provide an object detection device and an object detection method for efficiently and simply selecting an image for creating training data on the basis of the number of detected objects.SOLUTION: An object detection device is provided with: a detection unit for detecting an object from each of a plurality of input images using a dictionary; a receiving unit for displaying, on a display device, a graph indicating a relationship between the input images and the number of subregions in which the objects are detected, and displaying, on the display device, in order to create training data, one of the plurality of input images in accordance with a position on the graph received through operation of an input device; a generation unit for generating the training data from the input image; and a learning unit for learning the dictionary using the training data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an object detection device that detects an object from an image using a dictionary, and the like.

Background Art

[0002] Devices that analyze a captured image obtained by a imaging device and detect an object have been proposed or put into practical use. For example, Patent Document 1 discloses a technique for analyzing an image obtained by imaging a substrate in a manufacturing process of a substrate such as a printed wiring board and detecting an object such as a scratch on the substrate. Patent Document 2 discloses a technique for analyzing a monitor image on a road obtained by imaging with a surveillance camera and detecting an object such as a vehicle.

[0003] In order to perform the above-described detection, it is necessary to train an object detection device to analyze an image or video. And for this training, teacher data (also called training data) is required. Teacher data is a pair of input and output cases. There are two types of teacher data: positive cases and negative cases. Both positive cases and negative cases are required to correctly train an object detection device. However, creating appropriate teacher data requires a lot of time and effort. For this reason, several techniques for assisting in the creation of teacher data have been proposed.

[0004] For example, the above Patent Document 1 (see FIG. 9) discloses a method for creating teacher data necessary for detecting an object such as a scratch on a substrate. In this method, a region having a luminance value different from that of a non-defective product is extracted from an image of a printed wiring board and displayed on a display, and the selection of the region and the input of its class (also called a category) are received from a user using a keyboard and a mouse. Specifically, the user selects a specific one of a plurality of existing regions by clicking the mouse, and then selects a desired class from a pull-down menu displayed at the time of this selection by clicking the mouse.

[0005] In addition, Patent Document 2 described above discloses a method for creating teaching data necessary for detecting an object such as a vehicle traveling on a road. In this method, a series of operations are automatically performed, in which the area of an object is cut out from an arbitrary captured image using a dictionary, a predetermined feature amount is extracted from the area, and the dictionary is learned based on the extracted feature amount.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0007] There may be many images in which the area of an object is detected from an image obtained by imaging with a camera using a dictionary. In this case, there is a user's desire to select the order of images to be confirmed in order to create teaching data based on the number of detected objects. Currently, in order to confirm the number of objects detected in an image, each image has to be displayed on a display device for confirmation. However, this method requires a lot of time and effort to select the images to be processed based on the number of detected objects.

[0008] In view of the above problems, an object of the present invention is to provide an object detection device or the like that can efficiently and easily select an image for creating teaching data based on the number of detected objects.

Means for Solving the Problems

[0009] The first feature of the present invention is detection means for detecting an object from each of a plurality of input images, Receiving means for receiving an input of a position in a graph showing the relationship between the plurality of input images and the object; Display means for displaying, on a display device, a first image that is one of the plurality of input images according to the position; Selecting means for selecting a class; Receiving means for receiving, in the first image, a designation of a partial region associated with the class; Generating means for generating data to be used in machine learning based on the partial region and the class; An object detection device comprising the above.

[0010] A second feature of the present invention is that a computer detects an object from each of a plurality of input images, receives an input of a position in a graph showing the relationship between the plurality of input images and the object, displays, on a display device, a first image that is one of the plurality of input images according to the position, selects a class, receives, in the first image, a designation of a partial region associated with the class, and generates data to be used in machine learning based on the partial region and the class. This is an object detection method.

[0011] A third feature of the present invention is that a process of detecting an object from each of a plurality of input images, a process of receiving an input of a position in a graph showing the relationship between the plurality of input images and the object, a process of displaying, on a display device, a first image that is one of the plurality of input images according to the position, a process of selecting a class, a process of receiving, in the first image, a designation of a partial region associated with the class, and a process of generating data to be used in machine learning based on the partial region and the class, are executed by a computer. This is a program.

Advantages of the Invention

[0012] The present invention can efficiently and easily select an image for creating teaching data based on the number of detected objects.

Brief Description of Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10A

Figure 10B

Figure 11

Figure 12A

Figure 12B

Figure 13A

Figure 13B

Figure 14

Figure 15

Figure 16A

Figure 16B

Figure 17A

Figure 17B

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23A

Figure 23B

Figure 24A

Figure 24B

Figure 25

Figure 26

Figure 27

Embodiments for Carrying Out the Invention

[0014] Next, embodiments of the present invention will be described in detail with reference to the drawings. In the following description of the drawings, the same or similar parts are denoted by the same or similar reference numerals. However, the drawings schematically represent the configurations in the embodiments of the present invention. Furthermore, the embodiments of the present invention described below are examples, and can be appropriately modified within the scope of the same essence. [First Embodiment] Referring to FIG. 1, an object detection device 100 according to a first embodiment of the present invention analyzes a monitor video on a road obtained by imaging with a surveillance camera 110 and detects an object. In this embodiment, the objects to be detected are two types: motorcycles and automobiles. That is, the object detection device 100 is a type of multi-class classifier and detects two specific types of objects from an image.

[0015] The object detection device 100 mainly includes a detection unit 101, a dictionary 102, a detection result DB (detection result database) 103, a reception unit 104, a generation unit 105, a teacher data memory 106, and a learning unit 107. In addition, a display device 120 and an input device 130 are connected to the object detection device 100. The object detection device 100 may be configured by an information processing device 230 and a storage medium that stores a program 240, as shown in FIG. 26, for example. The information processing device 230 includes an arithmetic processing unit 210 such as one or more microprocessors and a storage unit 220 such as a semiconductor memory or a hard disk. The storage unit 220 stores the dictionary 102, the detection result DB 103, the teacher data memory 106, and the like. The program 240 is read from an externally computer-readable recording medium into the memory when the object detection device 100 is started up or the like, and by controlling the operation of the arithmetic processing unit 210, functional means such as the detection unit 101, the reception unit 104, the generation unit 105, and the learning unit 107 are realized on the arithmetic processing unit 210.

[0016] The display device 120 is composed of a screen display device such as an LCD (Liquid Crystal Display) or a PDP (Plasma Display Panel), and displays various information such as detection results on the screen according to an instruction from the object detection device 100.

[0017] The input device 130 is an operation input device such as a keyboard or a mouse. The input device 130 detects the operator's operation as an input and outputs it to the object detection device 100. In the present embodiment, a mouse is used as the input device 130. Also, the mouse to be used is assumed to be capable of performing two types of mouse gestures: left click and right click.

[0018] The detection unit 101 inputs the images obtained by imaging with the surveillance camera 110 one by one in chronological order, and detects an object from each of the input images using the dictionary 102. The detection unit 101 stores the detection result in the detection result DB 103. One image refers to one frame of the image captured by the surveillance camera 110. A series of frames output from the surveillance camera 110 are assigned consecutive frame numbers for identification.

[0019] FIG. 2 is an explanatory diagram of the concept of detecting an object from an image. The detection unit 101 sets a search window 112 for object detection for the image 111, and extracts the feature amount of the image within the search window 112. Further, the detection unit 101 calculates a likelihood indicating the similarity between the extracted feature amount and the dictionary 102, and detects an object by determining whether the image within the search window 112 is an object or not based on the calculated likelihood. When the detection unit 101 detects an object, it associates the detection result having the information of the partial region where the object is detected and the class described later with the identification information of the image 111 and stores it in the detection result DB 103. As the information of the partial region of the object, the position information of the search window 112 determined to be an object on the image 111 is used. Here, as the position information, when the search window 112 is rectangular, for example, the coordinate values of the upper left and lower right vertices are used. Also, the classes related to the object are a total of three classes: class 1 representing the class of two-wheeled vehicles, class 2 representing the class of four-wheeled vehicles, and class 0 representing neither of them. Further, the identification information of the image 111 is, for example, the frame number of the image.

[0020] The detection unit 101 changes the position and size of the search window 112, and repeats the same operation as described above to search for objects with different positions and sizes existing in the image 111 without omission. In the example shown in FIG. 2, three objects are detected from the image 111. The detection unit 101 detects the first object with the search window at the position and size indicated by reference numeral 113, the second object with the search window at the position and size indicated by reference numeral 114, and the third object with the search window at the position and size indicated by reference numeral 115. In FIG. 2, the horizontal direction of the image is the X-axis and the vertical direction is the Y-axis. After finishing the detection process for one image 111, the detection unit 101 repeats the same process for the next image.

[0021] FIG. 3 shows an example of the detection result DB103. The detection result DB103 in this example records the detection results corresponding to the frame numbers. The detection result is composed of the information of the partial region of the object and the class. For example, the detection result DB103 records a pair of the information of the partial regions of three objects and the class corresponding to the frame number 001. One of them is a pair of the partial region with the coordinates of the upper left vertex being (x1, y1) and the coordinates of the lower right vertex being (x2, y2) and class 2, another one is a pair of the partial region with the coordinates of the upper left vertex being (x3, y3) and the coordinates of the lower right vertex being (x4, y4) and class 1, and the remaining one is a pair of the partial region with the coordinates of the upper left vertex being (x5, y5) and the coordinates of the lower right vertex being (x6, y6) and class 1.

[0022] Also, the detection result DB103 shown in FIG. 3 has a column for recording the corrected class for each detection result. The corrected class corresponds to the class input by the reception unit 104. At the time when the detection unit 101 outputs the detection result to the detection result DB103, all the columns of the corrected class are NULL (indicated by - in FIG. 3).

[0023] The reception unit 104 visualizes the detection results of the images stored in the detection result DB103 and displays them on the display device 120, and receives input of corrections from the operator.

[0024] FIG. 4 shows an example of a reception screen 121 that the reception unit 104 displays on the display device 120. In this example, the reception screen 121 displays a graph 122 at the lower part of the screen, a slider 123 above it, and an image window 124 further above it.

[0025] The graph 122 is a graph showing the relationship between a plurality of images including the image 111 and the number of partial regions where an object is detected in each screen. When the horizontal direction of the reception screen 121 is the X-axis and the vertical direction is the Y-axis, the X-axis of the graph 122 indicates the frame number of the image, and the Y-axis indicates the number of partial regions where an object is detected, that is, the object detection number. That is, the reception unit 104 uses, as the X-axis of the graph 122, a column in which the input images corresponding to the detection results stored in the detection result DB 103 are arranged in ascending or descending order of the frame number. However, the X-axis may be other than a column in which the input images are arranged in ascending or descending order of the frame number. For example, the X-axis may be a column in which the input images are arranged by the number of partial regions of the object (that is, the object detection number). In this case, the reception unit 104 sorts the input images of the detection result 103 stored in the detection result DB 103 in ascending or descending order by the object detection number, and uses the column of the sorted input images as the X-axis of the graph 122. Also, in FIG. 4, the graph 122 is a bar graph, but the type of the graph 122 is not limited to the bar graph, and may be other types such as a line graph.

[0026] The slider 123 is a GUI (Graphical User Interface) for selecting a position on the X-axis of the graph 122, that is, for selecting an image. The slider 123 includes a slide bar 123a, and by operating the slide bar 123a left and right with the input device 130, a position on the X-axis of the graph 122 is selected.

[0027] The image window 124 displays one image 111 selected by the operation of the slider 123 among a plurality of images in frame units. On the image 111 displayed in the image window 124, a display for emphasizing the partial region of the detected object is added. In the example of FIG. 4, a rectangular frame 125 indicating the outer periphery of the partial region of the object is displayed. However, the display for emphasizing the partial region of the object is not limited to the rectangular frame 125, and any display form such as a display form that makes the luminance of the entire partial region higher or lower than other regions or a display form that displays hatching on the entire partial region may be used.

[0028] In addition, the display for emphasizing the partial region of the detected object is displayed in different display forms according to the class of the detected object. In FIG. 4, the partial region of the object of class 1 is displayed by a dashed-line frame 125, and the partial region of the object of class 2 is displayed by a solid-line frame 125. In addition to making the line types of the frame 125 different, the display color of the frame 125 may be made different. Also, a numerical value indicating the class may be displayed near the partial image.

[0029] Generally, on the slider of the GUI, a scale and a label serving as a guide for the operator's operation are displayed. The reception screen 121 shown in FIG. 4 displays a graph 122 indicating the number of object detections for each image instead of the scale and the label. By displaying such a graph 122, the operator can easily select one image from a plurality of images based on the number of object detections.

[0030] The reception unit 104 receives the selection of a partial region and the input of a class for the selected partial region by operating the input device 130 on the image 111 displayed on the display device 120. The reception unit 104 receives the selection of a partial region based on the position of the image 111 where the operation of the input device 130 is performed, and receives the input of a class based on the type of the operation. In the present embodiment, the input device 130 is a mouse, and the types of operations are left click and right click. The left click means an operation of single-clicking the left button of the mouse. The right click means an operation of single-clicking the right button of the mouse. The reception unit 104 receives the selection of a partial region having the position of the image 111 where the left click or right click is performed within the region. Further, the reception unit 104 receives the input of class 0 for the selected partial region when the left click is performed, and receives the input of a class obtained by incrementing the current class when the right click is performed. Here, the class obtained by incrementing the current class means class 2 when the current class is class 1, class 0 when the current class is class 2, and class 1 when the current class is class 0, respectively. The reception unit 104 records the received class for the selected partial region in the column of the corrected class in the detection result DB 103.

[0031] The generation unit 105 generates teacher data from the image of the partial region selected by the reception unit 104 and the input class, and stores the teacher data in the teacher data memory 106. The generation unit 105 generates teacher data for each partial region in which any one of classes 0, 1, and 2 is recorded in the column of the corrected class in the detection result DB 103 shown in FIG. 3. The teacher data is an input / output pair of the image of the partial region or the feature amount extracted from the image of the partial region and the class recorded in the column of the corrected class.

[0032] The learning unit 107 learns the dictionary 102 using the teacher data stored in the teacher data memory 106. Since the method of learning the dictionary 102 using the teacher data is widely known, its description is omitted. The dictionary 102 is upgraded by learning with the teacher data. The detection unit 101 uses this upgraded dictionary 102 to detect an object from each image input again from the surveillance camera 110.

[0033] Figure 5 is a flowchart showing the operation of this embodiment. Hereinafter, the operation of this embodiment will be described with reference to Figure 5.

[0034] The detection unit 101 of the object detection device 100 inputs one by one in chronological order the images obtained by imaging with the surveillance camera 110 (step S101). Next, the detection unit 101 detects an object from the input image using the dictionary 102 by the method described with reference to Figure 2, and saves the detection result in the detection result DB 103 (step S102). If there is a next image obtained by imaging with the surveillance camera 110 (step S103), the detection unit 101 returns to step S100 and repeats the same processing as described above. As a result, the detection results of a plurality of images as shown in Figure 3 are accumulated in the detection result DB 103. When the processing of the detection unit 101 for a series of images captured by the surveillance camera 110 is completed, the processing by the reception unit 104 is started.

[0035] The reception unit 104 displays a reception screen 121 as shown in Figure 4 on the display device 120 and receives an input operation of the input device (step S104). The details of this step S104 will be described later.

[0036] When the reception by the reception unit 104 is completed, the generation unit 105 generates teacher data from the image of the partial area selected by the reception unit 104 and the input class (step S105). Next, the learning unit 107 learns the dictionary 102 using the teacher data stored in the teacher data memory 106 (step S106).

[0037] FIG. 6 is a flowchart showing the details of the operation of the reception unit 104 performed in step S104. Hereinafter, the operation of the reception unit 104 will be described in detail with reference to FIG. 6.

[0038] The reception unit 104 calculates the number of object detections per image from the detection result DB 103 (step S111). Referring to FIG. 3, since three partial regions are detected in the image with frame number 001, the reception unit 104 sets the number of object detections for the image with frame number 001 to 3. Similarly, the number of object detections for images with other frame numbers is calculated.

[0039] Next, the reception unit 104 displays the initial reception screen 121 on the display device 120 (step S112). As shown in FIG. 4, the initial reception screen 121 displays a graph 122, a slider 123, and an image window 124. The graph 122 displays the number of object detections for each frame number calculated in step S111. The slider 123 has a slide bar 123a placed at a predetermined position. The image window 124 displays the image of the object number on the X-axis of the graph 122 indicated by the slide bar 123a. Also, the partial regions of the objects detected in the image displayed in the image window 124 are emphasized by frames 125 of line types corresponding to the classes of the partial regions.

[0040] Next, the reception unit 104 determines whether the slide bar 123a of the slider 123 has been operated (step S113), whether a click has been performed on the image in the image window (step S114), and whether the reception end condition has been satisfied (step S115). The reception end condition may be, for example, that a reception end command has been input from the operator, or that no input operation has been performed for a certain period of time or more.

[0041] When the reception unit 104 detects a left click or a right click on the image in the image window 124 (YES in step S114), it determines whether a partial region including the click position exists in the image (step S116). This determination is made by examining whether the coordinate value of the clicked position is within any partial region of the image displayed in the image window 124. If the partial region including the click position does not exist (NO in step S116), the reception unit 104 ignores the click. On the other hand, if the partial region including the click position exists (YES in step S116), the reception unit 104 determines that the partial region including the click position has been selected, and determines whether the type of click is a left click or a right click in order to determine the input of the class (step S117).

[0042] If it is a left click, the reception unit 104 accepts the input of class 0 for the selected partial region (step S118). Next, the reception unit 104 updates the correction class corresponding to the selected partial region in the detection result in the detection result DB103 to the accepted class 0 (step S119). Next, the reception unit 104 hides the frame 125 displayed in the selected partial region of the image displayed in the image window 124 (step S120).

[0043] Also, if a right click is detected, the reception unit 104 receives an input of the class obtained by incrementing the current class of the selected partial region (step S121). The current class of the selected partial region is the class of the detection result when the correction class corresponding to the selected partial region in the detection result in the detection result DB 103 is NULL, and is the class described in the correction class when the correction class is not NULL. Next, the reception unit 104 updates the correction class corresponding to the selected partial region in the detection result in the detection result DB 103 to the class after the increment (step S122). Next, the reception unit 104 updates the display of the frame 125 displayed in the selected partial region of the image displayed in the image window 124 (step S123). Specifically, if the class after the increment is class 2, the reception unit 104 displays the frame 125 with a solid line. Also, if the class after the increment is class 0, the reception unit 104 hides the frame 125. Also, if the class after the increment is class 1, the reception unit 104 displays the frame 125 with a dashed line.

[0044] On the other hand, when the reception unit 104 detects that the slider 123a has been operated (step S113), the reception unit 104 updates the slider 123a on the reception screen 121 and the image in the image window 124 (step S124). Specifically, the reception unit 104 moves the display position of the slider 123a according to the input operation. Also, the reception unit 104 displays, in the image window 124, an image of the object number on the X-axis of the graph 122 indicated by the slider 123a after the movement.

[0045] Furthermore, the reception unit 104 determines whether a partial region on the image displayed in the image window 124 has the same position and the same pre-input class as the partial region of the image for which the class was input in the past. When it is determined that they have the same position and the same pre-input class, the reception unit 104 hides the frame 125 that emphasizes the partial region. For example, the positions of the third partial images in frame numbers 001, 002, and 003 in FIG. 3 are all rectangles at the same position with (x5, y5) and (x6, y6) as the upper left vertex and the lower right vertex, respectively, and their pre-input classes are also the same class 1. Therefore, for example, after the class of the third partial image in the image of frame number 001 is changed from class 1 to class 0, when the images of frame numbers 002 and 003 are displayed in the image window 124, the frame 125 is not displayed for the third partial images of the images of frame numbers 002 and 003. However, class 1 is still recorded in the classes of their detection results.

[0046] As described above, according to the present embodiment, an image for creating teacher data based on the number of detected objects can be efficiently and easily selected. More specifically, one image can be easily selected from a plurality of images based on the number of object detections. The reason is that the reception unit 104 selects an appropriate input image for displaying the detection result on the display device 120. The reception unit 104 displays a graph 122 showing the relationship between the input image and the number of partial regions where objects are detected on the display device 120, and selects the input image according to the position on the graph 122 received by operating the slider 123. Thus, the present embodiment can select an image to be processed based on the number of object detections. For this reason, according to the skill level and preference of the operator, an image with many detected objects can be preferentially processed. On the contrary, in the present embodiment, it is easy to preferentially process an image with a small number of detected objects or an image with an intermediate number of detected objects. Generally, preferentially processing an image with many detected objects shortens the total working time. However, in the case of an image with a lot of overlap between partial regions, it is not easy for an inexperienced operator to confirm.

[0047] Furthermore, according to the present embodiment, high-quality teacher data can be efficiently created. The reason is that the reception unit 104 displays the input image on the display device 120 with a display that emphasizes the partial region of the detected object, and accepts the selection of the partial region and the input of the class for the selected partial region by a click which is one operation of the input device 130.

[0048] Also generally, when a surveillance camera 110 with a fixed field of view captures images of two-wheeled vehicles and four-wheeled vehicles traveling on a road at regular intervals, immobile objects commonly appear at the same position in a plurality of consecutive frame images. For this reason, for example, when a puddle formed on a road is erroneously detected as a two-wheeled vehicle, the same misrecognition result appears at the same position in a plurality of consecutive frame images. It is sufficient to generate teacher data by correcting any one of the same misrecognition results. However, if the same misrecognition results that have already been corrected are highlighted over a large number of frame images, it becomes rather troublesome. In the present embodiment, the reception unit 104 makes the frame 125 that emphasizes the partial region having the same position and the same pre-input class as the partial region of the image for which the class has been input in the past invisible. Thereby, such troublesomeness can be eliminated.

[0049] [Modification Example of the First Embodiment] Next, various modification examples in which the configuration of the first embodiment is changed will be described.

[0050] In the first embodiment described above, the input of class 0 is accepted by a left click of a mouse gesture, and the input of class increment is accepted by a right click. However, as exemplified in the list of FIG. 7, various combinations are possible for the gesture type for accepting the input of class 0 and the gesture type for accepting the input of class increment.

[0051] For example, in No. 1-1 shown in FIG. 7, the reception unit 104 receives an input of class 0 by a left click of the mouse gesture and receives an input of increment of the class by a double click. A double click means an operation of continuously performing two single clicks on the left button of the mouse.

[0052] Also, in No. 1-2 shown in FIG. 7, the reception unit 104 receives an input of class 0 by a left click of the mouse gesture and receives an input of increment of the class by a drag and drop. A drag and drop means an operation of moving (dragging) the mouse while pressing the left button of the mouse and releasing (dropping) the left button at another location. When using drag and drop, it is determined whether a drag and drop is performed on the image window in step S114 of FIG. 6, and in step S116, for example, it is determined whether there is a partial region including the start point and the drop point of the drag.

[0053] Similar to the first embodiment, the above No. 1-1 and 1-2 use mouse gestures. In contrast, No. 1-3 to 1-6 shown in FIG. 7 use touch gestures. A touch gesture means a touch operation using a part of the human body such as a fingertip performed on a touch panel. There are many types of touch gestures, such as a tap (a light tapping operation with a finger or the like), a double tap (a double tapping operation with a finger or the like), a flick (a flicking or swiping operation with a finger or the like), a swipe (a tracing operation with a finger or the like), a pinch in (a pinching or narrowing operation with multiple fingers or the like), and a pinch out (a spreading operation with multiple fingers or the like). When using a touch gesture, a touch panel is used as the input device 130. The use of a touch panel is also useful, for example, when the present invention is implemented in a mobile device such as a smartphone or a tablet terminal.

[0054] In numbers 1-3 of FIG. 7, the reception unit 104 receives a Class 0 input by flicking and receives an input for incrementing the class by swiping. Also, in number 1-4 of FIG. 7, the reception unit 104 receives a Class 0 input by flicking leftward and receives an input for incrementing the class by flicking rightward. In number 1-5 of FIG. 7, the reception unit 104 receives a Class 0 input by pinching in and receives an input for incrementing the class by pinching out. In number 1-6 of FIG. 7, the reception unit 104 receives a Class 0 input by tapping and receives an input for incrementing the class by double-tapping. These are examples, and any other combination of gestures may be used.

[0055] In the above first embodiment and numbers 1-1 to 1-6 of FIG. 7, the reception unit 104 received a Class 0 input by a first type of gesture and received an input for incrementing the class by a second type of gesture. However, the reception unit 104 may be configured to receive an input for decrementing instead of incrementing. That is, in the above first embodiment and numbers 1-1 to 1-6 of FIG. 7, the reception unit 104 may receive a Class 0 input by a first type of gesture and receive an input for decrementing the class by a second type of gesture.

[0056] In the above first embodiment, the reception unit 104 received a Class 0 input by a left click of a mouse gesture and received an input for incrementing the class by a right click. However, the class and the operation may be associated one-to-one, and the reception unit 104 may be configured to receive inputs for Class 0, Class 1, and Class 2 by specific single operations respectively. As exemplified in the list of FIG. 8, various combinations are possible for the gesture types for receiving inputs for each class.

[0057] For example, in the numbers 1-11 shown in FIG. 8, the reception unit 104 receives an input of class 0 by a left click of a mouse gesture, receives an input of class 1 by a right click, and receives an input of class 2 by a double click. Also, in the number 1-12 shown in FIG. 8, the reception unit 104 receives inputs of class 0 and class 1 by the same mouse gestures as those in 1-11, and receives an input of class 2 by drag and drop. The above are examples of using mouse gestures. In contrast, the numbers 1-13 to 1-15 shown in FIG. 8 use touch gestures. For example, in the number 1-13 shown in FIG. 8, the reception unit 104 receives an input of class 0 by a flick, receives an input of class 1 by a tap, and receives an input of class 2 by a double tap. Also, in the number 1-14 shown in FIG. 8, the reception unit 104 receives an input of class 0 by a leftward flick, receives an input of class 1 by an upward flick, and receives an input of class 2 by a downward flick. Also, in the number 1-15 shown in FIG. 8, the reception unit 104 receives an input of class 0 by a tap, receives an input of class 1 by a pinch-in, and receives an input of class 2 by a pinch-out.

[0058] In the above first embodiment, the number of classes is three (class 0, class 1, class 2), but it can also be applied to a two-class object detection device with two classes (class 0, class 1) and a multi-class object detection device with four or more classes.

[0059] [Second Embodiment] A second embodiment of the present invention will be described for an object detection device that enables selection of a partial region of an object that could not be mechanically detected from an input image and input of a class for the partial region by an operation of an input device on the input image. In this embodiment, compared with the first embodiment, the function of the reception unit 104 is different, and the rest is basically the same as the first embodiment. Hereinafter, borrowing FIG. 1 which is a block diagram of the first embodiment, the function of the reception unit 104 suitable for the second embodiment will be described in detail.

[0060] The reception unit 104 visualizes the detection result DB 103 of the image and displays it on the display device 120, and accepts input of corrections from the operator.

[0061] FIG. 9 shows an example of a reception screen 126 that the reception unit 104 displays on the display device 120. The reception screen 126 in this example differs from the reception screen 121 shown in FIG. 4 only in that a rectangular frame indicating the outer periphery of the partial region is not displayed at the location of the two-wheeled vehicle in the image 111 displayed in the image window 124. The reason the rectangular frame is not displayed is that the detection unit 101 failed to detect the two-wheeled vehicle at that location when detecting an object from the image 111 using the dictionary 102. In order to improve such detection omissions, the object detection device needs to create teacher data, which is an input / output pair of the partial region of the two-wheeled vehicle in the image 111 and its class, and learn the dictionary 102.

[0062] The reception unit 104 accepts selection of a partial region and input of a class for the selected partial region by an operation of the input device 130 on the image 111 displayed on the display device 120. Here, in the present embodiment, the selection of the partial region means both selection of an existing partial region and creation of a new partial region. The reception unit 104 accepts selection of the partial region based on the position of the image 111 where the operation of the input device 130 is performed, and accepts input of the class based on the type of the operation. In the present embodiment, the input device 130 is a mouse, and the types of operations are left click, right click, drag and drop in the lower right direction, and drag and drop in the lower left direction.

[0063] The drag-and-drop in the lower right direction means an operation of moving (dragging) the mouse from a certain location (starting point) a on the image 111 while pressing the left button of the mouse as shown by the arrow in FIG. 10A, and then releasing (dropping) the left button at another location (ending point) b existing in the lower right direction of the starting point a. Also, the drag-and-drop in the lower left direction means an operation of moving (dragging) the mouse from a certain location (starting point) a on the image 111 while pressing the left button of the mouse as shown by the arrow in FIG. 10B, and then releasing (dropping) the left button at another location (ending point) b existing in the lower left direction of the starting point a.

[0064] The reception unit 104 receives the selection of a partial region having the position of the image 111 clicked with the left or right button within the region. Also, when the left button is clicked on the selected partial region, the reception unit 104 receives an input of class 0, and when the right button is clicked, it receives an input of a class obtained by incrementing the current class. Here, the class obtained by incrementing the current class means class 2 when the current class is class 1, class 0 when the current class is class 2, and class 1 when the current class is class 0, respectively. The reception unit 104 records the received class for the selected partial region in the modified class column of the detection result 103. The above is the same as when the left or right button is clicked in the first embodiment.

[0065] Also, when the reception unit 104 detects a drag-and-drop in the lower right direction, as shown by the dashed line in FIG. 10A, it calculates a rectangular region with the starting point a as the upper left vertex and the ending point b as the lower right vertex as the partial region, and receives an input of class 1 for that partial region. Also, when the reception unit 104 detects a drag-and-drop in the lower left direction, as shown by the dashed line in FIG. 10B, it calculates a rectangular region with the starting point a as the upper right vertex and the ending point b as the lower left vertex as the partial region, and receives an input of class 2 for that partial region. Also, the reception unit 104 records the newly calculated partial region and the information of its class in the detection result DB103.

[0066] FIG. 11 is a flowchart showing details of the operation of the reception unit 104. In FIG. 11, steps S211 to S224 are the same as steps S111 to S124 in FIG. 6. Hereinafter, with reference to FIG. 11, the operation of the reception unit 104 will be described in detail centering on the differences from FIG. 6.

[0067] When the reception unit 104 detects a drag-and-drop on the image in the image window 124 (YES in step S225), it determines the direction of the drag-and-drop (step S226). If the reception unit 104 detects a drag-and-drop in the lower-right direction, it calculates a rectangular partial region with the starting point as the upper-left vertex and the ending point as the lower-right vertex as described with reference to FIG. 10A, and accepts an input of class 1 (step S227). Next, the reception unit 104 records the pair of the calculated partial region and class 1 as one detection result in the detection result DB103 (step S228). Next, the reception unit 104 updates the image window 124 of the reception screen 126 (step S229). Then, the reception unit 104 returns to the process of step S213.

[0068] Also, if the reception unit 104 detects a drag-and-drop in the lower-left direction (step S226), it calculates a rectangular partial region with the starting point as the upper-right vertex and the ending point as the lower-left vertex as described with reference to FIG. 10B, and accepts an input of class 2 (step S230). Next, the reception unit 104 records the pair of the calculated partial region and class 2 as one detection result in the detection result DB103 (step S231). Next, the reception unit 104 updates the image window 124 of the reception screen 126 (step S232). Then, the reception unit 104 returns to the process of step S213.

[0069] FIG. 12A and FIG. 12B are explanatory diagrams of step S228 by the reception unit 104. FIG. 12A shows the detection result DB103 generated by the detection unit 101, and FIG. 12B shows the detection result DB103 after the reception unit 104 adds a record of a pair of a newly calculated partial region and class 1. The newly added partial region is a rectangle with the upper left vertex being (x3, y3) and the lower right vertex being (x4, y4), and its class is class 1. In this way, the class of the partial region to be added is recorded in the modified class column. By doing so, the generation unit 105 generates teacher data from the added partial region and class. In the case of step S231 by the reception unit 104, the difference is that class 2 is recorded in the modified class.

[0070] FIG. 13A and FIG. 13B are explanatory diagrams of step S229 by the reception unit 104. FIG. 13A shows the image window 124 before update, and FIG. 13B shows the image window 124 after update. In the image window 124 after update, a rectangular frame 125 with a dashed line that emphasizes the newly added partial region is drawn. In the case of step S232 by the reception unit 104, the difference is that a rectangular frame that emphasizes the newly added partial region is drawn with a solid line in the image window 124 after update.

[0071] As described above, according to this embodiment, an image for creating teacher data based on the number of detected objects can be efficiently and easily selected, and thus high-quality teacher data can be efficiently created. The reason is that the reception unit 104 displays the input image on the display device 120 with a display that emphasizes the partial region of the detected object, and accepts the selection of the partial region and the input of the class for this selected partial region by a click, which is one operation of the input device 130. Also, the reception unit 104 detects one operation of drag-and-drop in the lower right direction or the lower left direction on the image, and accepts the calculation of a new partial region and the input of class 1 or class 2.

[0072] [Modification Example of the Second Embodiment] Next, various modifications of the configuration of the second embodiment of the present invention will be described.

[0073] In the above second embodiment, the reception unit 104 receives an input of class 0 by a left click of a mouse gesture, receives an input of class increment by a right click, calculates a partial area of class 1 by a drag-and-drop in the lower right direction, and calculates a partial area of class 2 by a drag-and-drop in the lower left direction. However, as exemplified in the list of FIG. 14, various combinations of gesture types are possible. The gesture types include reception of an input of class 0, reception of an input of class increment, reception of calculation of a partial area of class 1, and reception of calculation of a partial area of class 2.

[0074] For example, in No. 2-1 shown in FIG. 14, the reception unit 104 receives an input of class 0 by a left click of a mouse gesture and receives an input of class increment by a double click. Further, the reception unit 104 receives reception of calculation of a partial area of class 1 by a drag-and-drop in the lower right direction and receives reception of calculation of a partial area of class 2 by a drag-and-drop in the lower left direction.

[0075] Also, in the number 2-2 shown in FIG. 14, the reception unit 104 receives an input of class 0 by a left click of the mouse gesture, and receives an input of increment of the class by a right click. Further, the reception unit 104 receives a calculation of a partial region of class 1 by a drag-and-drop in the upper left direction, and receives a calculation of a partial region of class 2 by a drag-and-drop in the upper right direction. Here, the drag-and-drop in the upper left direction is a drag-and-drop in the direction opposite to the arrow in FIG. 10A, which means an operation of moving (dragging) the mouse from a certain location (starting point) b on the image 111 with the left button of the mouse pressed, and releasing (dropping) the left button at another location (ending point) a existing in the upper left direction of the starting point b. Also, the drag-and-drop in the upper right direction is a drag-and-drop in the direction opposite to the arrow in FIG. 10B, which means an operation of moving (dragging) the mouse from a certain location (starting point) b on the image 111 with the left button of the mouse pressed, and releasing (dropping) the left button at another location (ending point) a existing in the upper right direction of the starting point b.

[0076] Also, in the number 2-3 shown in FIG. 14, compared with the number 2-2, the operation of receiving the calculation of the partial regions of class 1 and class 2 is different. In the number 2-3, the reception unit 104 receives a calculation of a partial region of class 1 by a left double click, and receives a calculation of a partial region of class 2 by a right double click. Here, the left double click means continuously clicking the left button of the mouse twice. Also, the right double click means continuously clicking the right button of the mouse twice.

[0077] FIG. 15 is an explanatory diagram of a method for calculating a partial region of class 2 from the position of a right double-click performed on the image 111 by the reception unit 104. When the reception unit 104 detects that a right double-click has been performed at the position of point c on the image 111, assuming that the center of gravity of a typical class 2 object (a four-wheeled vehicle in this embodiment) 131 coincides with point c, it calculates the circumscribed rectangle 132 of the object as the partial region. When the surveillance camera 110 with a fixed imaging field captures an object, there is a correlation between each point of the object and each point on the image. Therefore, the circumscribed rectangle of a typical-sized object with a certain point on the image as the center of gravity is almost uniquely determined. Furthermore, the circumscribed rectangle of a typical class 1 object (a two-wheeled vehicle in this embodiment) with a certain point on the image as the center of gravity is also almost uniquely determined. Thus, the reception unit 104 calculates the partial region of class 1 from the position of a left double-click performed on the image 111 by the same method.

[0078] In the example of FIG. 15, the operator is configured to double-click the right or left button on the input device 130 at the center of gravity position of the object. However, it is not limited to the center of gravity position, and any predetermined position may be used. For example, for a two-wheeled vehicle, the center of the front wheel may be double-clicked on the right or left, and for a four-wheeled vehicle, the middle between the front and rear wheels may be double-clicked on the right or left. Also, in the case of an object detection device for detecting a person, the operator is configured to double-click the left button on the input device 130 at the head of the person. In this case, the reception unit 104 estimates the person region from the head position and calculates the circumscribed rectangle of the estimated person region as the partial region.

[0079] The above numbers 2-1 to 2-3 use mouse gestures as in the second embodiment. In contrast, the numbers 2-4 to 2-7 shown in FIG. 14 use touch gestures as follows.

[0080] In No. 2-4 of FIG. 14, the reception unit 104 receives a Class 0 input by a flick, receives an input for incrementing the class by a tap, receives a calculation of a partial region of Class 1 by a swipe in the lower right direction, and receives a calculation of a partial region of Class 2 by a swipe in the lower left direction.

[0081] The swipe in the lower right direction means an operation of tracing with a fingertip from a certain location (starting point) a on the image 111 to another location (ending point) b existing in the lower right direction of the starting point a, as shown by the arrow in FIG. 16A. Also, the swipe in the lower left direction means an operation of tracing with a fingertip from a certain location (starting point) a on the image 111 to another location (ending point) b existing in the lower left direction of the starting point a, as shown by the arrow in FIG. 16B. When the reception unit 104 detects a swipe in the lower right direction, as shown by the broken line in FIG. 16A, it calculates a rectangular region with the starting point a as the upper left vertex and the ending point b as the lower right vertex as the partial region, and receives a Class 1 input for that partial region. Also, when the reception unit 104 detects a swipe in the lower left direction, as shown by the broken line in FIG. 16B, it calculates a rectangular region with the starting point a as the upper right vertex and the ending point b as the lower left vertex as the partial region, and receives a Class 2 input for that partial region.

[0082] In No. 2-5 of FIG. 14, the reception unit 104 receives a Class 0 input by a flick to the left, receives an input for incrementing the class by a flick to the right, receives a calculation of a partial region of Class 1 by a swipe downward, and receives a calculation of a partial region of Class 2 by a swipe upward.

[0083] The downward swipe means an operation of tracing with a fingertip from a certain location (starting point) a on the image 111 to another location (ending point) b existing downward from the starting point a, as shown by the arrow in Fig. 17A. The upward swipe means an operation of tracing with a fingertip from a certain location (starting point) b on the image 111 to another location (ending point) a existing upward from the starting point b, as shown by the arrow in Fig. 17B. When the reception unit 104 detects a downward swipe, it calculates a rectangle as shown by the dashed line in Fig. 17A as a partial region of class 1. The rectangle is line-symmetric with respect to the swipe line, and the vertical length is equal to the length of the swipe. The horizontal length is obtained by multiplying the length of the swipe by a predetermined ratio. As the predetermined ratio, for example, the ratio of the horizontal length to the vertical length of the circumscribed rectangle of an object of class 1 with a typical shape (a two-wheeled vehicle in this embodiment) can be used. When the reception unit 104 detects an upward swipe, it calculates a rectangle as shown by the dashed line in Fig. 17B as a partial region of class 2. The rectangle is line-symmetric with respect to the swipe line, and the vertical length is equal to the length of the swipe. The horizontal length is obtained by multiplying the length of the swipe by a predetermined ratio. As the predetermined ratio, the ratio of the horizontal length to the vertical length of the circumscribed rectangle of an object of class 2 with a typical shape (a four-wheeled vehicle in this embodiment) can be used.

[0084] In No. 2-6 of Fig. 14, the reception unit 104 receives an input of class 0 by pinching in, receives an input of class increment by pinching out, receives a calculation of a partial region of class 1 by a downward-right swipe, and receives a calculation of a partial region of class 2 by a downward-left swipe.

[0085] In No. 2-7 of Fig. 14, the reception unit 104 receives an input of class 0 by tapping, receives an input of class increment by double-tapping, receives a calculation of a partial region of class 1 by a downward-right swipe, and receives a calculation of a partial region of class 2 by a downward-left swipe. These are examples, and any other combination of gestures may be used.

[0086] In the above-described second embodiment and in Nos. 2-1 to 2-7 of FIG. 14, the reception unit 104 received an input of class 0 by a gesture of the first type and received an input of class increment by a gesture of the second type. However, it may be configured to receive an input of decrement instead of increment. That is, in the above-described second embodiment and in Nos. 2-1 to 2-7 of FIG. 14, the reception unit 104 may be configured to receive an input of class decrement by a gesture of the second type.

[0087] In the above-described second embodiment, the reception unit 104 received an input of class 0 by a left click of a mouse gesture and received an input of class increment by a right click. However, the class and the operation may be associated one-to-one, and the reception unit 104 may be configured to receive an input of class 0, class 1, and class 2 by a specific one operation, respectively. As exemplified in the list of FIG. 18, various combinations are possible for the gesture type for receiving an input of each class and the gesture type for receiving a calculation of a partial region of each class.

[0088] For example, in No. 2-11 shown in FIG. 18, the reception unit 104 receives an input of class 0 by a left click of a mouse gesture and receives an input of class 1 by a right click. Further, the reception unit 104 receives an input of class 2 by a double click, receives a calculation of a partial region of class 1 by a drag-and-drop in the lower right direction, and receives a calculation of a partial region of class 2 by a drag-and-drop in the lower left direction.

[0089] Also, No. 2-12 shown in FIG. 18 is different from No. 2-11 in the gesture type for receiving a calculation of a partial region of class 1 and class 2. In No. 2-12, the reception unit 104 receives a calculation of a partial region of class 1 by a drag-and-drop in the upper left direction and receives a calculation of a partial region of class 2 by a drag-and-drop in the upper right direction.

[0090] Also, in No. 2-13 shown in FIG. 18, the reception unit 104 receives class 0 input by a flick of a touch gesture, receives class 1 input by a tap, receives class 2 input by a double tap, receives calculation of a partial area of class 1 by a swipe in the lower right direction, and receives calculation of a partial area of class 2 by a swipe in the lower left direction.

[0091] Also, No. 2-14 shown in FIG. 18 differs from No. 2-13 in the gesture types for receiving inputs of classes 0, 1, and 2. In No. 2-14, the reception unit 104 receives class 0 input by a flick to the left, receives class 1 input by a flick upward, and receives class 2 input by a flick downward.

[0092] Also, in No. 2-15 shown in FIG. 18, the reception unit 104 receives class 0 input by a tap, receives class 1 input by a pinch-in, receives class 2 input by a pinch-out, receives calculation of a partial area of class 1 by a swipe downward, and receives calculation of a partial area of class 2 by a swipe upward.

[0093] In the above-described second embodiment, the number of classes was three (class 0, class 1, class 2), but the present invention can also be applied to a two-class object detection device having two classes (class 0, class 1) and a multi-class object detection device having four or more classes.

[0094] [Third Embodiment] In the third embodiment of the present invention, an object detection device will be described which receives inputs of different classes for the same operation depending on whether or not there is an existing partial area having the position of the operated image within the area. In this embodiment, the function of the reception unit 104 is different from that of the second embodiment, and the rest is basically the same as that of the second embodiment. Hereinafter, with reference to FIG. 1 which is a block diagram of the first embodiment, the function of the reception unit 104 suitable for the third embodiment will be described in detail.

[0095] The reception unit 104 visualizes the detection result DB 103 of the image and displays it on the display device 120, and accepts input of corrections from the operator.

[0096] The reception unit 104 accepts the selection of a partial region and the input of a class for the selected partial region by an operation of the input device 130 on the image 111 displayed on the display device 120. Here, in the present embodiment, the selection of a partial region means both the selection of an existing partial region and the selection of a new partial region. The reception unit 104 accepts the selection of a partial region based on the position of the image 111 where the operation of the input device 130 is performed, and accepts the input of a class based on the type of the operation. Also, in the present embodiment, different class inputs are accepted for the same operation depending on whether or not there is an existing partial region having the position of the image where the operation is performed within the region. In the present embodiment, the input device 130 is a mouse, and the types of one operation are left click and right click.

[0097] The reception unit 104 accepts the selection of an existing partial region having the position of the image 111 where the left click or right click is performed within the region. Also, the reception unit 104 accepts the input of class 0 when the left click is performed for the selected partial region, and accepts the input of the class obtained by incrementing the current class when the right click is performed. Here, the class obtained by incrementing the current class means class 2 when the current class is class 1, class 0 when the current class is class 2, and class 1 when the current class is class 0, respectively. The reception unit 104 records the received class for the selected partial region in the correction class column of the detection result DB 103. The above is the same as when the left click or right click is performed in the first embodiment.

[0098] In addition, when there is no existing partial region within the region that has the position of the image 111 clicked on the left or right, the reception unit 104 calculates a new rectangular partial region based on the clicked position, and accepts an input of class 1 for a left click and an input of class 2 for a right click for the partial region. As a method for calculating a new partial region based on the clicked position, the same method as the method for calculating the partial region of an object based on the double-clicked position described with reference to FIG. 15 can be used.

[0099] FIG. 19 is a flowchart showing the details of the operation of the reception unit 104. In FIG. 19, steps S211 to S224, S226 to S232 are substantially the same as steps S211 to S224, S226 to S232 in FIG. 11. The difference between FIG. 19 and FIG. 11 is that when it is determined in step S216 that there is no existing partial region that includes the clicked position, the process proceeds to step S226. Note that step S226 in FIG. 11 determines the direction (type) of drag and drop, while step S226 in FIG. 19 determines the type of click. Hereinafter, with reference to FIG. 19, the operation of the reception unit 104 will be described in detail centering on the differences from FIG. 11.

[0100] The reception unit 104 detects a left click or a right click on the image in the image window 124 (YES in step S214). At this time, when the reception unit 104 determines that there is no partial region including the click position in the image (NO in step S216), it determines whether the type of click is a left click or a right click in order to determine the class input (step S226).

[0101] If it is a left click, the reception unit 104 calculates a new rectangular partial region and accepts an input of class 1 (step S227). Next, the reception unit 104 records the pair of the calculated partial region and class 1 as one detection result in the detection result DB103 (step S228). Next, the reception unit 104 updates the image window 124 on the reception screen 126 (step S229). Then, the reception unit 104 returns to the process of step S213.

[0102] Also, if it is a right click (step S226), the reception unit 104 calculates a new rectangular partial area and receives the input of class 2 (step S230). Next, the reception unit 104 records the pair of the calculated partial area and class 2 as one detection result in the detection result DB103 (step S231). Next, the reception unit 104 updates the image window 124 of the reception screen 126 (step S232). Then, the reception unit 104 returns to the process of step S213.

[0103] As described above, according to the present embodiment, an image for creating teacher data based on the number of detected objects can be efficiently and easily selected, and thus high-quality teacher data can be efficiently created. The reason is that the reception unit 104 displays the input image on the display device 120 with a display that emphasizes the partial area of the detected object, and accepts the selection of the partial area and the input of the class for the selected partial area by a single click operation of the input device 130. Also, the reception unit 104 detects a single click operation such as a left click or a right click at a position that does not overlap with the existing partial area, and accepts the calculation of a new partial area and the input of class 1 or class 2.

[0104] [Modification Example of the Third Embodiment] In the above-described third embodiment, two types of gestures, a left click and a right click, are used, but other combinations of mouse gestures or combinations of touch gestures can be used.

[0105] [Fourth Embodiment] In the fourth embodiment of the present invention, an object detection device will be described in which the input of the correct class for the misrecognized partial region and the selection of the correct partial region of the object can be performed by operating an input device on the input image. This embodiment is basically the same as the second embodiment except that the function of the reception unit 104 is different from that of the second embodiment. Hereinafter, with reference to FIG. 1 which is a block diagram of the first embodiment, the function of the reception unit 104 suitable for the fourth embodiment will be described in detail.

[0106] The reception unit 104 visualizes the detection result DB 103 of the image and displays it on the display device 120, and receives an input for correction from the operator.

[0107] FIG. 20 shows an example of the reception screen 127 that the reception unit 104 displays on the display device 120. The reception screen 127 in this example is different only in the position of the rectangular frame indicating the outer periphery of the partial region of the four-wheeled vehicle in the image 111 displayed in the image window 124 as compared with the reception screen 121 shown in FIG. 4. The reason for the difference in the position of the rectangular frame is that when the detection unit 101 detects an object from the image 111 using the dictionary 102, a puddle of water on the road or the like is erroneously detected as a part of the partial region of the object. In this example, although the partial region is erroneously detected, the class is correctly determined. However, there is also a case where both the partial region and the class are erroneously detected. In order to improve such an erroneous detection, it is necessary to create teacher data which is an input / output pair of the erroneously detected partial region in the image 111 and its class 0, and teacher data which is an input / output pair of the correct partial region of the four-wheeled vehicle in the image 111 and its class 2, and learn the dictionary 102.

[0108] The reception unit 104 receives the selection of a partial region and the input of a class for the selected partial region by an operation of the input device 130 on the image 111 displayed on the display device 120. Here, in the present embodiment, the selection of a partial region means both the selection of an existing partial region and the selection of a new partial region. The reception unit 104 receives the selection of a partial region based on the position of the image 111 where the operation of the input device 130 is performed, and receives the input of a class based on the type of the operation. Also, in the present embodiment, when a selection operation of a new partial region that partially overlaps with an existing partial region is performed, the reception unit 104 receives the input of class 0 for the existing partial region and receives the input of a class corresponding to the type of the operation for the new partial region. In the present embodiment, the input device 130 is a mouse, and the types of operations are left click, right click, drag and drop in the lower right direction, and drag and drop in the lower left direction.

[0109] The reception unit 104 receives the selection of a partial region having the position of the image 111 clicked with the left or right button within the region. Also, the reception unit 104 receives the input of class 0 when the left button is clicked and receives the input of a class obtained by incrementing the current class when the right button is clicked for the selected partial region. Here, the class obtained by incrementing the current class means class 2 when the current class is class 1, class 0 when the current class is class 2, and class 1 when the current class is class 0, respectively. The reception unit 104 records the received class for the selected partial region in the modified class column of the detection result DB 103. The above is the same as when the left or right button is clicked in the first embodiment.

[0110] Further, when the reception unit 104 detects a drag-and-drop in the lower right direction, as shown by the dashed line in FIG. 10A, it calculates a rectangular area with the starting point a as the upper left vertex and the ending point b as the lower right vertex as a partial area, and accepts an input of class 1 for that partial area. Also, when the reception unit 104 detects a drag-and-drop in the lower left direction, as shown by the dashed line in FIG. 10B, it calculates a rectangular area with the starting point a as the upper right vertex and the ending point b as the lower left vertex as a partial area, and accepts an input of class 2 for that partial area. Further, the reception unit 104 records the newly calculated partial area and the information of its class as one detection result in the detection result DB 103. The above is the same as when drag-and-drop in the lower right direction and lower left direction is performed in the second embodiment.

[0111] Furthermore, the reception unit 104 accepts an input of class 0 for an existing partial area that partially overlaps with the newly calculated partial area. For example, as shown in FIG. 21, when the reception unit 104 calculates a partial area 125a that partially overlaps with the frame 125 for an image 111 in which a frame 125 highlighting the partial area of class 2 is displayed due to a drag-and-drop in the lower left direction (or lower right direction), it accepts an input of class 0 for the existing partial area related to the frame 125.

[0112] FIG. 22 is a flowchart showing the details of the operation of the reception unit 104. In FIG. 22, steps S311 to S327, S330 are the same as steps S211 to S227, S230 in FIG. 11. Hereinafter, with reference to FIG. 22, the operation of the reception unit 104 will be described in detail centering on the differences from FIG. 11.

[0113] When the reception unit 104 receives a class 1 input for the newly calculated partial region (step S327), it determines whether there is an existing partial region that partially overlaps with this new partial region on the image 111 (step S341). This determination is made by checking whether there is a detection result of an existing partial region that partially overlaps with the new partial region in the detection result DB103 related to the image 111. If such an existing partial region exists, the reception unit 104 receives a class 0 input for the partial region (step S342) and proceeds to step S343. If such an existing partial region does not exist, the reception unit 104 skips the process of step S342 and proceeds to step S343. In step S343, the reception unit 104 updates the detection result DB103 according to the reception results of steps S327 and S342. That is, based on the reception result of step S327, the reception unit 104 records the pair of the newly calculated partial region and class 1 as one detection result in the detection result DB103. Also, based on the reception result of step S342, the reception unit 104 records class 0 in the correction class of the existing partial region. Next, the reception unit 104 updates the image 111 on the image window 124 according to the update of the detection result DB103 in step S343 (step S344). Then the reception unit 104 returns to the process of step S313.

[0114] Also, when the reception unit 104 receives an input of class 2 for the newly calculated partial region (step S330), it determines whether there is an existing partial region that partially overlaps with this new partial region on the image 111 (step S345). If the reception unit 104 determines that such an existing partial region exists, it receives an input of class 0 for the partial region (step S346) and proceeds to step S347. If such an existing partial region does not exist, the reception unit 104 skips the process of step S346 and proceeds to step S347. In step S347, the reception unit 104 updates the detection result DB103 according to the reception results of steps S330 and S346. That is, based on the reception result of step S330, the reception unit 104 records the pair of the newly calculated partial region and class 2 as one detection result in the detection result DB103. Also, based on the reception result of step S346, the reception unit 104 records class 0 in the correction class of the existing partial region. Next, according to the update of the detection result DB103 in step S347, the reception unit 104 updates the image 111 on the image window 124 (step S348). Then the reception unit 104 returns to the process of step S313.

[0115] FIGS. 23A and 23B are explanatory diagrams of step S347 by the reception unit 104. FIG. 23A shows the detection result DB103 generated by the detection unit 101, and FIG. 23B shows the updated detection result DB103 by the reception unit 104. In FIG. 23B, a rectangular partial region with the upper left vertex (x7, y7) and the lower right vertex (x8, y8) is newly added as class 2. Also, the class of the existing rectangular partial region with the upper left vertex (x1, y1) and the lower right vertex (x2, y2) is corrected from class 2 to class 0. In the case of step S343 by the reception unit 104, the difference is that the correction class of the newly added partial region becomes class 1.

[0116] FIG. 24A and FIG. 24B are explanatory diagrams of step S348 by the reception unit 104. FIG. 24A shows the image window 124 before update, and FIG. 24B shows the image window 124 after update. In the image window 124 after update, a rectangular frame 125 with a solid line that emphasizes the newly added partial region is drawn, and the frame that was displayed in the image window 124 before update and partially overlapped with it is made invisible. In the case of step S344 by the reception unit 104, the difference is that a rectangular frame that emphasizes the newly added partial region is drawn with a dashed line in the image window 124 after update.

[0117] As described above, according to this embodiment, an image for creating teacher data based on the number of detected objects can be efficiently and easily selected, and thus high-quality teacher data can be efficiently created. The reason is that the reception unit 104 displays the input image on the display device 120 with a display that emphasizes the partial region of the detected object, and accepts the selection of the partial region and the input of the class for this selected partial region by a click, which is one operation of the input device 130. Also, the reception unit 104 detects one operation of drag-and-drop in the lower right direction or the lower left direction on the image, and accepts the calculation of a new partial region and the input of class 1 or class 2. Further, the reception unit 104 accepts the input of class 0 for an existing partial region that partially overlaps with the newly calculated partial region.

[0118] [Modification Example of the Fourth Embodiment] For the configuration of the fourth embodiment of the present invention, the same modifications as those described in the modification example of the second embodiment can be made.

[0119] [Fifth Embodiment] A fifth embodiment of the present invention will be described. Referring to FIG. 25, an object detection apparatus 1 according to this embodiment detects an object from an input image 2. The object detection apparatus 1 has a dictionary 10 and, as main functional units, a detection unit 11, a reception unit 12, a generation unit 13, and a learning unit 14. Further, the object detection apparatus 1 is connected to a display device 3 and an input device 4.

[0120] The detection unit 11 detects an object from each of a plurality of input images 2 using the dictionary 10. The reception unit 12 displays on the display device 3 a graph showing the relationship between the input image 2 and the number of partial regions in which an object is detected, and displays on the display device 3 one of the plurality of input images for creating teacher data according to the position on the graph received by operating the input device 4. The generation unit 13 generates teacher data from the input image. The learning unit 14 learns the dictionary 10 using the teacher data generated by the generation unit 13 and upgrades the dictionary 10.

[0121] Next, the operation of the object detection apparatus 1 according to this embodiment will be described. The detection unit 11 of the object detection apparatus 1 detects an object from the input image 2 using the dictionary 10 and notifies the reception unit 12 of the detection result. The reception unit 12 displays on the display device 3 a graph showing the relationship between the input image 2 and the number of partial regions in which an object is detected based on the notified detection result. Further, the reception unit 12 displays on the display device 3 one of the plurality of input images for creating teacher data according to the position on the graph received by operating the input device 4. The generation unit 13 generates teacher data from the displayed input image and notifies the learning unit 14. The learning unit 14 learns the dictionary 10 using the teacher data generated by the generation unit 13 and upgrades the dictionary 10.

[0122] According to this embodiment, an image for creating teaching data can be efficiently and easily selected based on the number of detected objects. The reason is that the reception unit 12 displays on the display device 3 a graph showing the relationship between the input image 2 and the number of partial regions where objects are detected. Further, the reception unit 12 displays on the display device 3 one of the plurality of input images for creating teaching data according to the position on the graph received by the operation of the input device 4.

[0123] This embodiment is based on the above configuration, and various additional changes as follows are possible.

[0124] The reception unit 12 may be configured to display on the display device 3 a slide bar for selecting a position on the graph.

[0125] Also, the reception unit 12 may be configured such that one axis of the graph is a column in which a plurality of input images 2 are arranged in frame number order.

[0126] Also, the reception unit 12 may be configured such that one axis of the graph is a column in which a plurality of input images 2 are arranged according to the number of partial regions of the object.

[0127] Also, the reception unit 12 may be configured to receive the selection of the partial region according to the position of the input image 2 on which one operation is performed.

[0128] [Sixth Embodiment] The object detection device 300 according to the sixth embodiment of the present invention will be described with reference to FIG. 27. The object detection device 300 includes a detection unit 301, a reception unit 302, a generation unit 303, and a learning unit 304.

[0129] The detection unit 301 detects an object from each of a plurality of input images using a dictionary.

[0130] The reception unit 302 displays on a display device a graph showing the relationship between the input image and the number of partial regions where objects are detected, and displays on the display device one input image out of a plurality of input images for creating teacher data according to the position on the graph received by operating an input device.

[0131] The generation unit 303 generates teacher data from the input image.

[0132] The learning unit 304 learns a dictionary using the teacher data.

[0133] According to the present embodiment, an image for creating teacher data can be efficiently and easily selected based on the number of detected objects. The reason is that the reception unit 302 displays on the display device a graph showing the relationship between the input image and the number of partial regions where objects are detected. Further, the reception unit 302 displays on the display device one input image out of a plurality of input images for creating teacher data according to the position on the graph received by operating the input device.

[0134] As described above, the present invention has been described by taking the above-described embodiment as an exemplary example. However, the present invention is not limited to the above-described embodiment. That is, within the scope of the present invention, various aspects understandable by those skilled in the art can be applied.

[0135] This application claims priority based on Japanese Patent Application No. 2015-055927 filed on March 19, 2015, and incorporates the entire disclosure thereof herein.

Industrial Applicability

[0136] The present invention can be used in an object detection device that detects objects such as people and vehicles from an image on a road obtained by imaging with a surveillance camera using a dictionary.

Explanation of Reference Numerals

[0137] 1... Object detection device 2…Input image 3…Display device 4…Input device 10…Dictionary 11…Detection unit 12…Reception unit 13…Generation unit 14…Learning unit 100…Object detection device 101…Detection unit 102…Dictionary 103…Detection result DB 104…Reception unit 105…Generation unit 106…Teacher data memory 107…Learning unit 111…Image 112, 113, 114, 115…Search window 120…Display device 121…Reception screen 122…Graph 123…Slider 123a…Slider bar 124…Image window 125…Rectangular frame 125a…Partial region 126, 127…Reception screen 130…Input device 131…Object 132…Bounding rectangle

Claims

1. A detection means for detecting an object from each of a plurality of input images; a receiving means for receiving an input of a position in a graph showing a relationship between the plurality of input images and the object; a display means for displaying a first image, which is one of the plurality of input images, on a display device according to the position; A selection means for selecting a class; a receiving means for receiving a designation of a partial area associated with the class in the first image; A generation means for generating data to be used for machine learning based on the partial region and the class; An object detection apparatus comprising:

2. the data includes the location of the subregion along with the class; The object detection device according to claim 1 .

3. the data includes coordinates of vertices of the subregion; The object detection device according to claim 2 .

4. the data includes a frame number of the first image that includes the subregion; The object detection device according to claim 3 .

5. The subregion is rectangular. The object detection device according to claim 1 .

6. a display means for highlighting the partial region on the first image; The object detection device according to claim 1 .

7. the accepting means accepts the designation by operating an input device on the displayed first image. The object detection device according to claim 1 .

8. The computer Detecting an object in each of a plurality of input images; Accepting an input of a position in a graph showing a relationship between the plurality of input images and the object; displaying a first image of the plurality of input images on a display device according to the position; Select a class and Accepting a designation of a partial region associated with the class in the first image; generating data to be used for machine learning based on the subregion and the class; Object detection methods.

9. detecting an object from each of a plurality of input images; A process of receiving an input of a position in a graph showing a relationship between the plurality of input images and the object; displaying a first image of the plurality of input images on a display device according to the position; The process of selecting classes; receiving a designation of a sub-region in the first image associated with the class; and generating data to be used for machine learning based on the partial region and the class. program.

Citation Information

Patent Citations

  • Video display apparatus

    JP2007280325A

  • Image processing device, image processing method

    JP2012088787A

  • Image management device, method, and program, and capsule type endoscope system

    WO2012132840A1

  • Classification assisting apparatus, classifying apparatus, and program

    JP2003317082A

  • Object detection and identification unit and method for the same, and dictionary data generation method used for object detection and identification

    JP2014059729A