Information processing system, information processing method, and program

The object detection device facilitates efficient and accurate creation of teaching data by integrating detection, reception, and learning units to allow single-operation selection and input of partial regions and classes, addressing the inefficiencies of separate operations in existing systems.

JP2025103053AInactive Publication Date: 2025-07-08NEC CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025067369
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2015-03-19
Filing Date
2025-04-16
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing object detection devices require separate operations for selecting areas and inputting classes, leading to poor workability and potentially low-quality teaching data due to incorrect area and class associations.

Method used

An object detection device that integrates a detection unit, reception unit, generation unit, and learning unit to allow for the selection and input of partial regions and classes in a single operation, using a dictionary for efficient creation of high-quality teaching data.

Benefits of technology

Enables efficient creation of high-quality teaching data by allowing simultaneous selection and input of partial regions and classes, improving the accuracy and efficiency of the object detection process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025103053000001_ABST
    Figure 2025103053000001_ABST
Patent Text Reader

Abstract

To provide an object detection device or the like which efficiently generates good-quality training data.SOLUTION: An object detection device 300 comprises: a detection unit 301 which uses a dictionary to detect objects from an input image; a reception unit 302 which displays, on a display device, the input image accompanied by a display emphasizing partial areas of the detected objects, and receives, from one operation of an input device, a selection of a partial area and an input of a class relative to the selected partial area; a generation unit 303 which generates training data from an image of the selected partial area and the inputted class; and a learning unit 304 which uses the training data to learn the dictionary.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an object detection device that detects an object from an image using a dictionary, and the like.

Background Art

[0002] Devices that analyze a captured image obtained by a imaging device and detect an object have been proposed. Patent Document 1 discloses a technique for analyzing an image obtained by imaging a substrate in a manufacturing process of a substrate such as a printed wiring board and detecting an object such as a scratch on the substrate. Patent Document 2 discloses a technique for analyzing a monitor video on a road obtained by imaging with a surveillance camera and detecting an object such as a vehicle.

[0003] In the techniques disclosed in Patent Documents 1 and 2, in order to perform the above-described detection, an object detection device is trained. And for this training, teacher data (also called training data) is required. Teacher data is an example of a pair of input and output. There are two types of teacher data: positive examples and negative examples. In order to correctly train the object detection device, both positive examples and negative examples are required. However, creating appropriate teacher data requires a lot of time and effort.

[0004] Patent Document 1 discloses a technique for creating teacher data necessary for detecting an object such as a scratch on a substrate. In this technique, a region having a luminance value different from that of a non-defective product is extracted from an image of a printed wiring board and displayed on a display, and selection of the region and input of its class (also called a category) are received from a user using a keyboard and a mouse. Specifically, the user selects a specific one of a plurality of existing regions by clicking the mouse, and then selects a desired class from a pull-down menu displayed at the time of this selection by clicking the mouse.

[0005] Patent Document 2 discloses a technique for creating teaching data necessary for detecting objects such as vehicles traveling on a road. In this technique, the area of an object is cut out from an arbitrary captured image using a dictionary, predetermined feature amounts are extracted from the area, and the dictionary is learned based on the extracted feature amounts. These series of operations are automatically performed.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0007] However, in the technique disclosed in Patent Document 1, since the user needs to perform the selection of the area used for creating the teaching data and the input of the class for the selected area by separate operations of the input device, the workability is poor.

[0008] On the other hand, in the technique disclosed in Patent Document 2, since the cutting out of the object area, the extraction of the feature amounts from the area, and the learning of the dictionary using the extracted feature amounts are automatically performed, the teaching data can be created efficiently. However, since the area and class of the object cut out using the dictionary are not always correct, the quality of the teaching data deteriorates.

[0009] An object of the present invention is to solve the above-described problems. That is, to provide an object detection device or the like that efficiently creates high-quality teaching data.

Means for Solving the Problems

[0010] In order to solve the above-described problems, a first feature of the present invention is A detection unit that detects an object from an input image using a dictionary, A reception unit that attaches a display for emphasizing a partial region of the detected object to display the input image on a display device, and accepts selection of the partial region and input of a class for the selected partial region by one operation of an input device, A generation unit that generates teaching data from the image of the selected partial region and the input class, A learning unit that learns the dictionary with the teaching data, and an object detection device having the above.

[0011] A second feature of the present invention is, detecting an object from an input image using a dictionary, attaching a display for emphasizing a partial region of the detected object to display the input image on a display device, accepting selection of the partial region and input of a class for the selected partial region by one operation of an input device, generating teaching data from the image of the selected partial region and the input class, and an object detection method of learning a dictionary with the teaching data.

[0012] A third feature of the present invention is, a computer, a detection unit that detects an object from an input image using a dictionary, a reception unit that attaches a display for emphasizing a partial region of the detected object to display the input image on a display device, and accepts selection of the partial region and input of a class for the selected partial region by one operation of an input device, a generation unit that generates teaching data from the image of the selected partial region and the input class, a learning unit that learns the dictionary with the teaching data, and a recording medium storing a program for causing the above to function.

Advantages of the Invention

[0013] Since the present invention has the above-described configuration, high-quality teacher data can be efficiently created.

Brief Description of Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10A

Figure 10B

Figure 11

Figure 12A

Figure 12B

Figure 13A

Figure 13B

Figure 14

Figure 15

Figure 16A

Figure 16B

Figure 17A

Figure 17B

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23A

Figure 23B

Figure 24A

Figure 24B

Figure 25

Figure 26

Figure 27

Embodiments for Carrying Out the Invention

[0015] Next, embodiments of the present invention will be described in detail with reference to the drawings. In the following description of the drawings, the same or similar parts are denoted by the same or similar reference numerals. However, the drawings schematically represent the configurations in the embodiments of the present invention. Furthermore, the embodiments of the present invention described below are examples, and can be appropriately modified within the scope of the same essence. [First Embodiment] Referring to FIG. 1, the object detection device 100 according to the first embodiment of the present invention has a function of analyzing a monitor video on a road obtained by imaging with a surveillance camera 110 and detecting an object. In this embodiment, the objects to be detected are two types: a two-wheeled vehicle and a four-wheeled vehicle. That is, the object detection device 100 is a type of multi-class classifier and detects two specific types of objects from an image.

[0016] The object detection device 100 mainly includes a detection unit 101, a dictionary 102, a detection result DB (detection result database) 103, a reception unit 104, a generation unit 105, a teacher data memory 106, and a learning unit 107. In addition, a display device 120 and an input device 130 are connected to the object detection device 100. The object detection device 100 may be constituted by, for example, an information processing device 230 and a storage medium for recording a program 240 as shown in FIG. 26. The information processing device 230 includes an arithmetic processing unit 210 such as one or more microprocessors and a storage unit 220 such as a semiconductor memory or a hard disk. The storage unit 220 stores the dictionary 102, the detection result DB 103, the teacher data memory 106, and the like. The program 240 is read into the memory from an external computer-readable recording medium when the object detection device 100 is started up or the like, and by controlling the operation of the arithmetic processing unit 210, functional means such as the detection unit 101, the reception unit 104, the generation unit 105, and the learning unit 107 are realized on the arithmetic processing unit 210.

[0017] The display device 120 is composed of a screen display device such as an LCD (Liquid Crystal Display) or a PDP (Plasma Display Panel), and has a function of displaying various information such as a detection result on the screen in response to an instruction from the object detection device 100.

[0018] The input device 130 is an operation input device such as a keyboard or a mouse. The input device 130 detects the operator's operation as an input and outputs it to the object detection device 100. In the present embodiment, a mouse is used as the input device 130. Also, the mouse to be used is assumed to be capable of performing two types of mouse gestures: left click and right click.

[0019] The detection unit 101 inputs the images obtained by imaging with the monitoring camera 110 one by one in chronological order, and detects an object from each of the input images using the dictionary 102. The detection unit 101 has a function of storing the detection result in the detection result DB 103. One image refers to one frame of the image captured by the monitoring camera 110. A series of frames output from the monitoring camera 110 are assigned consecutive frame numbers for identification.

[0020] FIG. 2 is an explanatory diagram of the concept of detecting an object from an image. The detection unit 101 sets a search window 112 for object detection for the image 111, and extracts the feature amount of the image within the search window 112. Further, the detection unit 101 calculates a likelihood indicating the similarity between the extracted feature amount and the dictionary 102, and detects an object by determining whether the image within the search window 112 is an object or not based on the calculated likelihood. When the detection unit 101 detects an object, it associates the detection result having the information of the partial region where the object is detected and the class described later with the identification information of the image 111 and stores it in the detection result DB 103. As the information of the partial region of the object, the position information of the search window 112 determined to be an object on the image 111 is used. Here, as the position information, when the search window 112 is rectangular, for example, the coordinate values of the upper left and lower right vertices are used. Also, the classes related to the object are a total of three classes: class 1 representing the class of two-wheeled vehicles, class 2 representing the class of four-wheeled vehicles, and class 0 representing neither of them. Further, the identification information of the image 111 is, for example, the frame number of the image.

[0021] The detection unit 101 changes the position and size of the search window 112 and repeats the same operation as described above to search for objects with different positions and sizes existing in the image 111 without omission. In the example shown in FIG. 2, three objects are detected from the image 111. The detection unit 101 detects the first object with the search window at the position and size indicated by reference numeral 113, the second object with the search window at the position and size indicated by reference numeral 114, and the third object with the search window at the position and size indicated by reference numeral 115. In FIG. 2, the horizontal direction of the image is the X-axis and the vertical direction is the Y-axis. When the detection process for one image 111 is completed, the detection unit 101 repeats the same process for the next image.

[0022] FIG. 3 is an example of the detection result DB103. The detection result DB103 in this example records the detection results corresponding to the frame numbers. The detection result is composed of information on the partial region of the object and the class. For example, the detection result DB103 records a pair of information on the partial regions of three objects and the class corresponding to the frame number 001. One of them is a pair of a partial region with the coordinates of the upper left vertex being (x1, y1) and the coordinates of the lower right vertex being (x2, y2) and class 2, another one is a pair of a partial region with the coordinates of the upper left vertex being (x3, y3) and the coordinates of the lower right vertex being (x4, y4) and class 1, and the remaining one is a pair of a partial region with the coordinates of the upper left vertex being (x5, y5) and the coordinates of the lower right vertex being (x6, y6) and class 1.

[0023] Also, the detection result DB103 shown in FIG. 3 has a column for recording the corrected class for each detection result. The corrected class corresponds to the class input by the reception unit 104. At the time when the detection unit 101 outputs the detection result to the detection result DB103, all the columns of the corrected class are NULL (indicated by - in FIG. 3).

[0024] The reception unit 104 visualizes the detection results of the images stored in the detection result DB103 and displays them on the display device 120, and accepts input of corrections from the operator.

[0025] FIG. 4 shows an example of a reception screen 121 that the reception unit 104 displays on the display device 120. In this example, the reception screen 121 displays a graph 122 at the lower part of the screen, a slider 123 at the upper part thereof, and an image window 124 at the further upper part.

[0026] The graph 122 is a graph showing the relationship between a plurality of images including the image 111 and the number of partial regions where an object is detected in each screen. When the horizontal direction of the reception screen 121 is the X-axis and the vertical direction is the Y-axis, the X-axis of the graph 122 indicates the frame number of the image, and the Y-axis indicates the number of partial regions where an object is detected, that is, the object detection number. That is, the reception unit 104 uses, as the X-axis of the graph 122, a column in which the input images corresponding to the detection results stored in the detection result DB 103 are arranged in ascending or descending order of the frame number. However, the X-axis may be other than a column in which the input images are arranged in ascending or descending order of the frame number. For example, the X-axis may be a column in which the input images are arranged by the number of partial regions of the object (that is, the object detection number). In this case, the reception unit 104 sorts the input images of the detection results 103 stored in the detection result DB 103 in ascending or descending order by the object detection number, and uses the column of the sorted input images as the X-axis of the graph 122. Also, in FIG. 4, the graph 122 is a bar graph, but the type of the graph 122 is not limited to the bar graph, and may be other types such as a line graph.

[0027] The slider 123 is a GUI (Graphical User Interface) for selecting a position on the X-axis of the graph 122, that is, for selecting an image. The slider 123 includes a slide bar 123a, and the position on the X-axis of the graph 122 is selected by operating the slide bar 123a left and right with the input device 130.

[0028] The image window 124 displays one image 111 selected by the operation of the slider 123 among a plurality of images in frame units. On the image 111 displayed in the image window 124, a display for emphasizing the partial region of the detected object is added. In the example of FIG. 4, a rectangular frame 125 indicating the outer periphery of the partial region of the object is displayed. However, the display for emphasizing the partial region of the object is not limited to the rectangular frame 125, and any display form such as a display form that makes the luminance of the entire partial region higher or lower than other regions, or a display form that displays hatching on the entire partial region may be used.

[0029] Also, the display for emphasizing the partial region of the detected object is displayed in different display forms according to the class of the detected object. In FIG. 4, the partial region of the object of class 1 is displayed with a dashed frame 125, and the partial region of the object of class 2 is displayed with a solid line frame 125. In addition to making the line types of the frame 125 different, the display color of the frame 125 may be made different. Also, a numerical value indicating the class may be displayed near the partial image.

[0030] Generally, on a GUI slider, a scale and a label serving as a guide for the operator's operation are displayed. The reception screen 121 shown in FIG. 4 displays a graph 122 indicating the number of object detections for each image instead of the scale and the label. By displaying such a graph 122, the operator can easily select one image from a plurality of images based on the number of object detections.

[0031] In addition, the reception unit 104 has a function of receiving the selection of a partial region and the input of a class for the selected partial region by operating the input device 130 on the image 111 displayed on the display device 120. The reception unit 104 receives the selection of a partial region based on the position of the image 111 where the operation of the input device 130 is performed, and receives the input of a class based on the type of the operation. In the present embodiment, the input device 130 is a mouse, and the types of operations are left click and right click. The left click means an operation of single-clicking the left button of the mouse. The right click means an operation of single-clicking the right button of the mouse. The reception unit 104 receives the selection of a partial region having the position of the image 111 where the left click or right click is performed within the region. In addition, the reception unit 104 receives the input of class 0 for the selected partial region when a left click is performed, and receives the input of a class obtained by incrementing the current class when a right click is performed. Here, the class obtained by incrementing the current class means class 2 when the current class is class 1, class 0 when the current class is class 2, and class 1 when the current class is class 0, respectively. The reception unit 104 records the received class for the selected partial region in the modified class column of the detection result DB 103.

[0032] The generation unit 105 generates teacher data from the image of the partial region selected by the reception unit 104 and the input class, and stores the teacher data in the teacher data memory 106. The generation unit 105 generates teacher data for each partial region in which any one of classes 0, 1, and 2 is recorded in the modified class column in the detection result DB 103 shown in FIG. 3. The teacher data is an input / output pair of the image of the partial region or the feature amount extracted from the image of the partial region and the class recorded in the modified class column.

[0033] The learning unit 107 learns the dictionary 102 using the teacher data stored in the teacher data memory 106. Since the method of learning the dictionary 102 using the teacher data is widely known, its description will be omitted. The dictionary 102 is upgraded by learning with the teacher data. The detection unit 101 uses this upgraded dictionary 102 to detect an object from each image input again from the surveillance camera 110.

[0034] FIG. 5 is a flowchart showing the operation of the present embodiment. Hereinafter, the operation of the present embodiment will be described with reference to FIG. 5.

[0035] The detection unit 101 of the object detection device 100 inputs one by one in chronological order the images obtained by imaging with the surveillance camera 110 (step S101). Next, the detection unit 101 detects an object from the input image using the dictionary 102 by the method described with reference to FIG. 2, and stores the detection result in the detection result DB 103 (step S102). If there is a next image obtained by imaging with the surveillance camera 110 (step S103), the detection unit 101 returns to step S100 and repeats the same process as the above-described process. As a result, the detection results of a plurality of images as shown in FIG. 3 are accumulated in the detection result DB 103. When the processing of the detection unit 101 for a series of images captured by the surveillance camera 110 is completed, the processing by the reception unit 104 is started.

[0036] The reception unit 104 displays a reception screen 121 as shown in FIG. 4 on the display device 120 and receives an input operation of the input device (step S104). Details of this step S104 will be described later.

[0037] When the reception by the reception unit 104 is completed, the generation unit 105 generates teacher data from the image of the partial region selected by the reception unit 104 and the input class (step S105). Next, the learning unit 107 learns the dictionary 102 using the teacher data stored in the teacher data memory 106 (step S106).

[0038] FIG. 6 is a flowchart showing details of the operation of the reception unit 104 performed in step S104. Hereinafter, the operation of the reception unit 104 will be described in detail with reference to FIG. 6.

[0039] The reception unit 104 calculates the number of object detections per image from the detection result DB 103 (step S111). Referring to FIG. 3, since three partial regions are detected in the image of frame number 001, the reception unit 104 sets the number of object detections of frame number 001 to 3. Similarly, the number of object detections of images with other frame numbers is calculated.

[0040] Next, the reception unit 104 displays the initial reception screen 121 on the display device 120 (step S112). As shown in FIG. 4, the initial reception screen 121 displays a graph 122, a slider 123, and an image window 124. The graph 122 displays the number of object detections per frame number calculated in step S111. The slider 123 has a slide bar 123a placed at a predetermined position. The image window 124 displays the image of the object number on the X-axis of the graph 122 indicated by the slide bar 123a. Also, the partial regions of the objects detected in the image displayed in the image window 124 are emphasized by frames 125 of line types corresponding to the classes of the partial regions.

[0041] Next, the reception unit 104 determines whether the slide bar 123a of the slider 123 has been operated (step S113), whether a click has been performed on the image in the image window (step S114), and whether the reception end condition has been satisfied (step S115). The reception end condition may be, for example, that a reception end command has been input from the operator, or that no input operation has been performed for a certain period of time or more.

[0042] When the reception unit 104 detects a left click or a right click on the image in the image window 124 (YES in step S114), it determines whether a partial region including the click position exists in the image (step S116). This determination is made by checking whether the coordinate value of the clicked position is within any partial region of the image displayed in the image window 124. If the partial region including the click position does not exist (NO in step S116), the reception unit 104 ignores the click. On the other hand, if the partial region including the click position exists (YES in step S116), the reception unit 104 determines that the partial region including the click position has been selected, and determines whether the type of click is a left click or a right click in order to determine the input of the class (step S117).

[0043] If it is a left click, the reception unit 104 receives the input of class 0 for the selected partial region (step S118). Next, the reception unit 104 updates the correction class corresponding to the selected partial region in the detection result in the detection result DB103 to the received class 0 (step S119). Next, the reception unit 104 makes the frame 125 displayed in the selected partial region of the image displayed in the image window 124 invisible (step S120).

[0044] Also, if a right click is detected, the reception unit 104 receives an input of a class obtained by incrementing the current class of the selected partial region (step S121). The current class of the selected partial region is the class of the detection result when the correction class corresponding to the selected partial region in the detection result DB103 is NULL, and is the class described in the correction class when the correction class is not NULL. Next, the reception unit 104 updates the correction class corresponding to the selected partial region of the detection result in the detection result DB103 to the class after the increment (step S122). Next, the reception unit 104 updates the display of the frame 125 displayed in the selected partial region of the image displayed in the image window 124 (step S123). Specifically, if the class after the increment is class 2, the reception unit 104 displays the frame 125 with a solid line. Also, if the class after the increment is class 0, the reception unit 104 hides the frame 125. Also, if the class after the increment is class 1, the reception unit 104 displays the frame 125 with a dashed line.

[0045] On the other hand, when the reception unit 104 detects that the slider 123a has been operated (step S113), the reception unit 104 updates the slider 123a on the reception screen 121 and the image in the image window 124 (step 124). Specifically, the reception unit 104 moves the display position of the slider 123a according to the input operation. Also, the reception unit 104 displays, in the image window 124, an image of the object number on the X-axis of the graph 122 indicated by the slider 123a after the movement.

[0046] Furthermore, the reception unit 104 determines whether a partial region on the image displayed in the image window 124 has the same position and the same pre-input class as the partial region of the image for which the class was input in the past. When it is determined that they have the same position and the same pre-input class, the reception unit 104 hides the frame 125 that emphasizes the partial region. For example, the positions of the third partial images in frame numbers 001, 002, and 003 in FIG. 3 are all rectangles at the same position with (x5, y5) and (x6, y6) as the upper left vertex and the lower right vertex, respectively, and their pre-input classes are also the same class 1. Therefore, for example, after the class of the third partial image in the image of frame number 001 is changed from class 1 to class 0, when the images of frame numbers 002 and 003 are displayed in the image window 124, the frame 125 is not displayed for the third partial images of the images of frame numbers 002 and 003. However, class 1 is still recorded in the classes of their detection results.

[0047] As described above, according to this embodiment, high-quality teacher data can be efficiently created. The reason is that the reception unit 104 displays the input image on the display device 120 with a display that emphasizes the partial region of the detected object, and accepts the selection of the partial region and the input of the class for the selected partial region by a click, which is an operation of the input device 130.

[0048] According to this embodiment, one image can be easily selected from a plurality of images based on the number of detected objects. The reason is that the reception unit 104 displays a graph 122 showing the relationship between the input image and the number of partial regions where objects are detected on the display device 120, and selects an input image for which the detection result is to be displayed on the display device 120 according to the position on the graph 122 received by operating the slider 123. In this embodiment, an image to be processed can be selected based on the number of detected objects. Therefore, depending on the skill level and preference of the operator, an image with a large number of detected objects can be preferentially processed, or conversely, an image with a small number of detected objects can be preferentially processed. Alternatively, the operator can easily preferentially process an image with a medium number of detected objects. Generally, preferentially processing an image with a large number of detected objects shortens the total working time, but in the case of an image with a lot of overlap between partial regions, it is not easy for an inexperienced operator to confirm it.

[0049] Also, generally, when a surveillance camera 110 with a fixed field of view captures images of two-wheeled vehicles or four-wheeled vehicles traveling on a road at regular intervals, stationary objects commonly appear at the same position in a plurality of consecutive frame images. Therefore, for example, when a puddle formed on the road is erroneously detected as a two-wheeled vehicle, the same misrecognition result appears at the same position in a plurality of consecutive frame images. It is sufficient to correct any one of the same misrecognition results to generate teacher data, and it is rather troublesome if the same misrecognition results that have already been corrected are highlighted over a large number of frame images. In this embodiment, the reception unit 104 makes invisible a frame 125 that highlights partial regions having the same position and the same pre-input class as the partial regions of the image for which a class input has been performed in the past. Thereby, such troublesomeness can be eliminated.

[0050] [Modifications of the First Embodiment] Next, various modifications in which the configuration of the first embodiment is changed will be described.

[0051] In the first embodiment described above, the reception unit 104 received the input of class 0 by a left click of the mouse gesture and received the input of increment of the class by a right click. However, as exemplified in the list of FIG. 7, various combinations are possible for the gesture type for receiving the input of class 0 and the gesture type for receiving the input of increment of the class.

[0052] For example, in No. 1-1 shown in FIG. 7, the reception unit 104 receives the input of class 0 by a left click of the mouse gesture and receives the input of increment of the class by a double click. A double click means an operation of continuously performing two single clicks on the left button of the mouse.

[0053] Also, in No. 1-2 shown in FIG. 7, the reception unit 104 receives the input of class 0 by a left click of the mouse gesture and receives the input of increment of the class by a drag and drop. A drag and drop means an operation of moving (dragging) the mouse while pressing the left button of the mouse and releasing (dropping) the left button at another location. When using drag and drop, it is determined whether a drag and drop is performed on the image window in step S114 of FIG. 6, and in step S116, for example, it is determined whether there is a partial region including the start point of the drag and the drop point.

[0054] The numbers 1-1 and 1-2 above use mouse gestures, just like in the first embodiment. In contrast, the numbers 1-3 to 1-6 shown in FIG. 7 use touch gestures. Touch gestures mean touch operations performed on a touch panel using a part of the human body such as a fingertip. Touch gestures include many types such as tap (a light tapping operation with a finger or the like), double tap (an operation of tapping twice with a finger or the like), flick (a flicking or swiping operation with a finger or the like), swipe (a tracing operation with a finger or the like), pinch in (a pinching or narrowing operation with multiple fingers or the like), and pinch out (a spreading operation with multiple fingers or the like). When using touch gestures, a touch panel is used as the input device 130. The use of a touch panel is also useful, for example, when the present invention is implemented in mobile devices such as smartphones and tablet terminals.

[0055] In the number 1-3 of FIG. 7, the reception unit 104 receives an input of class 0 by flicking and receives an input of class increment by swiping. Also, in the number 1-4 of FIG. 7, the reception unit 104 receives an input of class 0 by flicking to the left and receives an input of class increment by flicking to the right. In the number 1-5 of FIG. 7, the reception unit 104 receives an input of class 0 by pinch in and receives an input of class increment by pinch out. In the number 1-6 of FIG. 7, the reception unit 104 receives an input of class 0 by tap and receives an input of class increment by double tap. These are examples, and any other combination of gestures may be used.

[0056] In the above-described first embodiment and in numbers 1-1 to 1-6 of FIG. 7, the reception unit 104 received an input of class 0 by a first type of gesture and received an input of increment of the class by a second type of gesture. However, the reception unit 104 may be configured to receive an input of decrement instead of increment. That is, in the above-described first embodiment and in numbers 1-1 to 1-6 of FIG. 7, the reception unit 104 may receive an input of class 0 by a first type of gesture and receive an input of decrement of the class by a second type of gesture.

[0057] In the above-described first embodiment, an input of class 0 was received by a left click of a mouse gesture, and an input of increment of the class was received by a right click. However, the class and the operation may be associated one-to-one, and inputs of class 0, class 1, and class 2 may be received by specific one operation respectively. As exemplified in the list of FIG. 8, various combinations are possible for the gesture types for receiving inputs of each class.

[0058] For example, in numbers 1-11 shown in FIG. 8, the reception unit 104 receives an input of class 0 by a left click of a mouse gesture, receives an input of class 1 by a right click, and receives an input of class 2 by a double click. Also, in number 1-12 shown in FIG. 8, the reception unit 104 receives inputs of class 0 and class 1 by the same mouse gestures as in number 1-11, and receives an input of class 2 by drag and drop. The above are examples of using mouse gestures. In contrast, numbers 1-13 to 1-15 shown in FIG. 8 use touch gestures. For example, in number 1-13 shown in FIG. 8, the reception unit 104 receives an input of class 0 by a flick, receives an input of class 1 by a tap, and receives an input of class 2 by a double tap. Also, in number 1-14 shown in FIG. 8, the reception unit 104 receives an input of class 0 by a leftward flick, receives an input of class 1 by an upward flick, and receives an input of class 2 by a downward flick. Also, in number 1-15 shown in FIG. 8, the reception unit 104 receives an input of class 0 by a tap, receives an input of class 1 by a pinch-in, and receives an input of class 2 by a pinch-out.

[0059] In the first embodiment above, the number of classes was three (class 0, class 1, class 2), but it can also be applied to a two-class object detection device with two classes (class 0, class 1) and a multi-class object detection device with four or more classes.

[0060] [Second Embodiment] A second embodiment of the present invention will be described for an object detection device that enables selection of a partial region of an object that could not be mechanically detected from an input image and input of a class for the partial region by an operation of an input device on the input image. In this embodiment, the function of the reception unit 104 is different compared to the first embodiment, and the rest is basically the same as the first embodiment. Hereinafter, borrowing FIG. 1 which is a block diagram of the first embodiment, the function of the reception unit 104 suitable for the second embodiment will be described in detail.

[0061] The reception unit 104 has a function of visualizing the detection result DB 103 of the image and displaying it on the display device 120, and receiving an input of correction from the operator.

[0062] FIG. 9 shows an example of the reception screen 126 displayed on the display device 120 by the reception unit 104. The reception screen 126 in this example differs only in that a rectangular frame indicating the outer periphery of the partial region is not displayed at the location of the two-wheeled vehicle in the image 111 displayed in the image window 124 as compared with the reception screen 121 shown in FIG. 4. The reason why the rectangular frame is not displayed is that the detection unit 101 failed to detect the two-wheeled vehicle at that location when detecting the object from the image 111 using the dictionary 102. In order to improve such detection omission, the object detection device needs to create teacher data which is an input / output pair of the partial region of the two-wheeled vehicle in the image 111 and its class, and learn the dictionary 102.

[0063] The reception unit 104 has a function of receiving the selection of the partial region and the input of the class for the selected partial region by the operation of the input device 130 on the image 111 displayed on the display device 120. Here, in the present embodiment, the selection of the partial region means both the selection of the existing partial region and the creation of a new partial region. The reception unit 104 receives the selection of the partial region according to the position of the image 111 where the operation of the input device 130 is performed, and receives the input of the class according to the type of the operation. In the present embodiment, the input device 130 is a mouse, and the types of operations are left click, right click, drag and drop in the lower right direction, and drag and drop in the lower left direction.

[0064] The drag-and-drop in the lower right direction means an operation of moving (dragging) the mouse from a certain location (starting point) a on the image 111 while holding down the left button of the mouse as shown by the arrow in Fig. 10A, and then releasing (dropping) the left button at another location (ending point) b in the lower right direction of the starting point a. Also, the drag-and-drop in the lower left direction means an operation of moving (dragging) the mouse from a certain location (starting point) a on the image 111 while holding down the left button of the mouse as shown by the arrow in Fig. 10B, and then releasing (dropping) the left button at another location (ending point) b in the lower left direction of the starting point a.

[0065] The reception unit 104 receives the selection of a partial region having the position of the image 111 clicked by the left click or the right click within the region. Also, when the left click is made on the selected partial region, the reception unit 104 receives an input of class 0, and when the right click is made, it receives an input of the class obtained by incrementing the current class. Here, the class obtained by incrementing the current class means class 2 when the current class is class 1, class 0 when the current class is class 2, and class 1 when the current class is class 0 respectively. The reception unit 104 records the received class in the modified class column of the detection result 103 for the selected partial region. The above is the same as when the left click or the right click is made in the first embodiment.

[0066] Also, when the reception unit 104 detects a drag-and-drop in the lower right direction, as shown by the dashed line in Fig. 10A, it calculates a rectangular region with the starting point a as the upper left vertex and the ending point b as the lower right vertex as the partial region, and receives an input of class 1 for that partial region. Also, when the reception unit 104 detects a drag-and-drop in the lower left direction, as shown by the dashed line in Fig. 10B, it calculates a rectangular region with the starting point a as the upper right vertex and the ending point b as the lower left vertex as the partial region, and receives an input of class 2 for that partial region. Also, the reception unit 104 records the newly calculated partial region and the information of its class in the detection result DB103.

[0067] FIG. 11 is a flowchart showing the details of the operation of the reception unit 104. In FIG. 11, steps S211 to S224 are the same as steps S111 to S124 in FIG. 6. Hereinafter, with reference to FIG. 11, the operation of the reception unit 104 will be described in detail centering on the differences from FIG. 6.

[0068] When the reception unit 104 detects a drag-and-drop on the image in the image window 124 (YES in step S225), it determines the direction of the drag-and-drop (step S226). If the reception unit 104 detects a drag-and-drop in the lower-right direction, it calculates a rectangular partial region with the starting point as the upper-left vertex and the ending point as the lower-right vertex as described with reference to FIG. 10A, and accepts the input of class 1 (step S227). Next, the reception unit 104 records the pair of the calculated partial region and class 1 as one detection result in the detection result DB103 (step S228). Next, the reception unit 104 updates the image window 124 of the reception screen 126 (step S229). Then, the reception unit 104 returns to the process of step S213.

[0069] Also, if the reception unit 104 detects a drag-and-drop in the lower-left direction (step S226), it calculates a rectangular partial region with the starting point as the upper-right vertex and the ending point as the lower-left vertex as described with reference to FIG. 10B, and accepts the input of class 2 (step S230). Next, the reception unit 104 records the pair of the calculated partial region and class 2 as one detection result in the detection result DB103 (step S231). Next, the reception unit 104 updates the image window 124 of the reception screen 126 (step S232). Then, the reception unit 104 returns to the process of step S213.

[0070] FIG. 12A and FIG. 12B are explanatory diagrams of step S228 by the reception unit 104. FIG. 12A shows the detection result DB103 generated by the detection unit 101, and FIG. 12B shows the detection result DB103 after the reception unit 104 adds a record of a pair of a newly calculated partial region and class 1. The newly added partial region is a rectangle with the upper left vertex being (x3, y3) and the lower right vertex being (x4, y4), and its class is class 1. In this way, the class of the partial region to be added is recorded in the modified class column. By doing so, the generation unit 105 generates teacher data from the added partial region and class. In the case of step S231 by the reception unit 104, the difference is that class 2 is recorded in the modified class.

[0071] FIG. 13A and FIG. 13B are explanatory diagrams of step S229 by the reception unit 104. FIG. 13A shows the image window 124 before update, and FIG. 13B shows the image window 124 after update. In the image window 124 after update, a rectangular frame 125 with a dashed line that emphasizes the newly added partial region is drawn. In the case of step S232 by the reception unit 104, the difference is that a rectangular frame that emphasizes the newly added partial region is drawn with a solid line in the image window 124 after update.

[0072] As described above, according to the present embodiment, high-quality teacher data can be efficiently created. The reason is that the reception unit 104 displays the input image on the display device 120 with a display that emphasizes the partial region of the detected object, and accepts the selection of the partial region and the input of the class for the selected partial region by a click, which is one operation of the input device 130. Further, the reception unit 104 detects one operation of drag-and-drop in the lower right direction or the lower left direction on the image, and accepts the calculation of a new partial region and the input of class 1 or class 2.

[0073] [Modifications of the Second Embodiment] Next, various modifications in which the configuration of the second embodiment is changed will be described.

[0074] In the above-described second embodiment, the reception unit 104 receives an input of class 0 by a left click of a mouse gesture, receives an input of increment of the class by a right click, calculates a partial region of class 1 by a drag-and-drop in the lower right direction, and calculates a partial region of class 2 by a drag-and-drop in the lower left direction. However, as exemplified in the list of FIG. 14, various combinations are possible for the gesture type for receiving an input of class 0, the gesture type for receiving an input of increment of the class, the gesture type for receiving the calculation of the partial region of class 1, and the gesture type for receiving the calculation of the partial region of class 2.

[0075] For example, in No. 2-1 shown in FIG. 14, the reception unit 104 receives an input of class 0 by a left click of a mouse gesture and receives an input of increment of the class by a double click. Further, the reception unit 104 receives the calculation of the partial region of class 1 by a drag-and-drop in the lower right direction and receives the calculation of the partial region of class 2 by a drag-and-drop in the lower left direction.

[0076] Also, in No. 2-2 shown in FIG. 14, the reception unit 104 receives an input of class 0 by a left click of a mouse gesture, receives an input of increment of the class by a right click, receives the calculation of the partial region of class 1 by a drag-and-drop in the upper left direction, and receives the calculation of the partial region of class 2 by a drag-and-drop in the upper right direction. Here, the drag-and-drop in the upper left direction is a drag-and-drop in the direction opposite to the arrow in FIG. 10A, and means an operation of moving (dragging) the mouse from a certain location (starting point) b on the image 111 while holding down the left button of the mouse and releasing (dropping) the left button at another location (ending point) a existing in the upper left direction of the starting point b. Also, the drag-and-drop in the upper right direction is a drag-and-drop in the direction opposite to the arrow in FIG. 10B, and means an operation of moving (dragging) the mouse from a certain location (starting point) b on the image 111 while holding down the left button of the mouse and releasing (dropping) the left button at another location (ending point) c existing in the upper right direction of the starting point b.

[0077] Also, as for number 2-3 shown in FIG. 14, the operation of receiving the calculation of the partial regions of class 1 and class 2 is different as compared with number 2-2. In number 2-3, the reception unit 104 receives the calculation of the partial region of class 1 by a left double click and receives the calculation of the partial region of class 2 by a right double click. Here, a left double click means continuously clicking the left button of the mouse twice. Also, a right double click means continuously clicking the right button of the mouse twice.

[0078] FIG. 15 is an explanatory diagram of a method for calculating the partial region of class 2 from the position of the right double click performed by the reception unit 104 on the image 111. When the reception unit 104 detects that a right double click has been performed at the position of point c on the image 111, assuming that the center of gravity of a typical object of class 2 (a four-wheeled vehicle in this embodiment) 131 coincides with point c, it calculates the circumscribed rectangle 132 of the object as the partial region. When the surveillance camera 110 with a fixed imaging field captures a subject, since there is a correlation relationship between each point of the subject and each point on the image, the circumscribed rectangle of an object of a typical size with a certain point on the image as the center of gravity is almost uniquely determined. Furthermore, the circumscribed rectangle of a typical object of class 1 (a two-wheeled vehicle in this embodiment) with a certain point on the image as the center of gravity is also almost uniquely determined. Thereby, the reception unit 104 calculates the partial region of class 1 from the position of the left double click performed on the image 111 by a similar method.

[0079] In the example of FIG. 15, the operator is configured to double - click the center - of - gravity position of the object with the right or left button via the input device 130. However, it is not limited to the center - of - gravity position, and any predetermined position may be used. For example, in the case of a two - wheeled vehicle, the center of the front wheel may be double - clicked with the right or left button, and in the case of a four - wheeled vehicle, the middle between the front and rear wheels may be double - clicked with the right or left button. Also, in the case of an object detection device that detects a person, the operator can be configured to double - click the head of the person via the input device 130. In this case, the reception unit 104 estimates the person area from the head position and calculates the circumscribed rectangle of the estimated person area as a partial area.

[0080] The above numbers 2 - 1 to 2 - 3 use mouse gestures as in the second embodiment. On the other hand, numbers 2 - 4 to 2 - 7 shown in FIG. 14 use touch gestures as follows.

[0081] In number 2 - 4 of FIG. 14, the reception unit 104 receives an input of class 0 by a flick, receives an input of class increment by a tap, receives a calculation of a partial area of class 1 by a swipe in the lower - right direction, and receives a calculation of a partial area of class 2 by a swipe in the lower - left direction.

[0082] A swipe in the lower - right direction means an operation of tracing with a fingertip from a certain location (starting point) a on the image 111 to another location (ending point) b existing in the lower - right direction of the starting point a as shown by the arrow in FIG. 16A. Also, a swipe in the lower - left direction means an operation of tracing with a fingertip from a certain location (starting point) a on the image 111 to another location (ending point) b existing in the lower - left direction of the starting point a as shown by the arrow in FIG. 16B. When the reception unit 104 detects a swipe in the lower - right direction, as shown by the broken line in FIG. 16A, it calculates a rectangular area with the starting point a as the upper - left vertex and the ending point b as the lower - right vertex as a partial area, and receives an input of class 1 for that partial area. Also, when the reception unit 104 detects a swipe in the lower - left direction, as shown by the broken line in FIG. 16B, it calculates a rectangular area with the starting point a as the upper - right vertex and the ending point b as the lower - left vertex as a partial area, and receives an input of class 2 for that partial area.

[0083] In No. 2-5 of FIG. 14, the reception unit 104 receives a class 0 input by a left flick, receives a class increment input by a right flick, receives calculation of a partial region of class 1 by a downward swipe, and receives calculation of a partial region of class 2 by an upward swipe.

[0084] The downward swipe means an operation of tracing with a fingertip from a certain location (starting point) a on the image 111 to another location (ending point) b existing downward of the starting point a as shown by the arrow in FIG. 17A. The upward swipe means an operation of tracing with a fingertip from a certain location (starting point) b on the image 111 to another location (ending point) a existing upward of the starting point b as shown by the arrow in FIG. 17B. When the reception unit 104 detects a downward swipe, it calculates a rectangle as shown by the broken line in FIG. 17A as the partial region of class 1. The rectangle is line-symmetric with respect to the swipe line, and the vertical length is equal to the length of the swipe. The horizontal length is obtained by multiplying the length of the swipe by a predetermined ratio. As the predetermined ratio, for example, the ratio of the horizontal length to the vertical length of the circumscribed rectangle of an object of class 1 having a typical shape (a two-wheeled vehicle in this embodiment) can be used. When the reception unit 104 detects an upward swipe, it calculates a rectangle as shown by the broken line in FIG. 17B as the partial region of class 2. The rectangle is line-symmetric with respect to the swipe line, and the vertical length is equal to the length of the swipe. The horizontal length is obtained by multiplying the length of the swipe by a predetermined ratio. As the predetermined ratio, the ratio of the horizontal length to the vertical length of the circumscribed rectangle of an object of class 2 having a typical shape (a four-wheeled vehicle in this embodiment) can be used.

[0085] In No. 2-6 of FIG. 14, the reception unit 104 receives a class 0 input by a pinch-in, receives a class increment input by a pinch-out, receives calculation of a partial region of class 1 by a downward-right swipe, and receives calculation of a partial region of class 2 by a downward-left swipe.

[0086] In No. 2-7 of FIG. 14, the reception unit 104 receives a Class 0 input by tapping, receives an increment input of the class by double-tapping, receives the calculation of a partial area of Class 1 by a swipe in the lower right direction, and receives the calculation of a partial area of Class 2 by a swipe in the lower left direction. These are examples, and any other combination of gestures may be used.

[0087] In the second embodiment and in No. 2-1 to 2-7 of FIG. 14, the reception unit 104 receives a Class 0 input by a first type of gesture and receives an increment input of the class by a second type of gesture. However, the reception unit 104 may be configured to receive a decrement input instead of an increment. That is, in the second embodiment and in No. 2-1 to 2-7 of FIG. 14, the reception unit 104 may be configured to receive a decrement input of the class by a second type of gesture.

[0088] In the second embodiment, the reception unit 104 receives a Class 0 input by a left click of a mouse gesture and receives an increment input of the class by a right click. However, the class and the operation may be associated one-to-one, and the inputs of Class 0, Class 1, and Class 2 may be received by specific one operation respectively. The gesture types for receiving the inputs of each class and the gesture types for receiving the calculation of the partial area of each class can be in various combinations as exemplified in the list of FIG. 18.

[0089] For example, in No. 2-11 shown in FIG. 18, the reception unit 104 receives a Class 0 input by a left click of a mouse gesture, receives a Class 1 input by a right click, receives a Class 2 input by a double click, receives the calculation of a partial area of Class 1 by a drag-and-drop in the lower right direction, and receives the calculation of a partial area of Class 2 by a drag-and-drop in the lower left direction.

[0090] Also, in the number 2-12 shown in FIG. 18, the gesture types for receiving the calculation of the partial regions of Class 1 and Class 2 are different compared to the number 2-11. In the number 2-12, the reception unit 104 receives the calculation of the partial region of Class 1 by a drag-and-drop in the upper left direction and receives the calculation of the partial region of Class 2 by a drag-and-drop in the upper right direction.

[0091] Also, in the number 2-13 shown in FIG. 18, the reception unit 104 receives the input of Class 0 by a flick of a touch gesture, receives the input of Class 1 by a tap, receives the input of Class 2 by a double tap, receives the calculation of the partial region of Class 1 by a swipe in the lower right direction, and receives the calculation of the partial region of Class 2 by a swipe in the lower left direction.

[0092] Also, in the number 2-14 shown in FIG. 18, the gesture types for receiving the inputs of Class 0, 1, and 2 are different compared to the number 2-13. In the number 2-14, the reception unit 104 receives the input of Class 0 by a flick to the left, receives the input of Class 1 by a flick upward, and receives the input of Class 2 by a flick downward.

[0093] Also, in the number 2-15 shown in FIG. 18, the reception unit 104 receives the input of Class 0 by a tap, receives the input of Class 1 by a pinch-in, receives the input of Class 2 by a pinch-out, receives the calculation of the partial region of Class 1 by a swipe downward, and receives the calculation of the partial region of Class 2 by a swipe upward.

[0094] In the above second embodiment, the number of classes is three (Class 0, Class 1, Class 2), but it can also be applied to a two-class object detection device with two classes (Class 0, Class 1) and a multi-class object detection device with four or more classes.

[0095] [Third Embodiment] A third embodiment of the present invention will be described for an object detection device configured to receive different classes of inputs with the same operation depending on whether there is an existing partial region having the position of the image on which the operation has been performed within the region. In this embodiment, the function of the reception unit 104 is different from that of the second embodiment, and the rest is basically the same as that of the second embodiment. Hereinafter, with reference to FIG. 1 which is a block diagram of the first embodiment, the function of the reception unit 104 suitable for the third embodiment will be described in detail.

[0096] The reception unit 104 has a function of visualizing the detection result DB 103 of the image and displaying it on the display device 120, and receiving an input for correction from the operator.

[0097] The reception unit 104 has a function of receiving the selection of a partial region and the input of a class for the selected partial region by an operation of the input device 130 on the image 111 displayed on the display device 120. Here, in this embodiment, the selection of a partial region means both the selection of an existing partial region and the selection of a new partial region. The reception unit 104 receives the selection of a partial region based on the position of the image 111 on which the operation of the input device 130 has been performed, and receives the input of a class based on the type of the operation. Also, in this embodiment, different class inputs are received for the same operation depending on whether there is an existing partial region having the position of the image on which the operation has been performed within the region. In this embodiment, the input device 130 is a mouse, and the types of operations are left click and right click.

[0098] The reception unit 104 accepts the selection of an existing partial region that has the position of the image 111 clicked on the left or right within the region. Also, for the selected partial region, when clicked on the left, the reception unit 104 accepts an input of class 0, and when clicked on the right, it accepts an input of a class obtained by incrementing the current class. Here, the class obtained by incrementing the current class means class 2 when the current class is class 1, class 0 when the current class is class 2, and class 1 when the current class is class 0, respectively. The reception unit 104 records the accepted class in the modified class column of the detection result DB103 for the selected partial region. The above is the same as when clicked on the left or right in the first embodiment.

[0099] Also, when there is no existing partial region that has the position of the image 111 clicked on the left or right within the region, the reception unit 104 calculates a new rectangular partial region based on the clicked position, and for that partial region, accepts an input of class 1 when clicked on the left and an input of class 2 when clicked on the right, respectively. As the method for calculating a new partial region based on the clicked position, the same method as the method for calculating the partial region of an object based on the double-clicked position described with reference to FIG. 15 can be used.

[0100] FIG. 19 is a flowchart showing the details of the operation of the reception unit 104. In FIG. 19, steps S211 to S224, S226 to S232 are substantially the same as steps S211 to S224, S226 to S232 in FIG. 11. The difference between FIG. 19 and FIG. 11 is that when it is determined in step S216 that there is no existing partial region that includes the clicked position, the process proceeds to step S226. Note that step S226 in FIG. 11 determines the direction (type) of drag and drop, while step S226 in FIG. 19 determines the type of click. Hereinafter, with reference to FIG. 19, the operation of the reception unit 104 will be described in detail centering on the differences from FIG. 11.

[0101] The reception unit 104 detects a left click or a right click on the image in the image window 124 (YES in step S214). At this time, if the reception unit 104 determines that the partial region including the click position does not exist in the image (NO in step S216), it determines whether the click type is a left click or a right click in order to judge the input of the class (step S226).

[0102] If it is a left click, the reception unit 104 calculates a new rectangular partial region and accepts the input of class 1 (step S227). Next, the reception unit 104 records the pair of the calculated partial region and class 1 as one detection result in the detection result DB103 (step S228). Next, the reception unit 104 updates the image window 124 of the reception screen 126 (step S229). Then, the reception unit 104 returns to the process of step S213.

[0103] If it is a right click (step S226), the reception unit 104 calculates a new rectangular partial region and accepts the input of class 2 (step S230). Next, the reception unit 104 records the pair of the calculated partial region and class 2 as one detection result in the detection result DB103 (step S231). Next, the reception unit 104 updates the image window 124 of the reception screen 126 (step S232). Then, the reception unit 104 returns to the process of step S213.

[0104] Thus, according to this embodiment, high-quality teacher data can be efficiently created. The reason is that the reception unit 104 displays the input image on the display device 120 with a display that emphasizes the partial region of the detected object, and accepts the selection of the partial region and the input of the class for the selected partial region by a click which is one operation of the input device 130. Also, it is because the reception unit 104 detects one operation such as a left click or a right click at a position that does not overlap with the existing partial region, and accepts the calculation of a new partial region and the input of class 1 or class 2.

[0105] [Modification Example of the Third Embodiment] In the above-described third embodiment, two types of gestures, i.e., left click and right click, were used. However, combinations of other types of mouse gestures or combinations of touch gestures can also be used.

[0106] [Fourth Embodiment] In the fourth embodiment of the present invention, an object detection device will be described in which correct class input for a misrecognized partial region and selection of the correct partial region of an object can be performed by operating an input device on an input image. In this embodiment, compared with the second embodiment, the function of the reception unit 104 is different, and the rest is basically the same as the second embodiment. Hereinafter, with reference to FIG. 1 which is a block diagram of the first embodiment, the function of the reception unit 104 suitable for the fourth embodiment will be described in detail.

[0107] The reception unit 104 has a function of visualizing the detection result DB 103 of the image and displaying it on the display device 120, and receiving input of correction from the operator.

[0108] FIG. 20 shows an example of a reception screen 127 displayed by the reception unit 104 on the display device 120. The reception screen 127 in this example differs only in the position of the rectangular frame indicating the outer periphery of the partial region of the four-wheel vehicle in the image 111 displayed in the image window 124, as compared with the reception screen 121 shown in FIG. 4. The reason for the difference in the position of the rectangular frame is that when the detection unit 101 detected an object from the image 111 using the dictionary 102, it misdetected a water puddle on the road or the like as part of the partial region of the object. In this example, although the partial region is misdetected, the class is correctly determined. However, there are also cases where both the partial region and the class are misdetected. In order to improve such misdetection, it is necessary to create teacher data which is an input / output pair of the misdetected partial region in the image 111 and its class 0, and teacher data which is an input / output pair of the correct partial region of the four-wheel vehicle in the image 111 and its class 2, respectively, and learn the dictionary 102.

[0109] The reception unit 104 has a function of receiving the selection of a partial region and the input of a class for the selected partial region by an operation of the input device 130 on the image 111 displayed on the display device 120. Here, in the present embodiment, the selection of a partial region means both the selection of an existing partial region and the selection of a new partial region. The reception unit 104 receives the selection of a partial region based on the position of the image 111 where the operation of the input device 130 is performed, and receives the input of a class based on the type of the operation. Also, in the present embodiment, when a selection operation of a new partial region that partially overlaps with an existing partial region is performed, the reception unit 104 receives the input of class 0 for the existing partial region and receives the input of a class corresponding to the type of the operation for the new partial region. In the present embodiment, the input device 130 is a mouse, and the types of the operations are left click, right click, drag and drop in the lower right direction, and drag and drop in the lower left direction.

[0110] The reception unit 104 receives the selection of a partial region having the position of the image 111 clicked with the left or right button within the region. Also, the reception unit 104 receives the input of class 0 when the left button is clicked and receives the input of a class obtained by incrementing the current class when the right button is clicked for the selected partial region. Here, the class obtained by incrementing the current class means class 2 when the current class is class 1, class 0 when the current class is class 2, and class 1 when the current class is class 0, respectively. The reception unit 104 records the received class for the selected partial region in the modified class column of the detection result DB 103. The above is the same as when the left or right button is clicked in the first embodiment.

[0111] When the reception unit 104 detects a drag-and-drop operation in the lower-right direction, it calculates a rectangular area with the starting point a as the upper-left vertex and the ending point b as the lower-right vertex as a partial area as shown by the dashed line in FIG. 10A, and accepts an input of class 1 for that partial area. When the reception unit 104 detects a drag-and-drop operation in the lower-left direction, it calculates a rectangular area with the starting point a as the upper-right vertex and the ending point b as the lower-left vertex as a partial area as shown by the dashed line in FIG. 10B, and accepts an input of class 2 for that partial area. Further, the reception unit 104 records the newly calculated partial area and the information of its class as one detection result in the detection result DB 103. The above is the same as when a drag-and-drop operation in the lower-right direction and a drag-and-drop operation in the lower-left direction are performed in the second embodiment.

[0112] Furthermore, the reception unit 104 has a function of accepting an input of class 0 for an existing partial area that partially overlaps with the newly calculated partial area. For example, as shown in FIG. 21, when a drag-and-drop operation in the lower-left direction (or lower-right direction) is performed on an image 111 in which a frame 125 highlighting a partial area of class 2 is displayed, and the reception unit 104 calculates a partial area 125a that partially overlaps with the frame 125, it accepts an input of class 0 for the existing partial area related to the frame 125.

[0113] FIG. 22 is a flowchart showing the details of the operation of the reception unit 104. In FIG. 22, steps S311 to S327 and S330 are the same as steps S211 to S327 and S230 in FIG. 11. Hereinafter, with reference to FIG. 22, the operation of the reception unit 104 will be described in detail centering on the differences from FIG. 11.

[0114] When the reception unit 104 receives an input of class 1 for the newly calculated partial region (step S327), it determines whether or not there is an existing partial region that partially overlaps with this new partial region on the image 111 (step S341). This determination is made by investigating whether or not there is a detection result of an existing partial region that partially overlaps with the new partial region in the detection result DB103 related to the image 111. If such an existing partial region exists, the reception unit 104 receives an input of class 0 for the partial region (step S342) and proceeds to step S343. If such an existing partial region does not exist, the reception unit 104 skips the process of step S342 and proceeds to step S343. In step S343, the reception unit 104 updates the detection result DB103 according to the reception results of steps S327 and S342. That is, based on the reception result of step S327, the reception unit 104 records a pair of the newly calculated partial region and class 1 as one detection result in the detection result DB103. Also, based on the reception result of step S342, the reception unit 104 records class 0 in the correction class of the existing partial region. Next, the reception unit 104 updates the image 111 on the image window 124 according to the update of the detection result DB103 in step S343 (step S344). Then, the reception unit 104 returns to the process of step S313.

[0115] Also, when the reception unit 104 receives an input of class 2 for the newly calculated partial region (step S330), it determines whether there is an existing partial region that partially overlaps with this new partial region on the image 111 (step S345). If the reception unit 104 determines that such an existing partial region exists, it receives an input of class 0 for the partial region (step S346) and proceeds to step S347. If such an existing partial region does not exist, the reception unit 104 skips the process of step S346 and proceeds to step S347. In step S347, the reception unit 104 updates the detection result DB103 according to the reception results of steps S330 and S346. That is, based on the reception result of step S330, the reception unit 104 records the pair of the newly calculated partial region and class 2 as one detection result in the detection result DB103. Also, based on the reception result of step S346, the reception unit 104 records class 0 in the correction class of the existing partial region. Next, according to the update of the detection result DB103 in step S347, the reception unit 104 updates the image 111 on the image window 124 (step S348). Then the reception unit 104 returns to the process of step S313.

[0116] Figures 23A and 23B are explanatory diagrams of step S347 by the reception unit 104. Figure 23A shows the detection result DB103 generated by the detection unit 101, and Figure 23B shows the updated detection result DB103 by the reception unit 104. In Figures 23A and 23B, a rectangular partial region with the upper left vertex at (x7, y7) and the lower right vertex at (x8, y8) is newly added as class 2. Also, the class of the existing partial region of the rectangle with the upper left vertex at (x1, y1) and the lower right vertex at (x2, y2) is corrected from class 2 to class 0. In the case of step S343 by the reception unit 104, the difference is that the correction class of the newly added partial region becomes class 1.

[0117] FIG. 24A and FIG. 24B are explanatory diagrams of step S348 by the reception unit 104. FIG. 24A shows the image window 124 before update, and FIG. 24B shows the image window 124 after update. In the image window 124 after update, a rectangular frame 125 with a solid line that emphasizes the newly added partial region is drawn, and the frame that was displayed in the image window 124 before update and partially overlapped with it is made invisible. In the case of step S344 by the reception unit 104, the difference is that a rectangular frame that emphasizes the newly added partial region is drawn with a broken line in the image window 124 after update.

[0118] As described above, according to this embodiment, high-quality teacher data can be efficiently created. The reason is that the reception unit 104 displays the input image on the display device 120 with a display that emphasizes the partial region of the detected object, and accepts the selection of the partial region and the input of the class for the selected partial region by a click, which is one operation of the input device 130. Also, it is because the reception unit 104 detects one operation of drag-and-drop in the lower right direction or the lower left direction on the image, and accepts the calculation of a new partial region and the input of class 1 or class 2. Further, it is because the reception unit 104 accepts the input of class 0 for an existing partial region that partially overlaps with the newly calculated partial region.

[0119] [Modification Example of the Fourth Embodiment] For the configuration of the above-described fourth embodiment, the same modifications as those described in the modification example of the second embodiment can be made.

[0120] [Fifth Embodiment] The fifth embodiment of the present invention will be described. Referring to FIG. 25, the object detection device 1 according to this embodiment has a function of detecting an object from the input image 2. The object detection device 1 has a dictionary 10, and as main functional units, has a detection unit 11, a reception unit 12, a generation unit 13, and a learning unit 14. Further, the object detection device 1 is connected to a display device 3 and an input device 4.

[0121] The detection unit 11 has a function of detecting an object from the input image 2 using the dictionary 10. The reception unit 12 displays the input image 2 on the display device 3 with a display that emphasizes the partial region of the object detected by the detection unit 11, and has a function of receiving the selection of the partial region and the input of the class for the selected partial region by one operation of the input device 4. The generation unit 13 has a function of generating teacher data from the image of the selected partial region and the input class. The learning unit 14 has a function of learning the dictionary 10 with the teacher data generated by the generation unit 13 and upgrading the dictionary 10.

[0122] Next, the operation of the object detection device 1 according to the present embodiment will be described. The detection unit 11 of the object detection device 1 detects an object from the input image 2 using the dictionary 10 and notifies the reception unit 12 of the detection result. The reception unit 12 displays the input image 2 on the display device 3 with a display that emphasizes the partial region of the object detected by the detection unit 11. Then, the reception unit 12 receives the selection of the partial region and the input of the class for the selected partial region by one operation of the input device 4, and notifies the result to the generation unit 13. The generation unit 13 generates teacher data from the image of the selected partial region and the input class and notifies the learning unit 14. The learning unit 14 learns the dictionary 10 with the teacher data generated by the generation unit 13 and upgrades the dictionary 10.

[0123] As described above, according to the present embodiment, high-quality teacher data can be efficiently created. The reason is that the reception unit 12 displays the input image 2 on the display device 3 with a display that emphasizes the partial region of the detected object. Further, the reception unit 12 receives the selection of the partial region and the input of the class for the selected partial region by a click which is one operation of the input device 4.

[0124] While the present embodiment is based on the configuration as described above, various additional changes as follows are possible.

[0125] The reception unit 12 may be configured to receive the selection of a partial region based on the position of the input image 2 for which one operation has been performed.

[0126] Further, the reception unit 12 may be configured to receive the selection of a partial region having the position of the input image 2 for which one operation has been performed within the region, and if there is no partial region having the position of the input image 2 for which one operation has been performed within the region, calculate the partial region of the object based on the position of the input image 2 for which one operation has been performed.

[0127] Further, the reception unit 12 may be configured to receive the input of a class according to the type of one operation.

[0128] Further, the reception unit 12 may be configured to receive the input of the first class in advance if the type of one operation is the first type, and receive the input of the class obtained by incrementing or decrementing the current class of the partial region if the type is the second type.

[0129] Further, the reception unit 12 may be configured to receive the input of different classes depending on whether or not there is a partial region having the position of the input image 2 for which one operation has been performed within the region.

[0130] Further, the reception unit 12 may be configured to receive the input of the first class when there is a partial region having the position of the input image 2 for which one operation has been performed within the region, and receive the input of the second class when there is no such partial region.

[0131] Further, the reception unit 12 may be configured to receive the input of the first class for the existing partial region when there is a partial region having the position of the input image 2 for which one operation has been performed within the region, calculate the partial region of the object based on the position of the input image for which one operation has been performed, and receive the input of the second class for this calculated partial region.

[0132] Further, the reception unit 12 may be configured to determine the second class according to the type of one operation.

[0133] Also, the input device 4 may be configured such that one operation thereof is one mouse gesture. Alternatively, the input device 4 may be configured such that one operation thereof is one touch gesture.

[0134] Also, when the partial region of the detected object has the same position and the same pre-input class as the partial region where the class was input in the past, the reception unit 12 may be configured to turn off the display that emphasizes the partial region of the detected object. [Sixth Embodiment] The object detection device 300 according to the sixth embodiment of the present invention will be described with reference to FIG. 27. The object detection device 300 includes a detection unit 301, a reception unit 302, a generation unit 303, and a learning unit 304.

[0135] The detection unit 301 detects an object from the input image using a dictionary.

[0136] The reception unit 302 displays the input image on the display device with a display that emphasizes the partial region of the detected object, and receives the selection of the partial region and the input of the class for the selected partial region by one operation of the input device.

[0137] The generation unit 303 generates teacher data from the image of the selected partial region and the input class.

[0138] The learning unit 304 learns the dictionary using the teacher data.

[0139] According to the present embodiment, high-quality teacher data can be efficiently created. The reason is that the reception unit 302 displays the input image on the display device with a display that emphasizes the partial region of the detected object, and receives the selection of the partial region and the input of the class for the selected partial region by one operation of the input device.

[0140] The present invention has been described by taking the above-described embodiments as exemplary examples. However, the present invention is not limited to the above-described embodiments. That is, within the scope of the present invention, various aspects understandable by those skilled in the art can be applied.

[0141] This application claims priority based on Japanese Patent Application No. 2015-055926 filed on March 19, 2015, and incorporates the entire disclosure thereof herein.

Industrial Applicability

[0142] The present invention can be used in an object detection device that detects objects such as people and vehicles from an image on a road obtained by imaging with a surveillance camera using a dictionary.

Explanation of Signs

[0143] 1…Object detection device 2…Input image 3…Display device 4…Input device 10…Dictionary 11…Detection unit 12…Reception unit 13…Generation unit 14…Learning unit 100…Object detection device 101…Detection unit 102…Dictionary 103…Detection result DB 104…Reception unit 105…Generation unit 106…Teacher data memory 107…Learning unit 111…Image 112, 113, 114, 115…Search window 120…Display device 121…Reception screen 122…Graph 123…Slider 123a…Slider bar 124…Image window 125…Rectangular frame 125a…Partial region 126, 127... Reception screen 130... Input device 131... Object 132... Bounding rectangle

Claims

1. Detection means for detecting an object from an input image, Control means for controlling to highlight a first region corresponding to the object and a second region corresponding to a position selected by a user in the input image, Comprising: The first region is associated with a first class, The second region is associated with a second class different from the first class, An object detection device.

2. In the process of highlighting, the control means controls to highlight the first region and the second region differently, The object detection device according to claim 1.

3. Further comprising generation means for generating teacher data, The control means receives an input of the second class for the second region, The generation means generates the teacher data based on at least the second region and the second class, The object detection device according to claim 1 or 2.

4. The object detection device according to any one of claims 1 to 3, wherein the second class indicates an object of a type different from the first class.

5. Detect an object from an input image, Control to highlight a first region corresponding to the object and a second region corresponding to a position selected by a user in the input image, The first region is associated with a first class, The second region is associated with a second class different from the first class, An object detection method.

6. Control to highlight the first region and the second region differently, The object detection method according to claim 5.

7. Receive an input of the second class for the second region, Generate teacher data based on at least the second region and the second class The object detection method according to claim 5 or 6.

8. The object detection method according to any one of claims 5 to 7, wherein the second class indicates an object of a type different from the first class.

9. A process of detecting an object from an input image, A process of highlighting a first region corresponding to the object and a second region corresponding to a position selected by a user in the input image, To be executed by a computer, The first region is associated with a first class, The second region is associated with a second class different from the first class, A program.

10. In the process of causing the highlighting, the first region and the second region are highlighted differently. The program according to claim 9. **Claim 11** A process of receiving an input of the second class for the second region, A process of generating teacher data based on at least the second region and the second class, The program according to claim 9 or 10, further causing a computer to execute. **Claim 12** The program according to any one of claims 9 to 11, wherein the second class indicates a different type of object from the first class.

Citation Information

Patent Citations

  • Sorting device and sorting method

    JP2006012069A

  • Video display apparatus

    JP2007280325A

  • Image processing device, image processing method

    JP2012088787A

  • Object detection device, object detection method and program

    JP2024097936A

  • Classification assisting apparatus, classifying apparatus, and program

    JP2003317082A