Image recognition device, image recognition system, image recognition program, and image recognition method
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2024-08-28
- Publication Date
- 2026-08-03
AI Technical Summary
【0008】 本発明によれば、対象物体の周辺に複数の物体領域が重なって検出された際に、検出された各物体領域に対して、正検出あるいは誤検出であることを判定することができる。
Smart Images

Figure 0007899267000001 
Figure 0007899267000002 
Figure 0007899267000003
Abstract
Description
Technical Field
[0006] , , , , , , , ,
[0005] ,
[0001] The present invention relates to an image recognition device that detects regions such as objects from images. [[ID=б]]
Background Art
[0002] As a method of image recognition for detecting regions such as target objects from images, in order to improve detection accuracy, a method of performing detection step by step is known. In step-by-step detection, first, the region of the entire object is detected from the image, and then, inside the region of the entire object, a process of further detecting a part of the object is performed.
[0003] However, if the detection result when detecting the region of the entire object from the image is incorrect, the detection results in subsequent steps will also be incorrect. Therefore, there is a need for a technique to determine whether the detected region has been correctly detected (correct detection) or whether the detected region has been erroneously detected (false detection).
[0004] In Patent Document 1, the target object to be detected is the hand of a wearer wearing a head-mounted display (hereinafter referred to as HMD). Further, by using at least any one of the movement of the device worn on the hand, the distance between the HMD and the hand, the moving direction of the hand, the extending direction of the hand, and the presence or absence of the device, when a hand other than the wearer is detected as the target object, a technique for determining that it is a false detection is disclosed.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, the technology disclosed in Patent Document 1 does not consider a method to reduce the possibility of incorrect determination of whether an object is correctly detected or incorrectly detected, even when multiple regions detected as containing an object overlap. In other words, when multiple regions detected as containing an object overlap, there is a risk of incorrectly determining that a region that is actually incorrectly detected is correctly detected. Therefore, the present invention aims to provide an image recognition device that reduces the possibility of incorrect determination of whether an object is correctly detected or incorrectly detected, even when multiple regions detected as containing an object overlap. [Means for solving the problem]
[0007] To achieve the above objective, the image recognition device of the present invention includes: a region detection means for detecting an object region containing an object presumed to be a hand from an image; a position detection means for detecting, for each object region, a detection target point that is presumed to be the position of a hand, including the joint points of the fingers and wrist, and the position of the detection target point on the object, when a plurality of the object regions are detected from a single image by the region detection means; and a determination means for determining, for each object region, whether or not the object region is a false detection, wherein the determination means determines the first object region to be a false detection if, in a first object region and a second object region different from the first object region detected from a single image by the region detection means, a specific detection target point is located within the first object region, among the detection target points detected from the second object region by the position detection means. [Effects of the Invention]
[0008] According to the present invention, when multiple object regions are detected overlapping around a target object, it is possible to determine whether each detected object region is a correct detection or a false detection. [Brief explanation of the drawing]
[0009] [Figure 1] This is a diagram illustrating the internal configuration of an image recognition device according to the first embodiment of the present invention. [Figure 2] This figure schematically represents an example of image data processed by the image recognition device according to the first embodiment of the present invention. [Figure 3] This diagram illustrates a flowchart showing the operation of an image recognition device according to the first embodiment of the present invention. [Figure 4] This figure schematically represents an example of image data processed by an image recognition device according to a second embodiment of the present invention. [Figure 5] This diagram illustrates a flowchart showing the operation of an image recognition device according to a second embodiment of the present invention. [Modes for carrying out the invention]
[0010] The embodiments will be described below with reference to the drawings. The same or equivalent components, members, and processes shown in each drawing will be denoted by the same reference numerals, and redundant explanations will be omitted as appropriate. Furthermore, some components, members, and processes will be omitted in each drawing. The following embodiments are not limiting to the present invention, and not all combinations of features described in these embodiments are essential to the solutions of the present invention. The configuration of the embodiments may be modified or changed as appropriate depending on the specifications of the device to which the present invention is applied and various conditions (operating conditions, operating environment, etc.). In the following embodiments, the same components will be denoted by the same reference numerals.
[0011] (First embodiment) <Internal Configuration of Image Recognition Device> Figure 1 is a block diagram illustrating a head-mounted display (HMD) 100, which is an image recognition device according to this embodiment. In this embodiment, an HMD is assumed as an example of a device that constitutes an image recognition device, but the present invention is not limited to these embodiments. For example, it may be a smartphone or tablet terminal equipped with a camera, or other devices such as a personal computer (PC) or digital camera.
[0012] The HMD100 is a head-mounted display device (electronic device) that can be worn on the user's head. The HMD100 is equipped with a camera to capture images of the area directly in front of the user and a display to show the images to the user. The HMD100's display shows a composite image, which is a combination of the image captured by the HMD100 of the area directly in front of the user and content such as computer graphics (CG) that is shaped according to the posture of the HMD100. This allows the user to experience virtual reality or mixed reality with their own eyes.
[0013] Referring to Figure 1, the internal configuration of the HMD100 will be explained. The HMD100 has a control unit 101, ROM 102, RAM 103, and imaging unit 104 connected to a system bus 105. The control unit 101, ROM 102, RAM 103, imaging unit 104, and system bus 105 are the hardware resources that constitute the HMD100. Note that each component of the HMD100 may be an image recognition system composed of individual hardware.
[0014] The control unit 101 controls various parts of the HMD 100 according to the input signals and the program described later. The control unit 101 has at least one CPU that executes the program stored in the ROM 102, and at least one other circuit. Alternatively, instead of the control unit 101 controlling the entire device, multiple hardware components may share the processing to control the entire device.
[0015] The ROM 102 is an electrically erasable and recordable non-volatile memory, which stores programs and the like executed by the control unit 101. When the HMD 100 is powered on, the control unit 101 reads the program from the ROM 102 and starts controlling the HMD 100. The ROM 102 is composed of, for example, a flash memory or the like.
[0016] The RAM 103 is used as a working area for the programs executed by the control unit 101. The RAM 103 is composed of, for example, a volatile memory (DRAM) using semiconductor elements or the like.
[0017] Here, the control unit 101, the ROM 102, and the RAM 103 have been described as individual hardware resources, but these functions may also be integrated and realized on a single LSI.
[0018] The imaging unit 104 consists of a stereo camera, and captures a color image or a monochrome image of a scene with the two mounted left and right cameras, and outputs a video signal to the system bus 105. This video signal is subjected to various image processes by the control unit 101 and stored in the RAM 103 as left and right image data. The imaging unit 104 is composed of, for example, an optical lens unit and an optical system for controlling an aperture, zoom, focus, etc., and an imaging element for converting the light (video) introduced through the optical lens unit into an electrical video signal. As the imaging element, generally, a CMOS imaging element (CMOS image sensor) using CMOS or a CCD imaging element (CCD image sensor) using CCD is used.
[0019] <Flow for determining whether each detected object region is a positive detection or a false detection> Next, a flowchart showing the operation of the HMD 100 according to the first embodiment will be described with reference to FIGS. 2(a), (b), (c), and FIG. 3. The processes shown in this flowchart are realized by the control unit 101 of the HMD 100 controlling each part of the HMD 100 according to an input signal or a program. That is, this process is realized by expanding the program recorded in the ROM 102 into the RAM 103 and executed by the control unit 101. The process shown in the flowchart of FIG. 3 is executed from the timing when the user starts an application on the HMD 100, and is executed every time the imaging unit 104 acquires an image (every time shooting is performed). The application is, for example, an application selected by the user on the home screen (home space) after starting the HMD 100, and examples of the application include an application that allows the user to interfere with a virtual object using a hand gesture. Note that the timing when this flowchart is executed is not limited to the timing when the user starts an application on the HMD 100. For example, it may be the timing when the user starts the HMD 100 or the timing when a virtual object is displayed in the mixed reality space.
[0020] Here, FIGS. 2(a), (b), and (c) are diagrams schematically showing an example of image data processed by the image recognition device of the present embodiment. In FIG. 2(a), the image 200 is an example of image data generated by the control unit 101 from the video signal acquired by the imaging unit 104, and the image 200 includes the hand 201 of the HMD wearer. The descriptions of FIGS. 2(b) and 2(c) will be described later.
[0021] In step S301, the control unit 101 acquires the video signal output from the imaging unit 104 to the system bus 105, generates image data, and proceeds to step S302. For example, the control unit 101 stores the image 200 in FIG. 2(a) in the RAM 103.
[0022] In step S302, the control unit 101 performs a process to detect a human hand, which is the target object to be detected, on the image data generated in step S301, and proceeds to step S303. The process for detecting a human hand is a region detection process that detects a region of a predetermined shape (hereinafter referred to as the object region) that is estimated to contain a hand. For example, the predetermined shape can be a rectangle. Furthermore, the smallest region of the predetermined shape that is estimated to contain a hand is detected. Note that, depending on the image, the rectangular region estimated to contain a hand may not be the smallest region. In the process for detecting a human hand, depending on how the subject etc. is depicted in the image data, only one object region may be detected, multiple object regions may be detected, or none may be detected. Note that the process for detecting a human hand may be implemented, for example, using a trained deep learning model or a rule-based algorithm. In this embodiment, the predetermined shape of the object region is a rectangle, but the predetermined shape is not necessarily limited to a rectangle; if all object regions have the same shape, they may be, for example, a polygon, a circle, or an ellipse.
[0023] Figure 2(b) is an example of the result of performing the processing in step S302 on image 200 in this embodiment, showing the detected object regions 211 and 221 superimposed on image 200. In this embodiment, the upper left of image 200 is the origin, the vertical downward direction is the v-coordinate axis, and the horizontal right direction is the u-coordinate axis. The v-coordinate 213 at the upper end of object region 211 is vt1, the v-coordinate 213 at the lower end is vb1, and the v-coordinate 223 at the upper end of object region 221 is vt2, with vt1=100, vb1=200, and vt2=150.
[0024] In step S303, the control unit 101 performs a position detection process to detect the position of the target point for all object regions detected in step S302, and proceeds to step S304. For example, the control unit 101 performs a position detection process to detect the positions of the fingertips of each finger and the joints of the hand, including the wrist (joint points), and proceeds to step S304. This joint point detection process may be implemented, for example, using a trained deep learning model or a rule-based algorithm. Here, for example, for each finger, at least 21 joint points are obtained, including four points from the fingertip to the base of the finger (DIP joint, PIP joint, MP joint, and IP joint) and one point at the wrist.
[0025] Figure 2(c) is an example of the result of performing the processing in step S303 on image 200 in the first embodiment, and is a composite image of the object region to be processed and the detected joint point positions on image 200. The black circles represent the joint point positions detected for the object region 211, and of the black circles, 212 represents the joint point position of the wrist. The white circles represent the joint point positions detected for the object region 221, and of the white circles, 222 represents the joint point position of the wrist.
[0026] The process from step S304 to step S311 is a loop process for determining object regions that have been falsely detected. Here, if a detected region is correctly detected, it is considered a positive detection; if a detected region is incorrectly detected, it is considered a false detection.
[0027] In step S304, the control unit 101 starts a loop process (hereinafter referred to as Loop A) to apply the processes from steps S305 to S310 to the object region detected in step S302. If no object region is detected in step S302, Loop A is not applied, and the process proceeds to step S312.
[0028] In step S305, the control unit 101 starts a loop process (hereinafter referred to as loop B) to apply the processes from steps S306 to S309 to the object region detected in step S302.
[0029] In step S306, the control unit 101 determines whether the object region selected in step S304 and the object region selected in step S305 are the same object region. If the control unit 101 determines that the object region selected in step S304 and the object region selected in step S305 are different object regions, it proceeds to step S307. If the control unit 101 determines that the object region selected in step S304 and the object region selected in step S305 are the same object region, it proceeds to step S310.
[0030] For example, if object region 211 is selected in step S304 and object region 221 is selected in step S305, the control unit 101 proceeds to step S307. Also, if object region 211 is selected in steps S304 and S305, or if object region 221 is selected in steps S304 and S305, the control unit 101 proceeds to step S310.
[0031] In step S307, the control unit 101 determines whether the wrist joint point position detected in S303 is within the object region selected in step S304. If the wrist joint point position detected in step S303 is not within the object region selected in step S304, the control unit 101 proceeds to step S310. If the wrist joint point position detected in step S303 is within the object region selected in step S304, the control unit 101 proceeds to step S308. In this embodiment, the wrist joint point position is used as the object to be determined, but the object to be determined is not necessarily limited to the wrist joint point position. Also, the object to be determined is not necessarily limited to one point, but may be multiple points. For example, it may be joint point positions including the fingertips, or multiple joint point positions including the wrist joint point position and joint points including the fingertips.
[0032] In this embodiment, referring to Figure 2(c), we assume that object region 211 is selected in step S304 and object region 221 is selected in step S305. In this case, since the wrist joint point position 212 detected in S303 is within the area of object region 221, the control unit 101 proceeds to step S308. Also referring to Figure 2(c), we assume that object region 221 is selected in step S304 and object region 211 is selected in step S305. In this case, since the wrist joint point position 222 detected in S303 is not within the area of object region 211, the control unit 101 proceeds to step S310.
[0033] In step S308, the control unit 101 determines whether the object region selected in step S305 is below a predetermined position (threshold) relative to the object region selected in step S304. If the object region selected in step S305 is below the predetermined position relative to the object region selected in step S304, the control unit 101 proceeds to step S309. If the object region selected in step S305 is not below the predetermined position relative to the object region selected in step S304, the control unit 101 proceeds to step S310. In this embodiment, the predetermined position (threshold) is set to the v-coordinate of the object region selected in step S304, which is 1 / 4 of the vertical width of the object region selected in step S304 relative to the v-coordinate of the upper end of the object region selected in step S304. The v-coordinate of the predetermined position is then compared with the v-coordinate of the upper end of the object region selected in step S305.
[0034] For example, if object region 211 is selected in step S304 and object region 221 is selected in step S305, then, as described above, the v-coordinate of the upper end of object region 211 is vt1=100. Also, the v-coordinate of the lower end of object region 211 is vb1=200, and the v-coordinate of the upper end of object region 221 is vt2=150. The vertical width of object region 211 selected in S304 is vb1-vt1=100 pixels, and the v-coordinate thv representing a predetermined position 1 / 4 of the vertical width of object region 211 is thv=vt1+(vb1-vt1) / 4=100+100 / 4=125. Since vt2 is below thv, the control unit 101 determines that object region 221 is below the predetermined position relative to object region 211 and proceeds to step S309.
[0035] The predetermined position (threshold) used in step S308 may be set based on the vertical or horizontal width of the object region selected in step S304. Alternatively, the predetermined position (threshold) used in step S308 may be set based on the vertical or horizontal width of the object region selected in step S305. Furthermore, the predetermined position (threshold) used in step S308 may be a predetermined position below the object region selected in step S304.
[0036] Note that the processing in step S308 may be omitted. That is, in step S307, if the wrist joint position detected in step S303 is within the object region selected in step S304, the control unit 101 may proceed to step S309. Here, as mentioned above, if the wrist joint position detected in step S303 is not within the object region selected in step S304, the process proceeds to step S310.
[0037] In step S309, the control unit 101 determines that the object region selected in step S305 is a false detection and proceeds to step S310. That is, the control unit 101 determines that the object region selected in step S305 is a region that does not accurately include a hand.
[0038] For example, if object region 221 is selected in step S305, the control unit 101 determines that object region 221 is a false detection and proceeds to step S310. The predetermined position used in the determination in step S308 may be, for example, the position of any of the detected object regions, or a position based on the length of one side of any of the detected object regions.
[0039] In step S310, the control unit 101 determines whether to terminate loop B. That is, if all object regions are selected in step S305, loop B is terminated and the process proceeds to step S311; otherwise, the process returns to step S305 and loop B continues.
[0040] For example, if the control unit 101 first selects object region 211 in step S305 and proceeds to step S310, then in step S310 it determines that object region 221 has not yet been selected and returns to step S305, continuing loop B. If object region 211 and object region 221 are selected in step S305, then it is assumed that all object regions have been selected and the process proceeds to step S311.
[0041] In step S311, the control unit 101 determines whether to terminate loop A. That is, if all object regions are selected in step S304, loop A is terminated and the process proceeds to step S312; otherwise, the process returns to step S304 and loop A continues.
[0042] For example, if the control unit 101 first selects object region 211 in step S304 and proceeds to step S311, then in step S311 it determines that object region 221 has not yet been selected and returns to S304, continuing loop A. If object region 211 and object region 221 are selected in step S304, then it is assumed that all object regions have been selected and proceeds to step S312.
[0043] In step S312, the control unit 101 determines that all object regions detected in step S302, excluding those determined to be false detections in S309, are correct detections, and proceeds to step S313.
[0044] In step S313, the control unit 101 determines whether to terminate the process shown in the flowchart of Figure 3. For example, if the user inputs a termination command via an operating device (not shown) or if the control unit 101 is unable to acquire image data, the control unit 101 determines to terminate the process shown in Figure 3. On the other hand, if the control unit 101 determines not to terminate in step S313, it proceeds to step S301.
[0045] As described above, in this embodiment, the control unit 101 detects an object region from the image data that includes the target object, a human hand, and detects the fingertips of each finger and the joint points including the wrist within that object region. If the wrist joint point among the joint points detected within one object region is located within the other object region, and the other object region is located below a predetermined position relative to the first object region, the control unit 101 determines that the other object region is a false detection. This makes it possible to determine whether each detected object region is a correct or false detection when multiple object regions are detected overlapping around the target object.
[0046] (Second Embodiment) Since the configuration of the image recognition device according to this embodiment is the same as that of the first embodiment, a description will be omitted.
[0047] The following flowchart illustrates the operation of the HMD 100 according to the first embodiment, with reference to Figures 4 and 5. The process shown in this flowchart is implemented by the control unit 101 of the HMD 100 controlling various parts of the HMD 100 according to input signals and programs. Specifically, this process is implemented by the control unit 101 executing a program stored in the ROM 102 after it has been loaded into the RAM 103. The process shown in the flowchart of Figure 5 is executed from the moment the user launches an application on the HMD 100, and is executed each time the imaging unit 202 acquires an image (each time a picture is taken). An application is, for example, an application selected by the user on the home screen (home space) after launching the HMD 100, and includes applications that allow the user to interact with virtual objects using hand gestures. The timing of the execution of this flowchart is not limited to the moment the user launches an application on the HMD 100. For example, it may be executed at the moment the user launches the HMD 100, or at the moment a virtual object is displayed in the mixed reality space.
[0048] The processes in steps S301 and S302 are the same as those in the flow chart in Figure 3, so their explanation will be omitted.
[0049] In step S503, the control unit 101 performs a process to detect the positions of each joint point, including the fingertips and wrists, and the sub-wrist point for all object regions detected in step S302, and then proceeds to step S304. Here, the sub-wrist point is defined as any point located between the wrist joint point and the elbow joint, and within the object region being processed. This process of detecting the positions of each joint point and the sub-wrist point may be implemented, for example, using a trained deep learning model, or it may be implemented using a rule-based algorithm.
[0050] Figure 4 is an example of the result of performing the processing in step S503 on image 200 in this embodiment, showing the object region to be processed and the positions of each detected joint point and wrist sub-point superimposed on image 200. The object region 211, the wrist joint point 212 within object region 211, object region 221, and the wrist joint point 222 within object region 221 are the same as in Figure 2(c), so their explanation is omitted. Also, in Figure 4, as in Figure 2(c), the black circles are the joint points detected for object region 211, and the white circles are the joint points detected for object region 221. Point 413 among the black circles represents the position of the wrist sub-point. Point 423 among the white circles represents the position of the wrist sub-point. The angle 414 is the angle of the arm, represented by the angle between the straight line connecting the wrist joint point position 212 and the wrist sub-point position 413, and the straight line drawn vertically from the wrist sub-point position 413. Angle 424 is the angle of the arm, represented by the angle between the line connecting the wrist joint point position 222 and the wrist base point position 423, and the line drawn perpendicularly from the wrist base point position 423. The coordinates of the wrist joint point position 212 are (u12, v12), the coordinates of the wrist base point position 413 are (u13, v13), the coordinates of the wrist joint point position 222 are (u22, v22), and the coordinates of the wrist base point position 423 are (u23, v23). Furthermore, let u12=100, u13=80, u22=50, u23=30, v12=160, v13=180, v22=210, and v23=230.
[0051] The processes in steps S304, S305, and S306 are the same as the flow shown in Figure 3, so their explanation is omitted.
[0052] In step S507, the control unit 101 determines whether the positions of the wrist joint point and the forearm of the wrist detected in S503 are within the object region selected in step S304. If the control unit 101 determines that the positions of the wrist joint point and the forearm of the wrist detected in S503 are not within the object region selected in step S305, the process proceeds to step S310. If the control unit 101 determines that the positions of the wrist joint point and the forearm of the wrist detected in S503 are within the object region selected in step S305, the process proceeds to step S508.
[0053] In this embodiment, for example, if object region 211 is selected in step S304 and object region 221 is selected in step S305, the wrist joint point position 212 and the wrist lower point position 413 detected in S503 are located within the area of object region 221. Therefore, the control unit 101 proceeds to step S508. Also, if object region 221 is selected in step S304 and object region 211 is selected in step S305, the wrist joint point position 222 and the wrist lower point position 423 detected in S503 are not located within the area of object region 221. Therefore, the control unit 101 proceeds to step S310.
[0054] In step S508, the control unit 101 determines the arm angle based on the positional relationship between the wrist joint point and the lower wrist point detected in S503 for the object region selected in step S304. The control unit 101 also determines the arm angle based on the positional relationship between the wrist joint point and the lower wrist point detected in S503 for the object region selected in step S305. The control unit 101 then compares the arm angle within the object region selected in step S304 with the arm angle within the object region selected in step S305 and determines whether the difference between the two angles is within a predetermined range (whether the absolute value of the difference between the two angles is less than a threshold). In this embodiment, the predetermined range is set to be between th_angle1 and th_angle2, with th_angle1 = -30° and th_angle2 = 30°.
[0055] For example, suppose object region 211 is selected in step S304 and object region 221 is selected in step S305. As mentioned above, the coordinates of the wrist joint point in object region 211 are (u12,v12)=(100,160), and the coordinates of the wrist lower point in object region 211 are (u13,v13)=(80,180). Also, the coordinates of the wrist joint point in object region 221 are (u22,v22)=(50,180), and the coordinates of the wrist lower point in object region 221 are (u23,v23)=(30,210). Furthermore, the angle 414 formed by the arms is arctan((u13-u12) / (v13-v12))=-45.0°, and the angle 424 formed by the arms is arctan((u23-u22) / (v23-v22))=-41.8°, so the difference between the angles of the two arms is 3.2°. Since this is between th_angle1 and th_angle2, the control unit 101 determines that the difference between the angles of the two arms is within a predetermined range and proceeds to step S309. Note that here, the angle formed by the line connecting the wrist joint point and the lower wrist point with respect to the vertical direction of the image was used as the arm angle for the aforementioned determination, but the angle used for the aforementioned determination is not limited to this. For example, the angle formed by the line connecting the wrist joint point and the lower wrist point with respect to the horizontal direction of the image may also be used as the arm angle for the aforementioned determination.
[0056] The processes in steps S309, S310, S311, and S312 are the same as the flow shown in Figure 3, so their explanation is omitted.
[0057] In step S513, if the control unit 101 determines that all of the object regions detected in step S302 are false detections, it proceeds to step S514. If even one object region is determined to be a correct detection, it proceeds to step S313.
[0058] In step S514, the control unit 101 identifies the object region that is at the top of all the object regions detected in step S302. Then, it determines that the object region identified as being at the top is a positive detection.
[0059] The process in step S313 is the same as the flow shown in Figure 3, so the explanation is omitted.
[0060] As described above, in this embodiment, the control unit 101 detects an object region from the image data that includes the target object, a human hand, and detects the positions of the fingertips of each finger, the joint points including the wrist, and the wrist sub-point within that object region. Next, if the positions of the wrist joint point and wrist sub-point detected in one object region are present in the other object region, the control unit compares the arm angles calculated based on the positional relationship between the wrist joint point and wrist sub-point in each object region. If the difference between the two angles is within a predetermined range, the control unit determines that the other object region is a false detection. However, if all detected object regions are determined to be false detections, the object region that is at the top of all object regions is determined to be a correct detection. This makes it possible to determine whether each detected object region is a correct or false detection when multiple object regions are detected overlapping around the target object.
[0061] (Other embodiments) Furthermore, the present invention can also be realized by performing the following process: that is, supplying software (program) that realizes the functions of the above-described embodiment to a system or device via a network or various storage media, and having the computer (or control unit or MPU, etc.) of the system or device read and execute the program code. In this case, the program and the storage medium storing the program constitute the present invention.
[0062] Although the present invention has been described in detail above based on its preferred embodiments, the present invention is not limited to these specific embodiments, and various forms that do not depart from the spirit of the invention are also included in the present invention. Some of the embodiments described above may be combined as appropriate.
[0063] Furthermore, each functional unit in each of the above embodiments (each modified example) may or may not be individual hardware. The functions of two or more functional units may be implemented by common hardware. Each of the multiple functions of a single functional unit may be implemented by individual hardware. Two or more functions of a single functional unit may be implemented by common hardware. In addition, each functional unit may or may not be implemented by hardware such as an ASIC, FPGA, or DSP. For example, the device may have a processor and a memory (storage medium) in which a control program is stored. The functions of at least some of the functional units of the device may be implemented by the processor reading and executing the control program from the memory.
[0064] The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.
[0065] Furthermore, in each of the examples described above, "processor" refers to a processor in a broad sense, including general-purpose processors (e.g., CPUs) and specialized processors (e.g., GPUs, ASICs, FPGAs, and programmable logic devices, etc.).
[0066] This embodiment includes the following configurations, methods, and programs.
[0067] [Configuration 1] A region detection means for detecting an object region containing a target object from an image, When the region detection means detects multiple object regions from a single image, the position detection means detects the position of the target point in the target object for each object region. If the region detection means detects multiple object regions from a single image, the system includes a determination means for determining whether each object region is a false detection. The determination means determines that the first object region is a false detection if, among the detection target points detected from the second object region by the position detection means, a specific detection target point exists within the first object region in the first object region, as detected by the region detection means from a single image of the first object region. An image recognition device characterized by the following features.
[0068] [Configuration 2] The region detection means detects the smallest region of a predetermined shape that includes the target object as the object region. The image recognition device according to configuration 1, characterized in that...
[0069] [Configuration 3] The predetermined shape is rectangular. The image recognition device according to configuration 2, characterized in that...
[0070] [Structure 4] The determination means determines that the first object region is a false detection if, in the first object region and the second object region detected from a single image by the region detection means, a specific detection target point among the detection target points detected from the second object region by the position detection means is located within the first object region, and the upper edge of the first object region is below a threshold position compared to the upper edge of the second object region. An image recognition device according to any one of configurations 1 to 3, characterized in that it is an image recognition device.
[0071] [Composition 5] The aforementioned object is a human hand, The detection target point is an articular point, which is the point from which the position of at least one of the joints of the hand and the fingertips is estimated. An image recognition device according to any one of configurations 1 to 4, characterized by the above.
[0072] [Composition 6] The aforementioned specific detection points do not include points that estimate the DIP, PIP, MP, and IP joints of each finger. The image recognition device according to configuration 5, characterized in that it is a picture recognition device.
[0073] [Composition 7] The aforementioned specific detection target point is a point that includes the wrist joint point. The image recognition device according to configuration 5 or 6, characterized by the above.
[0074] [Structure 8] The position detection means further detects a first point located between the wrist and the elbow joint as the target point for detection. An image recognition device according to any one of configurations 5 to 7, characterized by the above.
[0075] [Composition 9] The specific detection target point includes the wrist joint point and the first point. The image recognition device according to configuration 8, characterized in that it is a picture recognition device.
[0076] [Configuration 10] The system further includes an acquisition means for acquiring the orientation of the arm in the image based on the position of the wrist joint detected by the position detection means and the position of the first point. The determination means determines whether the first object region is a false detection if, in the first object region and the second object region detected from a single image by the region detection means, a specific detection target point among the detection target points detected from the second object region by the position detection means exists within the first object region, and the difference in angle between the orientation of the first arm in the first object region and the orientation of the second arm in the second object region, acquired by the acquisition means, is smaller than a threshold. The image recognition device according to configuration 9, characterized by the features described herein.
[0077] [Composition 11] The determination means determines that the first object region is a false detection if the specific detection target point of the second object region is located within the first object region, the specific detection target point of the first object region is located within the second object region, and the upper end of the second object region is located above the upper end of the first object region. An image recognition device according to any one of configurations 1 to 10, characterized in that it is an image recognition device.
[0078] [Composition 12] The determination means determines that the first object region is a false detection if the region detection means detects three or more object regions, including the first object region, from a single image, and at least one of the specific detection target points of each of the object regions other than the first object region is located within the first object region. An image recognition device according to any one of configurations 1 to 11, characterized by the above.
[0079] [method] A region detection step for detecting the object region containing the target object from the image, If multiple object regions are detected from a single image by the region detection step, a position detection step is performed to detect the position of the target point in the target object for each object region. If multiple object regions are detected from a single image by the region detection step, the region detection step includes a determination step to determine whether each object region is a false detection or not. In the determination step, if, in the first object region and the second object region detected from a single image by the region detection step, a specific detection target point among the detection target points detected from the second object region by the position detection step is located within the first object region, the first object region is determined to be a false detection. An image recognition method characterized by the following features.
[0080] [program] A program for causing a computer to function as one of the means of an image recognition device described in any one of items 1 to 12.
[0081] [system] A region detection device that detects the object region containing the target object from an image, When the region detection device detects multiple object regions from a single image, a position detection device detects the position of the target point in the target object for each object region. The region detection device has a determination device that, when multiple object regions are detected from a single image, determines whether each object region is a false detection or not. The determination device determines that the first object region is a false detection if, in the first object region and the second object region detected from a single image by the region detection device, a specific detection target point among the detection target points detected from the second object region by the position detection device is located within the first object region. An image recognition system characterized by the following features.
Claims
1. A region detection means for detecting an object region containing an object presumed to be a hand from an image, When the region detection means detects multiple object regions from a single image, for each object region, a position detection means detects the position of a target point estimated to be the position of a hand, including the joint points of the fingers and wrist, on the target object. If the region detection means detects multiple object regions from a single image, the system includes a determination means for determining whether each object region is a false detection. The determination means determines that the first object region is a false detection if, in the first object region and the second object region, which is different from the first object region, detected from the second object region by the position detection means, a specific detection target point exists within the first object region. An image recognition device characterized by the following features.
2. The region detection means detects the smallest region of a predetermined shape that includes the target object as the object region. The image recognition device according to feature 1.
3. The predetermined shape is rectangular. The image recognition device according to feature 2.
4. A region detection means for detecting an object region containing an object presumed to be a hand from an image, When the region detection means detects multiple object regions from a single image, for each object region, a position detection means detects the position of a target point estimated to be the position of a hand, including the joint points of the fingers and wrist, on the target object. If the region detection means detects multiple object regions from a single image, the system includes a determination means for determining whether each object region is a false detection. The determination means determines that the first object region is a false detection if, in the first object region and the second object region, which is different from the first object region, detected from a single image by the region detection means, a specific detection target point among the detection target points detected from the second object region by the position detection means is located within the first object region, and the upper edge of the first object region is below a threshold position compared to the upper edge of the second object region. An image recognition device characterized by the following features.
5. The aforementioned specific detection points do not include points that estimate the DIP, PIP, MP, and IP joints of each finger. The image recognition device according to feature 1.
6. The aforementioned specific detection target point is a point that includes the wrist joint point. The image recognition device according to feature 1.
7. The position detection means further detects a first point located between the wrist and the elbow joint as the target point for detection. The image recognition device according to feature 1.
8. The specific detection target point includes the wrist joint point and the first point. The image recognition device according to feature 7.
9. A region detection means for detecting an object region containing an object presumed to be a hand from an image, When the region detection means detects multiple object regions from a single image, for each object region, a position detection means detects the position of a target point estimated to be the position of a hand, including the joint points of the fingers and wrist, on the target object. When the region detection means detects multiple object regions from a single image, the determination means determines whether each object region is a false detection or not. The system includes an acquisition means for acquiring the orientation of the arm in the image based on the position of the wrist joint detected by the position detection means and the position of a first point located between the wrist and the elbow joint. The determination means determines that the first object region is a false detection if, in the first object region and the second object region different from the first object region detected from a single image by the region detection means, a specific detection target point among the detection target points detected from the second object region by the position detection means exists within the first object region, and the difference in angle between the orientation of the first arm in the first object region and the orientation of the second arm in the second object region, acquired by the acquisition means, is smaller than a threshold. An image recognition device characterized by the following features.
10. The determination means determines that the first object region is a false detection if the region detection means detects three or more object regions, including the first object region, from the single image, and at least one of the specific detection target points of each of the object regions other than the first object region is located within the first object region. The image recognition device according to feature 1.
11. A region detection step that detects an object region containing an object presumed to be a hand from an image, If multiple object regions are detected from a single image by the region detection step, a position detection step is performed to detect the location of the target point on the target object, which is estimated to be the position of each joint point of the hand, including the fingers and wrist, for each object region. If multiple object regions are detected from a single image by the region detection step, the method includes a determination step to determine whether each object region is a false detection or not. In the determination step, if, in the first object region and the second object region different from the first object region detected from a single image by the region detection step, a specific detection target point among the detection target points detected from the second object region by the position detection step exists within the first object region, the first object region is determined to be a false detection. An image recognition method characterized by the following features.
12. A program for causing a computer to function as one of the means of an image recognition apparatus according to any one of claims 1 to 10.
13. A region detection device that detects an object region containing an object presumed to be a hand from an image, When the region detection device detects multiple object regions from a single image, a position detection device detects the position of each object region, which is estimated to be a detection target point that is the position of each joint point of the hand, including the fingers and wrist, on the target object. The region detection device has a determination device that determines whether or not an object region is a false detection when multiple object regions are detected from a single image. The determination device determines that the first object region is a false detection if, in the first object region and a second object region different from the first object region detected by the region detection device from a single image, a specific detection target point among the detection target points detected from the second object region by the position detection device is located within the first object region. An image recognition system characterized by the following features.
14. A region detection step that detects an object region containing an object presumed to be a hand from an image, If multiple object regions are detected from a single image by the region detection step, a position detection step is performed to detect the position of a target point on the target object, which is estimated to be the position of the hand including the joint points of the fingers and wrist, for each object region. If multiple object regions are detected from a single image by the region detection step, the method includes a determination step to determine whether each object region is a false detection or not. In the determination step, if, in the first object region and the second object region, which is different from the first object region, detected from a single image by the region detection step, a specific detection target point among the detection target points detected from the second object region by the position detection step is located within the first object region, and the upper edge of the first object region is below a threshold position compared to the upper edge of the second object region, then it is determined that the first object region is a false detection. An image recognition method characterized by the following features.
15. A region detection device that detects an object region containing an object presumed to be a hand from an image, When the region detection device detects multiple object regions from a single image, a position detection device detects the position of a target point on the target object, which is estimated to be the position of the hand including the joint points of the fingers and wrist, for each object region. The region detection device has a determination device that determines whether or not an object region is a false detection when multiple object regions are detected from a single image. The determination device determines that the first object region is a false detection if, in the first object region and the second object region, which is different from the first object region, detected from a single image by the region detection device, a specific detection target point among the detection target points detected from the second object region by the position detection device is located within the first object region, and the upper edge of the first object region is below a threshold position compared to the upper edge of the second object region. An image recognition system characterized by the following features.
16. A region detection step that detects an object region containing an object presumed to be a hand from an image, If multiple object regions are detected from a single image by the region detection step, a position detection step is performed to detect the position of a target point on the target object, which is estimated to be the position of the hand including the joint points of the fingers and wrist, for each object region. If multiple object regions are detected from a single image by the region detection step, a determination step is performed to determine whether each object region is a false detection or not. The acquisition step includes obtaining the orientation of the arm in the image based on the position of the wrist joint detected by the position detection step and the position of a first point located between the wrist and the elbow joint, In the determination step, if, in the first object region and the second object region different from the first object region detected from a single image by the region detection step, a specific detection target point among the detection target points detected from the second object region by the position detection step exists within the first object region, and the difference in angle between the orientation of the first arm in the first object region and the orientation of the second arm in the second object region, acquired by the acquisition step, is smaller than a threshold, then it is determined that the first object region is a false detection. An image recognition method characterized by the following features.
17. A region detection device that detects an object region containing an object presumed to be a hand from an image, When the region detection device detects multiple object regions from a single image, a position detection device detects the position of a target point on the target object, which is estimated to be the position of the hand including the joint points of the fingers and wrist, for each object region. When the region detection device detects multiple object regions from a single image, a determination device is provided to determine whether each object region is a false detection or not. The device includes an acquisition device that acquires the orientation of the arm in the image based on the position of the wrist joint detected by the position detection device and the position of a first point located between the wrist and the elbow joint. The determination device determines that the first object region is a false detection if, in the first object region and the second object region different from the first object region detected from a single image by the region detection device, a specific detection target point among the detection target points detected from the second object region by the position detection device is located within the first object region, and the difference in angle between the orientation of the first arm in the first object region and the orientation of the second arm in the second object region, acquired by the acquisition device, is smaller than a threshold. An image recognition system characterized by the following features.