Region extraction device, region extraction method

By detecting the position of the elbow and wrist, and combining the color comparison of the hand side with the pose estimation model, the object region is accurately extracted, which solves the problem of object hiding or interference from multiple objects and achieves more accurate region extraction.

CN114158281BActive Publication Date: 2026-01-27RAKUTEN GROUP INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202080006866.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-07
Publication Date
2026-01-27
Estimated Expiration
2040-07-07

AI Technical Summary

Technical Problem

When shooting a product shelf from above, the object is sometimes hidden and cannot be identified, or when shooting from the side, objects other than the object can be identified, resulting in inaccurate area extraction.

Method used

By acquiring frame images, motion information, and detecting the positions of the human elbow and wrist, the region starting from the wrist is extracted based on this information. Combined with color comparison of the hand side and a pose estimation model, the object region is accurately extracted.

Benefits of technology

It improves the accuracy of extracting object regions from images, suppresses the extraction of uncaptured object regions, and appropriately sets object region boundaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114158281B_ABST
    Figure CN114158281B_ABST
Patent Text Reader

Abstract

A region extraction device acquires a first frame image and a second frame image which are continuous in time. The region extraction device acquires motion information indicating a region in which motion exists within the first frame image, based on the acquired first frame image and the second frame image. The region extraction device detects positions of an elbow and a wrist of a human body from a region in which motion exists indicated by the acquired motion information, based on the acquired first frame image. The region extraction device extracts a region corresponding to a portion of a hand side of the human body from the wrist from the region in which motion exists indicated by the acquired motion information, based on the detected positions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for extracting regions containing objects from images. Background Technology

[0002] Previously, techniques for identifying goods taken by a person from a place where goods are displayed were known. For example, Patent Document 1 discloses a market information collection device for estimating goods taken by a customer from a merchandise shelf. This information collection device slides a region across an image taken from the ceiling above the merchandise shelf between the customer and the customer, and calculates the similarity of the feature value of each region with the pre-calculated feature value of each item. The information collection device estimates the item whose similarity exceeds a threshold as the item included in the corresponding region.

[0003] Existing technical documents

[0004] Non-patent literature

[0005] Patent Document 1: Japanese Patent Application Publication No. 2016-201105 Summary of the Invention

[0006] The problem that the invention aims to solve

[0007] However, when shooting from above, the object may sometimes be hidden and unrecognizable. On the other hand, if the shot is taken from near a shelf or the exit of a product and facing outwards, a variety of objects other than the object will appear, and thus, sometimes these objects can be identified.

[0008] The present invention was made in view of the above problems. One example of this subject is to provide a region extraction device, region extraction method and region extraction program that can more accurately extract the region of the object in the image.

[0009] Methods for solving problems

[0010] To address the aforementioned issues, one aspect of the present invention is a region extraction device, characterized by comprising: a frame image acquisition unit that acquires a first frame image and a second frame image that are sequential in time; a motion information acquisition unit that acquires motion information representing a region in the first frame image that contains motion based on the acquired first frame image and the second frame image; a detection unit that detects the positions of the elbow and wrist of a human body from the region containing motion represented by the acquired motion information based on the acquired first frame image; and an extraction unit that extracts a region from the region containing motion represented by the acquired motion information that corresponds to the portion of the human body from the wrist on the hand side based on the detected positions.

[0011] Based on this aspect, motion information representing the region where motion exists within the first frame of the image can be obtained. If the object to be identified is grasped by a human hand, the object, hand, and arm may be moving within the image. Then, the positions of the human elbow and wrist are detected from the region where motion exists. Next, the region corresponding to the hand-side portion of the region where motion exists, starting from the wrist, is extracted. If the region where motion exists is divided into two parts along the wrist boundary, the portion without the elbow is the hand-side portion. The object grasped by the hand overlaps with the hand in the image. Therefore, by extracting the region corresponding to the hand-side portion, the region where the object was captured can be extracted more accurately.

[0012] Another aspect of the present invention is a region extraction device, characterized in that the extraction unit controls the extraction of the region corresponding to the hand-side portion based on a comparison result between the color of the hand-side portion and a predetermined skin tone.

[0013] When a hand is grasping an object, the hand-side portion of the area where the action occurs contains pixels with the object's color. Therefore, it's possible to extract colors other than skin tone from the hand-side portion. Based on this aspect, by comparing the color of the hand-side portion with a defined skin tone, it's possible to estimate whether the hand is grasping an object. Thus, since area extraction is controlled, it's possible to suppress the extraction of areas where no object has been captured.

[0014] Another aspect of the present invention is a region extraction device, characterized in that, when the difference between the color of the hand-side portion and the skin color exceeds a predetermined degree, the extraction unit extracts the region corresponding to the hand-side portion.

[0015] Based on this perspective, since the area is extracted when the color of the part on the hand side differs from the specified skin tone by more than a specified degree, it is possible to suppress the extraction of areas where no object has been photographed.

[0016] Another aspect of the present invention is a region extraction device, characterized in that the detection unit further detects the position of the joints and fingertips of the human finger from the region where the movement exists, and the extraction unit modifies the extracted region according to the position of the joints and fingertips of the finger.

[0017] Based on this, since the position of the object being grasped by the hand can be estimated based on the position of the detected finger joints and fingertips, the area where the object is being photographed can be set more appropriately.

[0018] Another aspect of the present invention is a region extraction device, characterized in that the extraction unit expands the extracted region in a direction from the joint of the finger to the fingertip.

[0019] The object being grasped by the hand tends to extend outwards from the hand in the direction of the fingertips. Based on this aspect, since the area expands in the direction of the fingertips, it is possible to more appropriately set the area where the object is being photographed.

[0020] Another aspect of the present invention is a region extraction device, characterized in that the detection unit uses a prescribed pose estimation model to detect the position of the elbow and the wrist.

[0021] Another aspect of the present invention is a region extraction device, characterized in that the region extraction device further comprises a training unit, which uses an image of the extracted region to train a model for recognizing objects within the image.

[0022] Based on this aspect, the area corresponding to the hand side within the region where movement occurs is used.

[0023] Images are used to train the model. Therefore, by using an image from the first frame that captures the part of the object being grasped by a hand during training, the model can be trained to recognize objects more appropriately.

[0024] Another aspect of the present invention is a region extraction device, characterized in that the region extraction device further comprises an output unit, which outputs object information representing objects existing in the extracted region by inputting an image of the extracted region into a predetermined model.

[0025] Based on this aspect, information representing the object grasped by the hand can be output from the image of the region corresponding to the hand side within the area where movement occurs. Therefore, since the recognition of objects not grasped by the hand can be prevented, it is possible to identify the objects that should have been recognized.

[0026] Another aspect of the present invention is a region extraction device, characterized in that the acquired motion information is dense optical flow.

[0027] Another aspect of the present invention is a region extraction method, characterized by the following steps performed by a computer: a frame image acquisition step, acquiring a first frame image and a second frame image that are sequential in time; an action information acquisition step, acquiring action information representing a region in the first frame image that contains action based on the acquired first frame image and the second frame image; a detection step, detecting the position of the elbow and wrist of a human body from the region containing action represented by the acquired action information based on the acquired first frame image; and an extraction step, extracting a region from the region containing action represented by the acquired action information that corresponds to the portion of the human body from the wrist on the hand side based on the detected position.

[0028] Another aspect of the present invention is a region extraction program, characterized in that a computer functions as the following units: a frame image acquisition unit that acquires a first frame image and a second frame image that are sequential in time; a motion information acquisition unit that acquires motion information representing a region in the first frame image that contains motion based on the acquired first frame image and the second frame image; a detection unit that detects the positions of the elbow and wrist of a human body from the region containing motion represented by the acquired motion information based on the acquired first frame image; and an extraction unit that extracts a region from the region containing motion represented by the acquired motion information that corresponds to the portion of the human body from the wrist on the hand side based on the detected positions.

[0029] Invention Effects

[0030] According to the present invention, it is possible to extract the area containing the object from the image more accurately. Attached Figure Description

[0031] Figure 1 This is a block diagram illustrating an example of the general structure of an image processing apparatus 1 according to one embodiment.

[0032] Figure 2 This is a diagram illustrating an example of the functional blocks of the system control unit 11 and GPU 18 of an image processing apparatus 1 according to one embodiment.

[0033] Figure 3 This is a diagram illustrating an example of the processing flow of the image processing apparatus 1.

[0034] Figure 4 This is a diagram illustrating an example of the effect of the operation of the image processing device 1.

[0035] Figure 5 (a) and (b) are diagrams showing examples of extraction of the region located on the side of hand 110.

[0036] Figure 6 This is a diagram showing an example of an extension of region 600.

[0037] Figure 7 This is a flowchart illustrating an example of the learning processing of the system control unit 11 and GPU 18 in the image processing apparatus 1.

[0038] Figure 8 This is a flowchart illustrating an example of the recognition processing of the system control unit 11 and the GPU 18 in the image processing apparatus 1. Detailed Implementation

[0039] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. The embodiments described below are those in which the present invention is applied to an image processing apparatus that performs learning for generating a model for recognizing objects captured in an image and using the generated model to recognize the objects. Object recognition may include recognizing or classifying objects present in the image. Furthermore, the apparatus for performing learning and the apparatus for recognizing objects may be different apparatuses.

[0040] [1. Structure of an image processing device]

[0041] First, use Figure 1 The structure of the image processing device 1 will be described. Figure 1 This is a block diagram illustrating an example of the general structure of the image processing apparatus 1 according to this embodiment. Figure 1 As shown, the image processing apparatus 1 includes a system control unit 11, a system bus 12, an input / output interface 13, a storage unit 14, a communication unit 15, an input unit 16, a display unit 17, a GPU (Graphics Processing Unit) 18, a GPU memory 19 (or video RAM), and a camera unit 20. The system control unit 11 and the input / output interface 13 are connected via the system bus 12. Examples of image processing apparatus 1 include server devices and personal computers.

[0042] The system control unit 11 consists of a CPU (Central Processing Unit) 11a, a ROM (Read-Only Memory) 11b, and a RAM (Random Access Memory) 11c.

[0043] The input / output interface 13 performs interface processing between the system control unit 11 and the storage unit 14, communication unit 15, input unit 16, display unit 17, GPU 18, GPU memory 19, and camera unit 20.

[0044] Storage unit 14 is configured, for example, by a hard disk drive or a solid-state drive. This storage unit 14 stores the generated model 2 and the training data used in generating model 2. The training data includes motion image data and ground truth labels for the categories of objects present in the motion image represented by the motion image data. Examples of motion image data formats include H.264 and MPEG-2. Storage unit 14 also stores the operating system, the model generation program, the object recognition program, etc. The training data and various programs can be obtained from a designated computer via a network, or recorded on a recording medium such as an optical disc, memory card, or disk and read via a drive device. If the device for generating model 2 and the device for object recognition are different devices, the exchange of the generated model 2 can be performed via a network or via a recording medium.

[0045] The communication unit 15 is composed of, for example, a network interface controller. The communication unit 15 connects to other computers via a defined network such as the Internet or a LAN (Local Area Network) and controls the communication status with those computers.

[0046] The input unit 16 receives the operator's commands and outputs the corresponding signals to the system control unit 11. Examples of input units 16 include keyboards, mice, and touch panels.

[0047] The display unit 17 is composed of, for example, a graphics controller and a display. The display unit 17 displays images, characters, and other information under the control of the system control unit 11. Examples of display panels include liquid crystal panels and organic EL (Light Emitting Diode) panels.

[0048] GPU 18 performs matrix operations and other machine learning operations under control from system control unit 11. GPU 18 pipelines multiple operations in parallel. GPU 18 is connected to GPU memory 19. GPU memory 19 stores the data used in the operations of GPU 18 and the operation results. Alternatively, when system control unit 11 performs all the machine learning operations, GPU 18 and GPU memory 19 are not required.

[0049] The camera unit 20 includes a digital camera equipped with, for example, a CCD (Charge-Coupled Device) sensor or a CMOS (Complementary Metal Oxide Semiconductor) sensor. The camera unit 20 captures moving images under the control of the system control unit 11. The camera unit 20 outputs moving image data representing the captured moving images to the system control unit 11 or the storage unit 14. If the learning apparatus and the object recognition apparatus are different, the learning apparatus may not include the camera unit 20. Furthermore, if object recognition is performed based on moving image data obtained from another computer or recording medium, rather than on moving images captured by the camera unit 20 in real time, the image processing apparatus 1 may also not include the camera unit 20.

[0050] The image processing device 1 may not have at least one of the input unit 16, display unit 17, GPU 18, GPU memory 19, and camera unit 20. At least one of these may be connected to the image processing device 1 via wired or wireless means.

[0051] [2. Functional Overview of the System Control Unit]

[0052] Next, use Figures 2 to 6 The functional overview of the system control unit 11 and GPU 18 is explained. Figure 2 This diagram illustrates an example of the functional blocks of the system control unit 11 and GPU 18 of the image processing apparatus 1 according to this embodiment. The system control unit 11 and GPU 18 cause the CPU 11a to read and execute various codes contained in the program stored in the storage unit 14, such as... Figure 2 As shown, it functions as a frame acquisition unit 111, a motion information acquisition unit 112, a joint detection unit 113, a region extraction unit 114, a training unit 115, and an object information output unit 116.

[0053] Figure 3This diagram illustrates an example of the processing flow of the image processing apparatus 1. The frame acquisition unit 111 acquires frame images that are sequential in time. A frame image is a static image contained within a moving image. The moving image that serves as the source for acquiring frame images is usually a moving image captured by the camera unit 20. However, when training model 2 as described later, the moving image that serves as the source for acquiring frame images may, for example, be pre-stored in the storage unit 14. Imagine that an object 100, which is to be identified, is captured in the moving image. The object 100 to be identified may be any part that is different from the human body. Examples of the object 100 include food, beverages, stationery, daily necessities, and groceries. Furthermore, imagine that the object 100 to be identified is grasped by a human hand 110. Typically, the moving image is captured when the hand 110 and arm 120 holding the object 100 are moving. For example, the moving image may also be captured when someone takes the object 100 from a place or intends to place the object 100 back in its original place. Therefore, it is imagined that object 100, hand 110 grasping object 100, and arm 120 are moving within the motion picture. At least one frame in the frame images contained in the motion picture may not include object 100. That is, object 100 may be drawn in or out. Furthermore, between some frame images, object 100 may not move at all. The captured motion picture contains frames that are consecutive in time. Frames that are consecutive in time refer, for example, frames captured at consecutive shooting times. For example, at a frame rate of 30fps, frames are captured at 1 / 30th of a second interval. The frame acquisition unit 111 may also acquire frame images sequentially from the motion picture data according to the shooting order. Figure 3 In this process, the frame acquisition unit 111 acquires, for example, frame t-1 and frame t. Frame t-1 is the (t-1)th frame image in the sequence of shooting of the frames contained in the moving image. Frame t is the tth frame image. Therefore, frames t-1 and frame t are consecutive in time.

[0054] The motion information acquisition unit 112 acquires motion information 200 representing the region 210 in frame t-1, indicating that motion exists in frame t-1, based on frame t-1 and frame t acquired by the frame acquisition unit 111. The motion region 210 may also be a region where a visual change occurs when the frame changes from frame t-1 to frame t. The motion region 210 may also refer to the region occupied by an object performing motion in frame t-1 when the frame changes. The object performing motion may be, for example, object 100, hand 110, arm 120, and / or other objects. Based on the above assumptions, it is generally assumed that the motion region 210 includes at least the regions occupied by object 100, hand 110, and arm 120. The motion information 200 may also include the coordinates of the motion region 210. Alternatively, the motion information 200 may include information indicating whether motion exists for each pixel in frame t-1. Alternatively, the motion information 200 may include vectors representing the direction and distance of movement for each pixel in frame t-1. The motion information 200 may also be, for example, optical flow. Dense optical flow is a type of optical flow. Dense optical flow represents the action region. Action information 200 can also be dense optical flow. Optical flow can also be generated using a model that includes a convolutional neural network (CNN). Examples of such models include FlowNet, FlowNet 2.0, and LiteFlowNet. Pre-learned models can also be used. Methods that do not use machine learning can also be employed to generate optical flow. Examples of such methods include block matching and gradient methods. Action information 200 can also be information different from optical flow. For example, action information 200 can be generated using inter-frame differencing or background differencing.

[0055] The joint detection unit 113 detects the positions of human joints from the region 210 where an action exists, represented by the action information 200 obtained by the action information acquisition unit 112, based on frame t-1 obtained by the frame acquisition unit 111. Specifically, the joint detection unit 113 detects the positions of the elbow 310 and wrist 320. The joint detection unit 113 may also use a human pose estimation model in detecting the positions of the elbow 310 and wrist 320. This model may, for example, include a CNN. Examples of pose estimation models include DeepPose, Convolutional Pose Machines, and HRNet. In addition to detecting the elbow 310 and wrist 320, the joint detection unit 113 may also detect the positions of the fingertips and finger joints from the action region 210. That is, the joint detection unit 113 may also detect the positions of the fingertips and finger joints constituting the hand 110. The fingers whose fingertips and joints are detected may be at least one of the thumb, index finger, middle finger, ring finger, and little finger. The joint being tested can be at least one of the first, second, and third joints.

[0056] The region extraction unit 114 extracts a region 600 from the region 210 containing movement, represented by the motion information 200 obtained by the motion information acquisition unit 112, based on the positions of the elbow 310 and wrist 320 detected by the joint detection unit 113. This region corresponds to the portion located on the hand 110 side of the body, starting from the wrist 320. Typically, the hand 110 and arm 120 can be divided into hand 110 and arm 120, centered on the wrist 320. For example, the region extraction unit 114 can also calculate a straight line 410 connecting the elbow 310 and wrist 320 at frame t-1 based on the detected positions. The region extraction unit 114 can also calculate a straight line 420 that intersects the straight line 410 at a right angle at the position of the wrist 320. With the straight line 420 as the boundary, the portion of the movement region 210 where the elbow 310 is located is the portion 220 on the arm 120 side. Furthermore, the portion of the movement region 210 without the elbow 310 is the portion 230 on the hand 110 side.

[0057] The region extraction unit 114 can also define a region 600 of a predetermined shape corresponding to the portion 230 on the hand 110 side when determining the portion 230 on the hand 110 side. Region 600 can also be the region surrounding the portion 230 on the hand 110 side. Therefore, if the hand 110 grasps the object 100, the region extraction unit 114 extracts the region surrounding the object 100 as region 600. Region 600 is, for example, a bounding box. The shape of region 600 can be, for example, a rectangle, or other shapes. The region extraction unit 114 can also determine, for example, the coordinates of each vertex in the region of the portion 230 on the hand 110 side where the interior angle is less than 180 degrees. The number of vertices determined can be 4, 3, or 5 or more. Figure 3 In the process, vertices 510, 520, 530, and 540 are determined. The region extraction unit 114 can also determine the minimum and maximum X-coordinates among all vertices, and the minimum and maximum Y-coordinates among all vertices. Then, the region extraction unit 114 can determine the coordinates of region 600 based on the determined X and Y coordinates. For example, the combination of the minimum X and Y coordinates becomes the coordinates of the upper-left vertex of region 600, and the combination of the maximum X and Y coordinates becomes the coordinates of the lower-right vertex of region 600. At frame t-1, the region extraction unit 114 extracts the set region 600 and obtains an image 610 corresponding to that region 600.

[0058] Figure 4 This is a diagram illustrating an example of the effect of the operation of the image processing device 1. Figure 4In frame t-1 shown, objects 100-1 and 100-2 are captured. Object 100-1 is grasped by hand 110. Object 100-2 is placed on a table. The image processing device 1 can identify either object 100-1 or 100-2. However, the object to be identified is object 100-1. Since the hand 110 and arm 120 holding object 100-1 are moving, object 100-1 also moves in the moving image. On the other hand, object 100-2 is not moving. Therefore, the motion information acquisition unit 112 acquires motion information 200 that shows the area occupied by object 100-1, hand 110, and arm 120 as motion area 210. The area occupied by object 100-2 is excluded from this motion area 210. Therefore, it is possible to prevent the extraction of the area where object 100-2, which should not be identified, is captured. In addition, the joint detection unit 113 detects the position of elbow 310 and wrist 320. The region extraction unit 114 can identify the location of the hand 110 from the motion area 210 based on the positions of the elbow 310 and wrist 320. Since it is assumed that the object is being held by the hand 110, the area of ​​the object to be identified can be extracted more accurately by determining the portion 230 on the side of the hand 110.

[0059] The region extraction unit 114 can also control the extraction of the region 600 corresponding to the portion 230 on the hand 110 side within the region 210 where movement exists, based on a comparison between the color of that portion and a predetermined skin tone. This control can also control whether to extract the image 610 corresponding to region 600. The region extraction unit 114 estimates whether the hand 110 is grasping an object based on the color comparison. The region extraction unit 114 may extract region 600 only if it is estimated that the hand 110 is grasping an object.

[0060] The color of the portion 230 located on the side of the hand 110 can, for example, be the average color of that portion 230. For example, the region extraction unit 114 can also calculate the average pixel value within the portion 230. The specified skin color can, for example, be the color of a human hand. For example, the brightness values ​​of the R, G, and B of the skin color can be pre-input into the image processing device 1 by the administrator of the image processing device 1. Alternatively, the image processing device 1 or other devices can calculate the average color of the hand from one or more images of the hand. The calculated average color value can also be pre-stored in the storage unit 14 as the value of the specified skin color.

[0061] For example, if the color difference between the portion 230 on the hand 110 side and a predetermined skin tone exceeds a predetermined level, the region extraction unit 114 can also extract the region 600. The region extraction unit 114 can also use a known algorithm to calculate the color difference. For example, the region extraction unit 114 can calculate the Euclidean distance. Alternatively, the region extraction unit 114 can calculate the difference in brightness values ​​for R, G, and B separately, and then sum the calculated differences in brightness values. The region extraction unit 114 can extract the region 600 only if the value of the color difference exceeds a predetermined threshold. If the hand 110 is grasping an object, the portion 230 on the hand 110 side may contain a relatively large number of pixels with colors other than skin tone. In this case, the average color of the portion 230 on the hand 110 side differs significantly from the skin tone. Therefore, it is possible to estimate whether the hand 110 is grasping the object 100.

[0062] Figure 5 (a) and Figure 5 (b) is a diagram showing an example of extraction of the region located on the side of hand 110. Figure 5 In frame t1-1 shown in (a), a hand 110 grasping an object 100 is captured. Here, the region extraction unit 114 determines the portion 230-1 on the side of the hand 110. The region extraction unit 114 calculates 45, 65, and 100 respectively as the brightness values ​​of R, G, and B of the average color of the portion 230-1. On the other hand, the brightness values ​​of R, G, and B for a given skin tone are 250, 180, and 100 respectively. In this case, since the color difference is greater than the predetermined level, the region extraction unit 114 extracts the region 600-1 surrounding the portion 230-1 on the side of the hand 110. On the other hand, in Figure 5 In frame t2-1 shown in (a), a hand 110 that is not grasping anything is captured. Here, the region extraction unit 114 determines the portion 230-2 on the side of the hand 110. The region extraction unit 114 calculates 230, 193, and 85 as the brightness values ​​of R, G, and B of the average color of the portion 230-2. In this case, since the color difference is smaller than a specified degree, the region extraction unit 114 does not extract the region 600-2 surrounding the portion 230-2 on the side of the hand 110.

[0063] When the joint detection unit 113 detects the positions of the joints and fingertips of a human's fingers, the region extraction unit 114 can also modify the extracted region 600. Since the position of the object 100 grasped by the hand 110 can be estimated to some extent based on the positions of the finger joints and fingertips, the region 600 can be modified accordingly. For example, the region extraction unit 114 can also expand the region 600 according to the direction from the finger joints to the fingertips. When the object 100 is being grasped by the hand 110, the object 100 usually overlaps with the fingers within frame t-1. Furthermore, the object 100 tends to extend in the direction of the fingertips. Therefore, by making the region 600 have a margin in the direction of the fingertips, the region 600 surrounding the object 100 can be appropriately set.

[0064] The direction from the finger joint to the fingertip can be any one of the following: from the first joint to the fingertip, from the second joint to the fingertip, or from the third joint to the fingertip. For example, if the first joint is detected, the region extraction unit 114 may preferentially use the direction from the first joint to the fingertip. If the second joint is detected but the first joint is not detected, the region extraction unit 114 may also use the direction from the second joint to the fingertip. If only the third joint is detected, the region extraction unit 114 may also use the direction from the third joint to the fingertip.

[0065] To accommodate the detection of joint and fingertip positions for multiple fingers individually, a priority can be pre-determined based on the direction of which finger is being targeted. For example, the priority can be determined in the order of index finger, middle finger, ring finger, little finger, and thumb. When the index finger is detected, the region extraction unit 114 can also determine the direction of region 600 expansion based on the joint and fingertip positions of the index finger. Similarly, when the middle finger is detected but not the index finger, the region extraction unit 114 can determine the direction of region 600 expansion based on the joint and fingertip positions of the middle finger. Alternatively, the region extraction unit 114 can synthesize a direction vector from the joint to the fingertip for each detected finger without using a priority. Then, the region extraction unit 114 can expand the region 600 based on the synthesized direction vector.

[0066] The region extraction unit 114 can also expand the area of ​​region 600 by a predetermined proportion relative to the original area of ​​region 600. Alternatively, the region extraction unit 114 can also expand the length of region 600 by a predetermined proportion relative to the length of its long or wide side.

[0067] The region extraction unit 114 may also expand the region 600 in the direction closest to the direction from the finger joint to the fingertip among the up, down, left, and right directions. Alternatively, the region extraction unit 114 may expand the region 600 in the direction corresponding to the X and Y components of the direction vector from the finger joint to the fingertip. For example, if the direction of the fingertip is the upper right direction, the region extraction unit 114 may expand the region 600 in the right and up directions. In this case, the region extraction unit 114 may also determine the ratio of the horizontal expansion amount of the region 600 to the vertical expansion amount of the region 600 based on the ratio of the X and Y components of the direction vector.

[0068] Figure 6 This is a diagram illustrating an example of an extension of region 600. In Figure 6 In frame t-1 shown, a hand 110 grasping the object 100 is captured. Here, the joint detection unit 113 detects the positions of the joints 710 and fingertips 720 of the index, middle, ring, and little fingers from the hand 110. In each finger, the direction 800 from the joint 710 to the fingertip 720 is approximately to the left. Therefore, the region extraction unit 114 can also expand the region 600 to the left by a predetermined proportion.

[0069] return Figure 3 The training unit 115 uses the image 610 of region 600 extracted by the region extraction unit 114 to train a model 2 for recognizing objects within an image. Model 2 can also be a classifier. Model 2 can also output object information 620 representing the probability of the presence of objects of each category in image 610. Model 2 can be a CNN. Examples of CNNs include ResNet, GoogleNet, AlexNet, and VGGNet. Since the image of the object 100 grasped by hand 110 is used in the training of model 2, model 2 can be generated to appropriately recognize the object 100 to be recognized. Here, in addition to the category of the object to be recognized, an "empty" category can also be defined. The "empty" category represents the category where hand 110 is not grasping anything. In dynamic images captured by the camera unit 20, sometimes a hand 110 is captured without grasping anything. To address such situations, an "empty" category is defined. The training unit 115 can also train model 2 using images 610 extracted from dynamic images of hands 110 grasping various categories of objects to be identified, and images 610 extracted from dynamic images of hands 110 not grasping anything. Furthermore, when model 2 is trained using a device other than image processing device 1, or when image processing device 1 uses a trained model to identify object 100, training unit 115 is not required in image processing device 1.

[0070] The object information output unit 116 inputs the image 610 of the region 600 extracted by the region extraction unit 114 into a predetermined model and outputs object information 620 representing the object 100 present in the extracted region 600. Thus, the object 100 is identified. The model used is a model for identifying objects within an image. This model outputs object information 620 representing the probability of the presence of each category of object in the image 610. This model can also be a classifier. This model can also be Model 2 trained by the training unit 115. Alternatively, this model can also be a model trained using a different method than the training by the training unit 115. For example, this model can also be a model trained using dynamic or static images of a hand 110 grasping an object of each category to be identified. The image processing device 1 can also determine the category of object 100 from the object information 620 output by the object information output unit 116, for example, the category with the highest probability of occurrence and whose probability of occurrence exceeds a predetermined threshold. In the case where the probability of occurrence of the "empty" category is the highest, the image processing device 1 can also determine that no object to be identified has been captured. The object information output unit 116 can also serve as the object recognition result, outputting the coordinates and dimensions of the region 600 in addition to the object information. Furthermore, when recognizing the object 100 using a device other than the image processing device 1, the object information output unit 116 is not required in the image processing device 1.

[0071] [3. Operation of the image processing device]

[0072] Next, use Figure 7 and Figure 8 The operation of the image processing device 1 will be explained. Figure 7 This is a flowchart illustrating an example of the learning process performed by the system control unit 11 and the GPU 18 in the image processing apparatus 1. The system control unit 11 and the GPU 18 perform the learning process according to the program code included in the program for model generation. For example, the learning process can also be performed according to instructions from the operator using the input unit 16.

[0073] like Figure 7 As shown, the frame acquisition unit 111 acquires the first set of motion image data and category labels contained in the training data stored in the storage unit 14 (step S101). Next, the frame acquisition unit 111 sets the frame number t to 1 (step S102). Next, the frame acquisition unit 111 acquires frame t from the acquired motion image data (step S103).

[0074] Next, the frame acquisition unit 111 increments the frame number t by 1 (step S104). The frame acquisition unit 111 acquires frame t from the acquired motion image data (step S105). Next, the motion information acquisition unit 112 acquires motion information 200 based on frame t-1 and frame t (step S106). For example, the motion information acquisition unit 112 acquires motion information 200 by inputting frame t-1 and frame t into a model for dense optical flow generation. Frame t-1 at this moment is the frame acquired in step S102.

[0075] Next, at frame t-1, the joint detection unit 113 detects the positions of the elbow 310 and wrist 320 respectively from the motion region 210 represented by the motion information 200 (step S107). For example, the joint detection unit 113 obtains the coordinates of the elbow 310 and wrist 320 respectively by inputting frame t-1 into the pose estimation model. The joint detection unit 113 extracts coordinates representing the positions within the motion region 210 from the obtained coordinates.

[0076] Next, the region extraction unit 114 determines the hand 110 side region 230 within the motion region 210 represented by the motion information 200 based on the acquired coordinates (step S108). For example, the region extraction unit 114 calculates the boundary line 420 passing through the wrist 320. The region extraction unit 114 divides the motion region 210 into two regions along the boundary line 420. The region extraction unit 114 identifies the region without the elbow 310 within these two regions as the hand 110 side portion 230.

[0077] Next, the region extraction unit 114 calculates the average color of the determined portion 230 on the hand 110 side. Then, the region extraction unit 114 calculates the difference between the average color of the portion 230 and a predetermined skin tone (step S109). Next, the region extraction unit 114 determines whether the calculated color difference is greater than a predetermined threshold (step S110). If the color difference is greater than the threshold (step S110: Yes), the region extraction unit 114 extracts the region 600 corresponding to the portion 230 on the hand 110 side (step S111). For example, the region extraction unit 114 determines the coordinates of each vertex of the portion 230. Based on the coordinates of all vertices, the region extraction unit 114 determines the minimum X and Y coordinates and the maximum X and Y coordinates. The region extraction unit 114 uses the determined coordinates to determine the coordinates of the region 600. In addition, if the category label obtained in step S101 is "empty", the region extraction unit 114 may omit steps S109 and S110 and always set the region 600.

[0078] Next, in frame t-1, the joint detection unit 113 detects the positions of the finger joints 710 and fingertips 720 from the motion region 210 represented by the motion information 200 (step S112). Alternatively, in step S107, the joint detection unit 113 may also detect the positions of the elbow 310 and wrist 320, as well as the positions of the finger joints 710 and fingertips 720.

[0079] Next, the region extraction unit 114 determines the direction from the joint 710 to the fingertip 720 based on the detected positions of the joint 710 and the fingertip 720 (step S113). For example, the region extraction unit 114 identifies the first joint and calculates the vector from the first joint to the fingertip. When multiple fingers are detected with joints and fingertips, the region extraction unit 114 determines, for example, which finger's direction to use based on priority. The region extraction unit 114 determines which direction (left or right) and how much to expand the region 600 based on the X component of the fingertip's direction vector. Furthermore, the region extraction unit 114 determines which direction (up or down) and how much to expand the region 600 based on the Y component of the direction vector. Next, the region extraction unit 114 expands the region 600 according to the determined direction and expansion amount, and obtains the coordinates of the expanded region 600 (step S114).

[0080] Next, the region extraction unit 114 extracts an image 610 corresponding to the set region 600 from frame t-1 (step S115). Then, the training unit 115 inputs the extracted image 610 into model 2 to obtain object information 620 (step S116). Next, the training unit 115 calculates the error between the obtained object information 620 and the category label obtained in step S101. Then, the training unit 115 updates the weights and biases of model 2 by backpropagating the calculated error (step S117). Furthermore, for ease of explanation, the weights are updated for each frame; however, for example, the weights may also be updated for each batch or each dynamic image data containing a predetermined number of frames.

[0081] Next, the training unit 115 determines whether frame t+1 exists in the acquired motion image data (step S118). If frame t+1 exists (step S118: yes), the process proceeds to step S104.

[0082] If there is no frame t+1 (step S118: No) or the color difference is not greater than a threshold (step S110: No), the training unit 115 determines whether there is the next dynamic image data in the training data (step S119). If there is the next dynamic image data (step S119: Yes), the frame acquisition unit 111 acquires the next set of dynamic image data and category label from the training data (step S120), and the process proceeds to step S102. On the other hand, if there is no next dynamic image data (step S119: No), the training unit 115 determines whether to end the learning process (step S120). For example, if the training unit 115 has performed a number of learning iterations equivalent to a predetermined time interval, it may also determine that the learning process has ended. Alternatively, the training unit 115 may also perform object recognition using test data and calculate the recognition error. If the average value of the calculated recognition error is less than a predetermined value, the training unit 115 may also determine that the learning process has ended. If the learning process has not ended (step S121: No), the process proceeds to step S101. If the learning process is complete (step S121: Yes), the learning process ends.

[0083] Figure 8 This is a flowchart illustrating an example of the recognition processing performed by the system control unit 11 and the GPU 18 of the image processing apparatus 1. Figure 8 In the middle, to and Figure 7 The same steps are labeled with the same numbers. Figure 8 The processing example shown is a case where an object is identified in real time based on a moving image captured by the camera unit 20. For example, using a method based on... Figure 7 The learning process shown is completed using Model 2 to perform recognition processing. The system control unit 11 and GPU 18 perform recognition processing according to the program code included in the object recognition program. For example, recognition processing can also be performed when the capture of a moving image by the camera unit 20 begins according to an instruction from the system control unit 11.

[0084] like Figure 8 As shown, the frame acquisition unit 111 sets the frame number t to 0 (step S201). Next, the frame acquisition unit 111 increments the frame number t by 1 (step S202). Next, the frame acquisition unit 111 acquires the latest frame as frame t from the imaging unit 20 (step S203).

[0085] Next, the frame acquisition unit 111 determines whether the frame number t is greater than 1 (step S204). If the frame number t is not greater than 1 (step S204: no), the process proceeds to step S202.

[0086] On the other hand, if the frame number t is greater than 1 (step S204: Yes), the following steps are executed. In step S110, if the color difference is greater than the threshold (step S110: Yes), proceed to step... Next, the object information output unit 116 inputs the image 610 extracted in step S115 into the model 2 and outputs object information 620 (step S205).

[0087] After step S205, or if the color difference is not greater than a threshold (step S110: No), the object information output unit 116 determines whether to end object recognition (step S206). The conditions for ending recognition can be predetermined according to the purpose of the image processing apparatus 1. If recognition has not ended (step S206: No), the process proceeds to step S202. On the other hand, if recognition has ended (step S206: Yes), the recognition process ends.

[0088] As explained above, according to this embodiment, the image processing apparatus 1 acquires frames t-1 and t that are sequential in time. Furthermore, based on the acquired frames t-1 and t, the image processing apparatus 1 acquires motion information 200 representing the region 210 where motion exists within frame t-1. Furthermore, based on the acquired frame t-1, the image processing apparatus 1 detects the positions of the elbow 310 and wrist 320 of the human body from the region 210 where motion exists, as represented by the acquired motion information 200. Furthermore, based on the detected positions, the image processing apparatus 1 extracts a region 600 from the region 210 where motion exists, representing the region 210 where motion exists, that corresponds from the wrist 320 to the portion 230 on the side of the human body's hand 110. Since the object 100 grasped by the hand 110 overlaps with the hand 110 in the image, by extracting the region 600 corresponding to the portion on the hand 110 side, the region 600 where the object 100 is captured can be extracted more accurately.

[0089] Here, the image processing device 1 can also control the extraction of the region 600 corresponding to the portion 230 on the hand 110 side based on the comparison result of the color of the portion 230 on the hand 110 side with a predetermined skin tone. In this case, by comparing the color of the portion 230 on the hand 110 side with the predetermined skin tone, it is possible to estimate whether the hand 110 is grasping the object 100. Therefore, since the extraction of the region 600 is controllable, it is possible to suppress the extraction of regions where the object 100 is not captured.

[0090] Here, if the color difference between the portion 230 on the hand 110 side and the skin tone exceeds a predetermined level, the image processing device 1 can also extract the region 600 corresponding to the portion 230 on the hand 110 side. In this case, it is possible to suppress the extraction of regions where the object 100 was not photographed.

[0091] Furthermore, the image processing device 1 can further detect the positions of the joints 710 and fingertips 720 of the human fingers from the area 210 where movement exists. Additionally, the image processing device 1 can modify the extracted area 600 based on the positions of the finger joints 710 and fingertips 720. In this case, since the position of the object 100 grasped by the hand 110 can be estimated based on the detected positions of the finger joints 710 and fingertips 720, the area 600 where the object 100 is captured can be set more appropriately.

[0092] Here, the image processing device 1 can also expand the extracted area 600 in the direction from the finger joint 710 to the fingertip 720. In this case, since the area 600 expands in the direction of the fingertip, the area where the object 100 is photographed can be set more appropriately.

[0093] In addition, the image processing device 1 can also use a prescribed pose estimation model to detect the position of the elbow 310 and the wrist 320.

[0094] Furthermore, the image processing device 1 can also use the image 610 of the extracted region 600 to train a model 2 for recognizing the object 100 within the image. In this case, the model is trained using the image 610 of the region 600 corresponding to the portion 230 on the hand 110 side of the region 210 where action occurs. Therefore, since the image 610 of frame t-1, which captures the portion of the object 100 grasped by the hand 110, is used during training, the model 2 can be trained to more appropriately recognize the object 100.

[0095] Furthermore, the image processing device 1 can also output object information 620 representing the object 100 present in the extracted region 600 by inputting the image 610 of the extracted region 600 into a predetermined model. In this case, information representing the object 100 grasped by the hand 110 is output from the image 610 of the region 600 corresponding to the portion 230 on the hand 110 side of the region 210 where the action occurs. Therefore, since the recognition of the object 100 not grasped by the hand can be prevented, the object 100 that should have been recognized can be identified.

[0096] In addition, the acquired motion information 200 can also be dense optical flow.

[0097] Label Explanation

[0098] 1: Image processing device;

[0099] 11: System Control Department;

[0100] 12: System bus;

[0101] 13: Input / output interfaces;

[0102] 14: Storage Department;

[0103] 15: Ministry of Communications;

[0104] 16: Input section;

[0105] 17: Display section;

[0106] 18: GPU;

[0107] 19: GPU memory;

[0108] 20: Camera Department;

[0109] 111: Frame Acquisition Unit;

[0110] 112: Motion Information Acquisition Unit;

[0111] 113: Joint Detection Section;

[0112] 114: Regional Extraction Department;

[0113] 115: Training Department;

[0114] 116: Object Information Output Section;

[0115] 2: Model.

Claims

1. A region extraction device, characterized in that, have: The frame image acquisition unit acquires the first frame image and the second frame image that are consecutive in time; The motion information acquisition unit acquires motion information based on the acquired first frame image and second frame image, the motion information indicating the area where motion exists in the first frame image; The detection unit detects the position of the elbow and wrist of the human body from the region where the action exists, represented by the acquired motion information, based on the acquired first frame image. as well as The extraction unit extracts, based on the detected elbow and wrist positions, a region from the wrist that corresponds to the hand side of the human body within the region of the action represented by the acquired motion information.

2. The region extraction device according to claim 1, characterized in that, The extraction unit controls the extraction of the area corresponding to the hand side based on a comparison between the color of the hand side portion and a specified skin tone.

3. The region extraction device according to claim 2, characterized in that, If the color difference between the hand-side portion and the skin tone exceeds a predetermined level, the extraction unit extracts the area corresponding to the hand-side portion.

4. The region extraction device according to any one of claims 1 to 3, characterized in that, The detection unit further detects the position of the joints and fingertips of the human fingers from the area where movement occurs. The extraction unit modifies the extraction area based on the position of the finger joints and fingertips.

5. The region extraction device according to claim 4, characterized in that, The extraction unit expands the extraction area in a direction from the joint of the finger to the fingertip.

6. The region extraction device according to any one of claims 1 to 3, characterized in that, The detection unit uses a prescribed pose estimation model to detect the position of the elbow and the position of the wrist.

7. The region extraction device according to any one of claims 1 to 3, characterized in that, The region extraction device also includes a training unit that uses images of the extracted regions to train a model for recognizing objects within the images.

8. The region extraction device according to any one of claims 1 to 3, characterized in that, The region extraction device also has an output unit that, by inputting an image of the extracted region into a specified model, outputs object information representing objects existing in the extracted region.

9. The region extraction device according to any one of claims 1 to 3, characterized in that, The obtained motion information is dense optical flow.

10. A region extraction method, characterized in that, The computer performs the following steps: The frame image acquisition step involves acquiring the first and second frame images that are consecutive in time. The motion information acquisition step involves acquiring motion information based on the acquired first frame image and second frame image, whereby the motion information indicates the region where motion exists within the first frame image. The detection step involves detecting the position of the elbow and wrist of the human body from the region where the action exists, represented by the acquired motion information, based on the first frame image obtained. as well as The extraction step involves extracting, based on the detected elbow and wrist positions, a region corresponding to the portion of the human body on the hand side from the wrist within the region where the action is present, represented by the acquired motion information.

11. The region extraction method according to claim 10, characterized in that, In the extraction step, the extraction of the region corresponding to the hand side is controlled based on the comparison result between the color of the hand side portion and the specified skin tone.

12. The region extraction method according to claim 11, characterized in that, In the extraction step, if the color difference between the hand-side portion and the skin tone exceeds a predetermined level, the area corresponding to the hand-side portion is extracted.

13. The region extraction method according to any one of claims 10 to 12, characterized in that, In the detection step, the positions of the joints and fingertips of the human fingers are further detected from the area where movement exists. In the extraction step, the extraction area is modified according to the position of the finger joints and fingertips.

14. The region extraction method according to claim 13, characterized in that, In the extraction step, the extraction area is expanded in the direction from the joint of the finger to the fingertip.

15. The region extraction method according to any one of claims 10 to 12, characterized in that, In the detection step, a prescribed pose estimation model is used to detect the position of the elbow and the position of the wrist.

16. The region extraction method according to any one of claims 10 to 12, characterized in that, The region extraction method also includes a training step in which an image of the extracted region is used to train a model for recognizing objects within the image.

17. The region extraction method according to any one of claims 10 to 12, characterized in that, The region extraction method also has an output step, in which an image of the extracted region is input into a specified model, and object information representing objects existing in the extracted region is output.

18. The region extraction method according to any one of claims 10 to 12, characterized in that, The obtained motion information is dense optical flow.

Citation Information

Patent Citations

  • Information processor and information processing method

    JP2016201105A

  • Multi-area real-time motion detection method based on monitoring video

    CN108764148A

  • Image recognition device and program

    JP2010113530A

  • Image processing device, image processing system, image processing program, and image processing method

    WO2018135326A1