Object gripping detection system
A two-stage detection process using first and second captured images accurately confirms object grasping by a person, addressing the limitations of conventional systems in detecting actual picking up of objects.
Patent Information
- Application Number
- JP2024079195
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-15
- Publication Date
- 2025-11-28
AI Technical Summary
Conventional systems struggle to accurately detect when a person is grasping an object, as they can only determine contact but not the actual act of picking up the object.
A two-stage detection process using first and second captured images to identify contact and confirm grasping, involving object detection, contact detection, and grip detection using 2D and 3D position information.
Accurately detects when a person is grasping an object by confirming the continuation of contact in subsequent images, enhancing the precision of object grasping detection.
Smart Images

Figure 2025173595000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a system for detecting the grasping of an object by a person. [Background technology]
[0002] Japanese Patent Application Laid-Open Publication No. 2022-165483 discloses a device for detecting contact between a person and an object using a camera image. This conventional device applies a trained model to the camera image to obtain positional information on the person and object contained in the camera image. The conventional device also calculates the degree of overlap between the person region and the object region based on the positional information on the person and object in the camera image. The conventional device further determines that the person has come into contact with the object if the degree of overlap is equal to or greater than a threshold.
[0003] In addition to JP 2022-165483 A, JP 2010-244413 A and Japanese Patent No. 7266145 A can be exemplified as documents showing the technical state of the technical field related to the present disclosure. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2022-165483 [Patent Document 2] Japanese Patent Application Laid-Open No. 2010-244413 [Patent Document 3] Patent No. 7266145 Summary of the Invention [Problem to be solved by the invention]
[0005] Consider a scenario in which a person (customer) visiting a store picks up an item displayed on a shelf. Conventional devices may be able to detect contact between the person and the shelf or item based on the degree of overlap between the person's area and the shelf or item area. However, it is also possible that the customer simply approaches the shelf, or reaches for the item but stops. In other words, conventional devices cannot detect whether the customer actually picked up an item on the shelf. Therefore, improvements are needed to accurately detect when a person is grasping an object.
[0006] An object of the present disclosure is to provide a technology that can accurately detect when a person is grasping an object. [Means for solving the problem]
[0007] The present disclosure relates to a system for detecting a person's grasp of an object, and has the following features. The system includes a storage device and a processing circuit. The storage device stores a first captured image of a target space and a second captured image captured in the target space after the first captured image. The processing device performs a first detection process using the first captured image and a second detection process using the detection result of the first detection process and the second captured image. The first detection process includes acquiring position information of a person and an object recognized from the first captured image, and detecting contact between the person and the object recognized from the first captured image based on the position information of the person and the object in the first captured image. The second detection process includes identifying the second captured image in which a person identical to a target person representing a person whose contact with an object was detected in the first detection process is recognized, acquiring positional information of the target person's hand and object recognized from the identified second captured image, and detecting, based on the positional information of the target person's hand and object in the identified second captured image, the target person's grasping of a contact object representing the object whose contact with the target person was detected in the first detection process. [Effects of the Invention]
[0008] According to the present disclosure, a first detection process is performed using a first captured image of a target space, and a person (target person) who is detected to have come into contact with an object recognized from the first captured image and an object (contacting object) who is detected to have come into contact with the target person are identified. Furthermore, a second detection process is performed using a second captured image captured in the target space after the first captured image, and the target person identified by the first detection process is detected to be holding the contacting object identified by the first detection process. Therefore, it is possible to accurately detect the target person's holding of the contacting object. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a conceptual diagram illustrating an overview of an embodiment of the present disclosure. [Figure 2] 1 is a block diagram showing a first configuration example of a system according to an embodiment. [Figure 3] FIG. 10 is a block diagram showing a second configuration example of the system according to the embodiment. [Figure 4] 10A to 10C are diagrams illustrating an example of a first detection process using 3D position information. [Figure 5] 10A to 10C are diagrams illustrating an example of contact detection processing using 3D position information. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In each drawing, the same or corresponding parts are denoted by the same reference numerals, and the description thereof will be simplified or omitted.
[0011] 1. Overview FIG. 1 is a conceptual diagram illustrating an overview of a system according to an embodiment of the present disclosure. FIG. 1 depicts a person HM and a product GD (products GD1, GD2, and GD3 in the example of FIG. 1) as an example of an object that can be grasped by the person HM. The person HM approaches a product shelf GS on which the product GD is displayed, picks up the product GD (product GD3 in the example of FIG. 1), and then leaves the product shelf GS. The person HM may simply approach the product shelf GS and then leave the product shelf GS without picking up the product GD. The person HM may approach the product shelf GS and reach for the product GD but then leave the product shelf GS without picking it up.
[0012] A series of actions of the person HM can be understood in relation to the time of image capture by the camera. In this embodiment, consider a camera CA1 that captures an image of a "target space" including a product shelf GS. By performing image analysis of a 2D image of the target space captured by the camera CA1 (hereinafter also referred to as a "first captured image IMG1"), the actions of the person HM at time T1, i.e., the actions of the person HM near the product shelf GS, can be understood.
[0013] In the embodiment, a camera CA2 that captures an image of the target space is also considered. The camera CA2 may be the same camera as the camera CA1, or may be a camera that is installed in the target space where the camera CA1 is installed but has a different imaging range from the camera CA1. By analyzing a 2D image of the target space captured by the camera CA2 (hereinafter also referred to as a "second captured image IMG2"), the behavior of the person HM at time T2, i.e., the behavior of the person HM at times after time T1, can be understood.
[0014] In this embodiment, a "first detection process" is performed using the first captured image IMG1 to detect contact between the person HM and the commodity GD. Furthermore, a "second detection process" is performed using the detection result of the first detection process and the second captured image IMG2 to detect the grip of the commodity GD by the person HM (hereinafter also referred to as the "target person TG") who was detected to be in contact with the commodity GD in the first detection process.
[0015] When only the first detection process is performed, it is possible to detect contact between the product GD and the person HM, but it is difficult to detect whether the person HM (i.e., the target person TG) actually picked up the product GD. In this regard, in the embodiment, the second detection process is performed in addition to the first detection process. Therefore, it is possible to accurately detect whether the person HM (i.e., the target person TG) who came into contact with the product GD in the first detection process actually picked up (held) the product GD.
[0016] The first detection process can also be considered a process of estimating whether the person HM is holding the product GD. The second detection process can also be considered a process of confirming whether the person HM is holding the product GD, when the person HM is estimated to be holding the product GD. As described above, according to the embodiment, a two-stage process is performed using the first captured image IMG1 and the second captured image IMG2, which makes it possible to accurately detect whether the person HM is holding the product GD.
[0017] 2. System configuration example 2-1. First configuration example Fig. 2 is a block diagram showing a first configuration example of a system according to an embodiment. The configuration example shown in Fig. 2 includes a camera CA and a data processing device 10. The camera CA and the data processing device 10 are connected via a communication network (not shown). The communication network is not particularly limited, and wired and wireless networks may be used.
[0018] A camera CA is installed in the target space. The installation position and shooting range of the camera CA are known. The camera CA acquires at least a first captured image IMG1 and a second captured image IMG2. The camera CA that acquires the first captured image IMG1 (for example, camera CA1 in FIG. 1) and the camera CA that acquires the second captured image IMG2 (for example, camera CA2 in FIG. 1) may be the same camera or different cameras. The first captured image IMG1 and the second captured image IMG2 may be images (frames) that constitute a video.
[0019] The data processing device 10 includes at least one processing circuit and at least one storage device. Examples of the processing circuit include a general-purpose processor, a specific-purpose processor, a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), and a field-programmable gate array (FPGA). Examples of the storage device include a hard disk drive (HDD), a solid state drive (SSD), a volatile memory, and a non-volatile memory.
[0020] The processing circuit deploys various programs stored in the storage device and processes various data stored in the storage device. Data processing by the processing circuit includes the first detection process and the second detection process described above. FIG. 2 illustrates an object detection unit 11 and a contact detection unit 12 as functional blocks for the first detection process. FIG. 2 also illustrates an object detection unit 13, an image identification unit 14, and a grip detection unit 15 as components for the second detection process. These functional blocks are realized, for example, by cooperation between the processing circuit and the storage device of the data processing device 10.
[0021] The object detection unit 11 performs an object detection process to detect an object from the first captured image IMG1. In the object detection process, for example, an object included in the first captured image IMG1 is detected using a YOLO (You Only Look Once) network, an SSD (Single Shot multibox Detector) network, or the like. Detection targets in the object detection process are static objects (e.g., product shelves GS, products GD) and dynamic objects. Examples of static objects include buildings, structures, and natural objects. Examples of dynamic objects include people HM, as well as robots, bicycles, and automobiles.
[0022] In the object detection process, detection information DET1 of the detection target is generated. The detection information DET1 includes, for example, ID information of the first captured image IMG1, timestamp information, image information of the detection target, type information of the detection target, and 2D position information of the detection target. The 2D position information includes position information of a bounding box assigned to surround the detection target. When the detection target is a person HM, information on features for human re-identification of this person HM (hereinafter also referred to as "ReID features") is extracted from the person image and added to the detection information DET1. Note that extraction of ReID features itself is well known. The detection information DET1 is transmitted to the contact detection unit 12.
[0023] The contact detection unit 12 performs a contact detection process to detect contact between the person HM and an object (for example, a product GD) based on the detection information DET1 received from the object detection unit 11. In the contact detection process, the degree of overlap between the 2D positions of the person HM and the object is calculated based on the 2D position information of these positions included in the detection information DET1. When detection information of the hand of the person HM is obtained, the degree of overlap between the 2D position of the hand of the person HM and the 2D position of the object may be calculated.
[0024] In the contact detection process, the overlapping degree of the 2D positions is also compared with a preset threshold. If the overlapping degree of the 2D positions is equal to or greater than the threshold, the person HM having the 2D position used in this overlapping degree comparison is identified as the target person TG. The contact detection unit 12 transmits the ReID feature amount of the target person TG identified in this way to the image identification unit 14 and the grip detection unit 15.
[0025] 2, the overlapping degree is calculated based on the 2D position information of the detection target, but the overlapping degree between the 3D position of the person HM (or the hand of the person HM) and the 3D position of the object may be calculated based on 3D position information obtained by adding depth information of the object separately acquired from the first captured image IMG1 to the 2D position information. An example of calculating the overlapping degree using the 3D position information will be described later.
[0026] The object detection unit 13 performs an object detection process to detect a person HM from the second captured image IMG2. The content of the object detection process by the object detection unit 13 is basically the same as that of the object detection process by the object detection unit 11. In the object detection process by the object detection unit 13, detection information DET2 of the detection target is generated. The detection information DET2 includes, for example, ID information of the second captured image IMG2, timestamp information, image information of the detection target, type information of the detection target, and 2D position information of the detection target. If the detection target is a person HM, information on the ReID feature of this person HM is added to the detection information DET2. The detection information DET2 is transmitted to the image specification unit 14 and the grip detection unit 15.
[0027] The image identification unit 14 performs an image identification process to identify the second captured image IMG2 in which the target person TG appears (hereinafter also referred to as the "second captured image IMG2_TG"). In the image identification process, the similarity between the ReID feature of the target person TG received from the contact detection unit 12 and the ReID feature of the person HM included in the detection information DET2 received from the object detection unit 13 is calculated.
[0028] In the image identification process, the similarity of the ReID features is also compared with a preset threshold. If the similarity of the ReID features is equal to or greater than the threshold, the second captured image IMG2, which shows the person HM having the ReID features used to calculate the overlap, is identified as the second captured image IMG2_TG. The image identification unit 14 transmits information about the second captured image IMG2_TG identified in this way to the grip detection unit 15.
[0029] The grip detection unit 15 performs a grip detection process to detect the grip of an object by the target person TG. In the grip detection process, the target person TG shown in the second captured image IMG2_TG is identified based on the ReID feature of the target person TG received from the contact detection unit 12 and the second captured image IMG2_TG received from the image identification unit 14, and further, 2D position information of the hand HD of this target person TG is acquired. The 2D position information of the hand HD includes position information of a bounding box assigned to surround the hand HD.
[0030] In the grip detection process, 2D position information of an object shown in the second captured image IMG2_TG is extracted based on the detection information DET2 received from the object detection unit 13 and the second captured image IMG2_TG received from the image specification unit 14. The 2D position information of the object is 2D position information of a detection target included in the detection information DET2, other than the person HM.
[0031] In the grasp detection process, an object having 2D position information matching the 2D position information of the hand HD is further searched for based on 2D position information of the hand HD of the target person TG captured in the second captured image IMG2_TG and 2D position information of the object captured in the second captured image IMG2_TG. This search is performed, for example, by calculating the degree of overlap between the 2D positions of the hand HD and the object and comparing it with a preset threshold. If the degree of overlap of the 2D positions is equal to or greater than the threshold, the object having the 2D position used in this overlap comparison is identified as the object having 2D position information matching the 2D position information of the hand HD. If such an object is identified, it is determined that the target person TG has grasped the object. If not, it is determined that the target person TG is not grasping the object.
[0032] 2-2. Second configuration example Fig. 3 is a block diagram showing a second configuration example of the system according to the embodiment. The configuration example shown in Fig. 3 includes a camera CA and a data processing device 10, similar to the configuration example shown in Fig. 2. The configuration of the camera CA and the basic configuration of the data processing device 10 are as described in Fig. 2.
[0033] In the configuration example shown in Fig. 3, the functional blocks of the data processing device 10 are depicted as object detection unit 13, object detection unit 11, contact detection unit 12, image identification unit 14, grip detection unit 15, object detection unit 16, and grip detection unit 17. The functional blocks from object detection unit 11 to grip detection unit 15 are as described in Fig. 2. The object detection unit 16 and grip detection unit 17 are functional blocks for the third detection process. Like the functional blocks from object detection unit 11 to grip detection unit 15, these functional blocks are realized, for example, by cooperation between the processing circuit and storage device of the data processing device 10.
[0034] The object detection unit 16 performs an object detection process to detect an object from the third captured image IMG3. The third captured image IMG3 is, for example, an image captured by the same camera CA that captured the first captured image IMG1, and is an image captured after the capture time (imaging time) of the first captured image IMG1. The content of the object detection process by the object detection unit 16 is basically the same as that of the object detection process by the object detection unit 11. In the object detection process by the object detection unit 16, detection information DET3 of the detection target is generated. The detection information DET3 includes, for example, ID information of the third captured image IMG3, timestamp information, image information of the detection target, type information of the detection target, and 2D position information of the detection target. The detection information DET3 is transmitted to the grip detection unit 17.
[0035] The grip detection unit 17 performs a grip detection process to detect the grip of an object by the target person TG. The grip detection process by the grip detection unit 15 directly detects the grip of an object by the target person TG, whereas the grip detection process by the grip detection unit 17 indirectly detects this grip, which is different between the two. The grip detection process by the grip detection unit 17 uses 2D position information of the contacting object and 2D position information of the detected object included in the detection information DET3 received from the object detection unit 16. Here, the "contacting object" is an object that has been detected to be in contact with the target person TG in the contact detection process by the contact detection unit 12. The 2D position information of the contacting object is identified by the contact detection unit 12 and transmitted to the grip detection unit 17.
[0036] In the grip detection process by the grip detection unit 17, it is determined whether or not the 2D position information of the detected object included in the detection information DET3 includes the 2D position information of the contacting object. If the 2D position information of the detected object included in the detection information DET3 includes the 2D position information of the contacting object, it means that the contacting object is still present at that 2D position. On the other hand, if the 2D position information of the contacting object is not included, it means that the contacting object is not present at that 2D position. In the grip detection process by the grip detection unit 17, it detects the grip of the contacting object by the target person TG based on the presence or absence of such 2D position information. The grip detection unit 17 transmits the grip detection result of the contacting object to the grip detection unit 15.
[0037] When the grip detection unit 15 receives the contacting object grip detection result from the grip detection unit 17, the grip detection unit 15 makes a final judgment about the grip of the contacting object by the target person TG based on the grip detection result and the judgment result of the grip detection process it has performed. For example, if the grip detection result of the contacting object by the grip detection unit 17 is "grasped" and the judgment result by the grip detection unit 15 is "grasped", it is judged that the target person TG has gripped the contacting object. If the grip detection result of the contacting object by the grip detection unit 17 is "grasped" and the judgment result by the grip detection unit 15 is "not gripped", it is judged that the target person TG is not gripping the contacting object. If the grip detection result of the contacting object by the grip detection unit 17 is "not gripped" and the judgment result by the grip detection unit 15 is "grasped", it is judged that the target person TG is not gripping the contacting object.
[0038] 3. First detection process using 3D position information Fig. 4 is a diagram illustrating an example of the first detection process using 3D position information. Fig. 4 illustrates functional blocks for the first detection process, including an object detection unit 21, a depth estimation unit 22, a 3D space generation unit 23, a posture estimation unit 24, a hand configuration estimation unit 25, and a contact detection unit 26. These functional blocks are realized, for example, by cooperation between a processing circuit and a storage device of the data processing device 10.
[0039] The object detection unit 21 performs an object detection process to detect an object from the first captured image IMG1. The content of the object detection process by the object detection unit 21 is basically the same as that of the object detection process by the object detection unit 11 described with reference to FIG. 2. However, in the object detection process by the object detection unit 21, detection of a person HM is not performed, and detection information DET1 including detection information of a dynamic object other than the person HM and detection information of a static object is generated. The detection information DET1 is transmitted to the 3D space generation unit 23.
[0040] The depth estimation unit 22 performs a depth generation process to generate depth information of an object captured in the first captured image IMG1. In the depth generation process, for example, the first captured image IMG1 (RGB image) is input to a machine learning model, and a depth image (distance image) is output. The depth information of the object captured in the first captured image IMG1 (i.e., distance information from the camera CA to the object) is generated based on the depth image output from the machine learning model. The depth information is transmitted to the 3D space generation unit 23.
[0041] The 3D space generation unit 23 performs a space generation process to generate a virtual target space VS that represents the real target space RS. The target space VS is expressed in the same world coordinate system (X, Y, Z) as the target space RS. In order to express the target space VS in the same world coordinate system (X, Y, Z) as the target space RS, a virtual camera corresponding to the camera CA that is installed in the target space RS is installed in the target space VS. The camera CA and the virtual camera have the same parameters, and these camera parameters have been calibrated in advance.
[0042] In the space generation process, a virtual object corresponding to an object existing in the target space RS is defined in the target space VS. The configuration of the virtual object (e.g., position, orientation, shape, size, etc.) is defined with reference to the detection information DET1 of the target object received from the object detection unit 21 and the depth information of the target object received from the depth estimation unit 22. Referring to Fig. 5, a product shelf GS is depicted on the left side of Fig. 5 as an object existing in the target space RS, and the target space VS in which a product shelf GS_VR is defined as a virtual object corresponding to the product shelf GS is depicted on the right side of Fig. 5.
[0043] The posture estimation unit 24 performs posture estimation processing to estimate the 3D posture (3D pose) of the person HM appearing in the first captured image IMG1. In the posture estimation processing, for example, a bounding box is assigned to the person HM appearing in the first captured image IMG1. Then, key points of the person HM are extracted from this bounding box to estimate the 3D posture of the person HM. The 3D posture is represented by lines connecting parts such as joints, head, hands, and feet. Note that the posture estimation processing is a well-known technique, and the method is not particularly limited. For example, MeTRAbs, TransPose, etc. are used in the posture estimation processing. Information on the 3D posture of the person HM is transmitted to the hand configuration estimation unit 25.
[0044] The hand configuration estimation unit 25 extracts information about the hand of the person HM from the 3D posture information of the person HM received from the posture estimation unit 24, and estimates the 3D configuration (position and size) of the hand. While the 3D posture of the person HM is represented by parts such as joints, head, hands, and feet and the tips connecting these parts, the 3D posture of the hand of the person HM is represented by the joints of the fingers of the hand and the lines connecting these joints. The hand configuration estimation unit 25 encloses the 3D posture of the entire hand of the person HM or the fingertips of some of the hand of the person HM in a 3D bounding box 3Dbbox, and transmits the 3D configuration information of this 3D bounding box 3Dbbox_HD to the contact detection unit 26.
[0045] The size of the bounding box 3D bbox differs when the 3D bounding box 3D bbox encloses the 3D pose of the entire hand of the person HM and when the 3D bounding box 3D bbox encloses the 3D pose of the fingertips of the hand of the person HM. This change in size is performed to reduce erroneous detection by the contact detection unit 26, which will be described later. The change in size is performed, for example, by determining whether or not the 3D pose of the hand of the person HM corresponds to a pose indicating a grasping state. If it is determined that the pose corresponds to a pose indicating a grasping state, a bounding box 3D bbox of a size that encloses the 3D pose of the entire hand is set. If it is determined that the pose does not correspond to a pose that indicates a grasping state (a pointing state in the example of FIG. 4), a bounding box 3D bbox of a size that encloses the 3D pose of the fingertips is set.
[0046] The contact detection unit 26 performs contact detection processing based on the 3D configuration information of the virtual object defined in the target space VS and the 3D configuration information of the hand of the person HM received from the hand configuration estimation unit 25. In the contact detection processing, for example, the index IoU (Intersection over Union) is used to calculate the overlap between the 3D bounding box surrounding the hand of the person HM and the 3D bounding box assigned to the virtual object.
[0047] Referring to Fig. 5, on the left side of Fig. 5, a product GD displayed on a product shelf GS is depicted as an object existing in a target space RS. On the right side of Fig. 5, a product GD_VR is depicted as a virtual object corresponding to the product GD, and each of these virtual objects is assigned a 3D bounding box 3Dbbox_GD. 3D configuration (position and size) information of the 3D bounding box 3Dbbox_GD surrounding the product GD_VR is calculated from detection information DET1 and depth information of the product GD.
[0048] The index IoU is calculated as the overlapping volume of the 3D bounding box 3Dbbox_GD surrounding the product GD_VR and the 3D bounding box 3Dbbox_HD surrounding the hand of the person HM. If there is a product GD_VR whose index IoU is equal to or greater than a threshold, the person HM is identified as the target person TG who has come into contact with this product GD_VR. Furthermore, the product GD_VR that has come into contact with the target person TG is identified as the contacting object described above. [Explanation of symbols]
[0049] 10...data processing device, 11, 13, 16, 21...object detection unit, 12, 26...contact detection unit, 14...image identification unit, 15, 17...grasp detection unit, 22...depth estimation unit, 23...3D space generation unit, 24...posture estimation unit, 25...hand configuration estimation unit, CA, CA1, CA2...camera, GD, GD1, GD2, GD3, GD_VR...product, GS, GS_VR...product shelf, HD...hand, HM...person, TG...target person, RS, VS...target space, IMG1...first captured image, IMG2...second captured image, IMG3...third captured image, DET1, DET2, DET3...detection information
Claims
1. A system for detecting the grasping of an object by a person, comprising: a storage device that stores a first captured image of the target space and a second captured image that is captured in the target space after the first captured image; a processing circuit that performs a first detection process using the first captured image and a second detection process using a detection result of the first detection process and the second captured image, The first detection process acquiring position information of a person and an object recognized from the first captured image; detecting contact between the person and the object recognized from the first captured image based on position information of the person and the object in the first captured image; The second detection process Identifying the second captured image in which a person identical to a target person representing a person whose contact with an object was detected in the first detection process is recognized; acquiring position information of the target person's hand and an object recognized from the specified second captured image; Detecting, based on position information of the target person's hand and an object in the specified second captured image, the target person's grasp of a contacting object indicating an object whose contact with the target person was detected in the first detection process; and 1. An object grasp detection system comprising:
2. 10. The system of claim 1, Detecting the grip of the contact object by the target person in the second detection process, determining that the target person is holding the contact object when the position information of the object in the specified second captured image includes information of an object recognized at the position of the target person's hand in the specified second captured image; and determining that the target person is not holding the contact object when the position information of the object in the specified second captured image does not include information of an object recognized at the position of the target person's hand in the specified second captured image; and 1. An object grasp detection system comprising:
3. 3. The system according to claim 1 or 2, The storage device further stores a third captured image that is captured in the target space after the first captured image, the processing circuit further performs a third detection process using the detection result of the first detection process and the third captured image; The third detection process acquiring position information of an object recognized from the third captured image; indirectly detecting the target person's grasping of the contacting object based on position information of the contacting object acquired in the first detection process and position information of an object recognized from the third captured image, Detecting movement of the contact object by the target person in the third detection process, determining that the target person is holding the contacting object when the position information of the contacting object is not included in the position information of the object recognized from the third captured image; determining that the target person is not holding the contacting object when the position information of the object recognized from the third captured image includes the position information of the contacting object; 1. An object grasp detection system comprising:
4. 3. The system according to claim 1 or 2, acquiring position information of the person recognized from the first captured image in the first detection process includes acquiring position information of a hand of the person recognized from the first captured image, detecting contact between a person and an object recognized from the first captured image in the first detection process, Calculating an overlap between the position of the hand recognized from the first captured image and the position of the object recognized from the first captured image; If there is an object whose overlapping degree is equal to or greater than a threshold, detecting that the object whose overlapping degree is equal to or greater than the threshold has come into contact with the person recognized from the first captured image; 1. An object grasp detection system comprising:
5. 3. The system according to claim 1 or 2, In the first detection process, acquiring position information of the person and the object recognized from the first captured image is performed in a three-dimensional virtual space corresponding to the target space, the three-dimensional virtual space being generated based on two-dimensional position information and depth information of the object recognized from the first captured image. An object grasp detection system comprising:
Citation Information
Patent Citations
Crime prevention support system
JP2005347905A
Method for recognizing object gripped by grip means
JP2010244413A
Projection device, projection method and computer program for projection
JP2017112565A
Information processing program, information processing method, and information processing apparatus
JP2023050826A
Detecting device, detecting system, detecting method, and detecting program
JP2022165483A