Information processing program, information processing method, and information processing device

JP7726391B2Active Publication Date: 2025-08-20FUJITSU LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2024522875
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2025-08-20
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

Conventional ROI acquisition methods for three-dimensional structures result in either overly wide or overly narrow regions of interest, leading to false detections or missed detections during inspections, particularly in the context of steel towers and bridges, due to inaccuracies in skeletal estimation and variations in work performed by individuals.

Method used

An information processing system that generates three-dimensional skeletal information from captured images, sets a first region of interest based on design data, and adjusts this region using hand part trajectory to define a more accurate second region of interest, leveraging existing object detection and skeletal estimation technologies.

Benefits of technology

Enables the acquisition of a more appropriate region of interest, reducing false positives and negatives in work detection by accurately identifying the object being used by individuals in the images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007726391000007
    Figure 0007726391000007
  • Figure 0007726391000008
    Figure 0007726391000008
  • Figure 0007726391000009
    Figure 0007726391000009
Patent Text Reader

Abstract

This information processing device: generates three-dimensional first skeleton information of a first person in a captured image by recognizing a skeleton of the first person included in a first video; sets, on the basis of pre-stored design data defining an area of an object present in a three-dimensional space, a first area of interest in the three-dimensional space in the captured image, for each of a plurality of objects in the captured image included in the first video of the captured image; using position information of a hand part in the first skeleton information of the captured image and the first area of interest in the captured image, identifies a first object that the first person in the captured image is using, from the plurality of objects in the captured image included in the first video of the captured image; acquires a position distribution of the hand part in the captured image in the three-dimensional space in the captured image, on the basis of a trajectory of movement of the hand part in the captured image in the first skeleton information of the captured image; and sets, on the basis of the position distribution in the captured image in which the hand part in the captured image is present, a second area of interest in the three-dimensional space in the captured image, the second area of interest indicating the use of the first object in the captured image by the person.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing program, an information processing method, and an information processing device. [Background technology]

[0002] When constructing three-dimensional structures such as steel towers and bridges in the steel structure industry, processes such as component processing, welding, and temporary assembly may be carried out between design using CAD (Computer Aided Design) and on-site assembly. In this case, after the temporary assembly process, the components are shipped to the site and assembled on-site. In the component processing process, the steel components of the three-dimensional structure are processed. In the welding process, the steel components are temporarily welded, the temporary welds are inspected, and the actual welding is performed. In the temporary assembly process, temporary assembly, disassembly, painting finish, and inspection are performed.

[0003] Because these three-dimensional structures are often large, one-off items, diagnosis during the welding and pre-assembly processes is often carried out by visual inspection, comparing the pre-welded or painted object with the 3D CAD model created during the design process.

[0004] If a defect in the welding process is overlooked during the diagnosis and is discovered during the temporary assembly work, a return to the temporary assembly process will be required, such as from the temporary assembly process to the component processing process. Also, if a defect in the temporary assembly process is overlooked during the diagnosis and is discovered during the on-site assembly work, a return to the on-site assembly process will be required.

[0005] Therefore, to prevent rework during temporary assembly and on-site assembly, a certain amount of work time is spent visually comparing the target object with the 3D CAD model. In this case, as the number of components increases, the number of locations to be inspected in the three-dimensional structure also increases, resulting in increased visual inspection time. Furthermore, it is not easy to train skilled workers who can perform highly accurate inspections in a short amount of time, and the training period for such skilled workers can sometimes last several years. This problem is not limited to inspections during construction work on steel towers, bridges, etc., but also arises in inspection work that involves comparing other three-dimensional structures with models represented by their model information.

[0006] Therefore, to streamline the confirmation process of comparing a three-dimensional structure with a model represented by model information, there is a technology that, for example, matches the edges of an object in a captured image with multiple three-dimensional (3D) line segments included in the CAD of the object, and displays the CAD by overlaying it on the captured image.In addition, because the position and orientation of the object are determined when overlaying the CAD on the captured image, this information can be used to extract 3D rectangular information surrounding the object, which can be used as an ROI (Region of Interest), also known as the target region, region of interest, or region of interest.

[0007] In this way, for example, if ROIs for objects can be extracted from images, i.e., captured images, taken by cameras installed in a factory, it will be possible to analyze the work of workers on those objects. More specifically, for example, by extracting ROIs for people and the objects they are working on from captured images and measuring the work content, work location, and time spent on that work, it will be possible to implement measures such as optimal allocation of people and things, and optimization of work instructions.

[0008] FIG. 1 is a diagram showing an example of a conventional task recognition technology. As shown in FIG. 1, in the conventional technology, an image captured by a surveillance camera installed in a factory is used as an input image, and an information processing device estimates a person's skeletal information 5 from the input image. The person manually draws an ROI 6 of an object on which the person will perform the task, and the information processing device estimates the task content using the positional relationship between the ROI 6 and the estimated skeletal information of the person. More specifically, for example, as shown in FIG. 1, if the skeleton of the person's hand is contained within the ROI 6, the information processing device determines that the person is touching the object and performing a task on the object. [Prior art documents] [Patent documents]

[0009] [Patent Document 1] Japanese Patent Application Publication No. 2017-091078 [Patent Document 2] Japanese Patent Application Publication No. 2019-139570 [Patent Document 3] Japanese Patent Publication No. 2021-177399 [Patent Document 4] US Patent Application Publication No. 2020 / 0082544 [Patent Document 5] US Patent Application Publication No. 2017 / 0084044 Summary of the Invention [Problem to be solved by the invention]

[0010] However, for example, when acquiring ROI at the part level for an object in a captured image, sufficient information may not be acquired. Figure 2 is a diagram showing an example of acquiring ROI for an object in a captured image using conventional technology. Figure 2 shows skeletal information 5 of a person and an ROI 6 of an object on which the person is performing work, extracted from the captured image using conventional technology.

[0011] The example on the left side of Figure 2 is an example in which the ROI 6 of the target object is acquired using, for example, conventional object recognition technology. In this case, only a rough ROI 6, a rectangle that contains the target object, is acquired, and the acquired ROI 6 is too wide. If the ROI 6 is too wide, there is a possibility of false detection, in which, for example, a person's work on an object other than the target object, as indicated by the skeleton information 5, is determined to be work on the target object.

[0012] On the other hand, the example on the right side of Figure 2 is an example in which ROI6 of an object is acquired using, for example, CAD design data. In this case, ROI6 can be acquired by narrowing the object down to the component level, but the acquired ROI6 is too narrow. If ROI6 is too narrow, there is a possibility that work on the object cannot be detected due to, for example, variations in work performed by people indicated by skeletal information 5 or the accuracy of skeletal estimation.

[0013] As such, in conventional technology, the ROI for an object obtained from a captured image may be too wide or too narrow, resulting in erroneous detection or failure to detect the object, and ultimately the work performed on the object.

[0014] According to one aspect, an object is to provide an information processing program, an information processing method, and an information processing device that are capable of acquiring a more appropriate ROI of an object from a captured image. [Means for solving the problem]

[0015] In one aspect, the information processing program causes a computer to perform the following processes: generate first three-dimensional skeletal information of the person in the first captured image by performing skeletal recognition on the first person included in the first video; set a first area of interest in the three-dimensional space of the captured image for each of multiple captured image objects included in the first video based on pre-stored design data that defines the areas of objects existing in three-dimensional space; identify a first object being used by the person in the first captured image from among multiple captured image objects included in the first video using position information of the hand part in the skeletal information of the first captured image and the area of interest in the first captured image; obtain a positional distribution of the hand part in the three-dimensional space of the captured image based on the trajectory of movement of the hand part in the captured image in the skeletal information of the first captured image; and set a second area of interest in the three-dimensional space of the captured image that indicates that the person is using the object in the first captured image based on the positional distribution of the hand part in the captured image. [Effects of the Invention]

[0016] In one aspect, a more appropriate ROI of the object can be obtained from the captured image. [Brief explanation of the drawings]

[0017] [Figure 1] FIG. 1 is a diagram showing an example of a conventional task recognition technique. [Figure 2] FIG. 2 is a diagram showing an example of ROI acquisition for an object in a captured image according to the prior art. [Figure 3] FIG. 3 is a diagram illustrating an example of a configuration of the information processing system according to the first embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of the configuration of the information processing device 10 according to the first embodiment. [Figure 5] FIG. 5 is a flowchart illustrating a flow of the 3D ROI setting process according to the first embodiment. [Figure 6] FIG. 6 is a diagram showing an example of acquiring an initial image. [Figure 7] FIG. 7 is a diagram illustrating an example of 3D position and orientation estimation of a device. [Figure 8] FIG. 8 is a diagram showing an example of coordinate transformation of the device. [Figure 9] FIG. 9 is a diagram showing an example of a three-dimensional rectangle extracted from a part in a device. [Figure 10] FIG. 10 is a diagram illustrating an example of skeleton information. [Figure 11] FIG. 11 is a diagram illustrating an example of determining the posture of the entire body. [Figure 12] FIG. 12 is a diagram illustrating an example of relative 3D skeletal coordinate estimation. [Figure 13] FIG. 13 is a diagram showing an example of the positional relationship between 2D / 3D coordinate systems. [Figure 14] FIG. 14 is a diagram showing an example of a work schedule. [Figure 15] FIG. 15 is a diagram showing an example of setting a 3D ROI. [Figure 16] FIG. 16 is a flowchart illustrating a flow of the motion determination process using the 3D ROI according to the second embodiment. [Figure 17] FIG. 17 is a diagram showing an example of motion determination using a 3D ROI. [Figure 18] FIG. 18 is a diagram illustrating an example of the hardware configuration of the information processing device 10. As shown in FIG. DETAILED DESCRIPTION OF THE INVENTION

[0018] Examples of the information processing program, information processing method, and information processing device according to the present embodiment will be described in detail below with reference to the accompanying drawings. Note that the present embodiment is not limited to these examples. Furthermore, the examples can be combined as appropriate within a consistent range. [Example]

[0019] First, an information processing system for carrying out this embodiment will be described. Fig. 3 is a diagram showing a configuration example of the information processing system according to Example 1. As shown in Fig. 3, the information processing system 1 is a system in which an information processing device 10 and a camera device 100 are connected via a network 50 so as to be able to communicate with each other.

[0020] The network 50 may be any of a variety of communication networks, whether wired or wireless, such as an intranet used within a factory. The network 50 may not be a single network, but may instead be configured such that an intranet and the Internet are connected via a network device such as a gateway or other device (not shown).

[0021] The information processing device 10 is, for example, an information processing device such as a desktop PC (Personal Computer), a notebook PC, or a server computer that is installed in a factory and used by workers, managers, and the like.

[0022] The information processing device 10 receives, from the camera device 100, a plurality of images captured by the camera device 100 of a predetermined imaging range, such as a predetermined work area in a factory. Strictly speaking, the plurality of images are images captured by the camera device 100, i.e., a series of frames of a video.

[0023] Furthermore, the information processing device 10 uses, for example, an existing object detection technology to identify objects such as workers (hereinafter simply referred to as "people") working in a factory and the equipment on which the people work from the captured image. Furthermore, the information processing device 10 uses an existing skeleton detection technology to generate skeleton information of the people identified from the captured image. Furthermore, the information processing device 10 identifies a first ROI in a three-dimensional space of the object identified from the captured image based on design data that defines the area of an object existing in a three-dimensional space, such as CAD. Hereinafter, the ROI in the three-dimensional space may be referred to as a "3D ROI," and the first ROI in the three-dimensional space may be referred to as a "first 3D ROI" or a "first region of interest."

[0024] Furthermore, the information processing device 10 acquires a position distribution in three-dimensional space where the hand parts are present, for example, based on the trajectory of the movement of the hand parts of the generated skeletal information. Then, the information processing device 10 sets a second ROI in three-dimensional space that indicates the person's use of the object, for example, based on the acquired position distribution. Note that the second ROI in three-dimensional space may hereinafter be referred to as a "second 3D ROI" or a "second region of interest."

[0025] 3, the information processing device 10 is shown as a single computer, but may be, for example, a distributed computing system configured with multiple computers. Furthermore, the information processing device 10 may be a cloud computer device managed by a service provider that provides cloud computing services.

[0026] The camera device 100 is, for example, a surveillance camera installed in a factory. Although one camera device 100 is shown in Fig. 1, in reality, for example, multiple camera devices may be installed in each work area in the factory. Images captured by the camera device 100 are transmitted to the information processing device 10 at any time or at a predetermined timing.

[0027] [Functional configuration of information processing device 10] Next, a functional configuration of the information processing device 10 that executes the present embodiment will be described. Fig. 4 is a diagram illustrating an example of the configuration of the information processing device 10 according to Example 1. As shown in Fig. 4, the information processing device 10 includes a communication unit 20, a storage unit 30, and a control unit 40.

[0028] The communication unit 20 is a processing unit that controls communication with other devices such as the camera device 100, and is a communication interface such as a USB (Universal Serial Bus) interface or a network interface card.

[0029] The storage unit 30 has a function of storing various data and programs executed by the control unit 40, and is realized by a storage device such as a memory or a hard disk. For example, the storage unit 30 stores a 3D object model for estimating the posture of an object that a person works on, a 3D human skeleton model for estimating three-dimensional skeletal information of a person, and the like.

[0030] The storage unit 30 also stores a plurality of captured images, which are a series of frames captured by the camera device 100. The storage unit 30 can also store position information within the captured images of people and objects identified from the captured images. The storage unit 30 also stores two-dimensional skeleton information of people identified from the captured images captured by the camera device 100.

[0031] The above information stored in the storage unit 30 is merely an example, and the storage unit 30 can store various information other than the above information.

[0032] The control unit 40 is a processing unit, such as a processor, that controls the entire information processing device 10. The control unit 40 includes an image acquisition unit, an object posture estimation unit, a target object identification unit, a human region detection unit, a 3D skeleton estimation unit, a target region adjustment unit, an ROI determination unit, etc. Each processing unit is an example of an electronic circuit included in the processor or an example of a process executed by the processor.

[0033] The image acquisition unit acquires from the camera device 100 a plurality of captured images, which are a series of frames captured by the camera device 100 .

[0034] The object pose estimation unit, for example, uses existing technology to extract edges of the image captured by the camera device 100. The object pose estimation unit also estimates the position and orientation in three-dimensional space of each object that may be a target of work by a person, for example, using 3D line segments of the object 3D model and corresponding edges in the captured image.

[0035] The human region detection unit uses, for example, an existing object detection algorithm to identify a person from an image captured by the camera device 100 and detect a bounding box that is the region of the person. Here, the existing object detection algorithm may be, for example, an object detection algorithm such as YOLO (You Only Look Once), SSD (Single Shot Multibox Detector), or Faster R-CNN (Convolutional Neural Network) that uses deep learning.

[0036] The 3D skeleton estimation unit uses existing technology such as a Cascaded Pyramid Network (CPN) to estimate the posture of a person from a partial image within a bounding box of the person identified from a captured image, and generates 2D skeleton information. The 3D skeleton estimation unit also uses, for example, existing technology, to estimate the 3D skeleton coordinates of each part of the person's 2D skeleton information relative to a reference position, such as the waist. Note that the 3D skeleton coordinates estimated by the 3D skeleton estimation unit are relative coordinates normalized from a reference position, and the dimensions between each part are not actual size. Therefore, in this embodiment, the scale of the person is estimated. Furthermore, since the 3D skeleton coordinates estimated by the 3D skeleton estimation unit are relative coordinates from a reference position, in this embodiment, absolute coordinates relative to world coordinates are calculated.

[0037] The 3D skeleton estimation unit also converts the person's relative 3D skeleton coordinates into absolute 3D skeleton information relative to world coordinates, for example, using a homography transformation matrix. The homography transformation matrix may be calculated, for example, based on the coordinates of four different points in a captured image of a factory and the world coordinates corresponding to each of the four points. The absolute 3D skeleton information may be calculated by calculating absolute 3D skeleton information for several parts, such as the waist and right foot, and using the calculated absolute coordinates of the waist and right foot to calculate the absolute coordinates of other parts. Hereinafter, the estimated absolute 3D skeleton information may be simply referred to as "3D skeleton information."

[0038] The target object identifying unit identifies a target object from among the objects whose positions and orientations have been estimated, for example, using the estimated 3D skeleton information, and identifies a first 3D ROI for the identified object.

[0039] The target region adjustment unit corrects the identified first 3D ROI using, for example, the estimated 3D skeletal information to adjust it to a more appropriate ROI. This is a process of correcting the first 3D ROI based on the position distribution of the hand parts in the 3D skeletal information, since the first 3D ROI identified using, for example, CAD design data may be too narrow compared to the actual working area.

[0040] The ROI determination unit is, for example, a process of determining and setting the corrected first 3D ROI as the second 3D ROI indicating that a person uses the object.

[0041] [Function details] Next, the 3D ROI setting process executed by the information processing device 10 will be described in detail with reference to Figs. 5 to 17. Fig. 5 is a flowchart showing the flow of the 3D ROI setting process according to the first embodiment.

[0042] 5, the information processing device 10 acquires from the storage unit 30 a captured image of a predetermined imaging range, such as a predetermined work area in a factory, captured by the camera device 100, and acquires an initial image (step S101). Strictly speaking, the captured image acquired from the storage unit 30 is a video of a series of operations in the predetermined work area (hereinafter, sometimes referred to as "training data").

[0043] The initial image acquired in step S101 is, for example, an initial frame in which the device is captured, acquired from a video that is training data. Fig. 6 is a diagram showing an example of initial image acquisition. An index 7 for determining the camera posture is captured in the initial image, and a specific position of the index 7 is set as the origin of the world coordinates.

[0044] Returning to the explanation of FIG. 5, next, the information processing device 10 estimates the position and orientation of the camera device 100 with respect to the world coordinate system using the index 7 by, for example, existing technology (step S102). More specifically, for example, the information processing device 10 estimates the position and orientation of the camera device 100 from a pair of the three-dimensional coordinates of each corner of the index 7 and the corresponding two-dimensional coordinates in the captured image. At this time, the position and orientation of the camera device 100 is obtained in the form of a 3×3-dimensional rotation matrix and a three-dimensional translation vector. Furthermore, the position and orientation of the camera device 100 is estimated, for example, by using the DLT method, which is one of the existing methods for estimating a homography matrix, and a homography matrix H lg (3 rows, 3 columns) is estimated.

[0045] Next, the information processing device 10 reads a 3D model of the device from, for example, the object 3D model stored in the storage unit 30 (step S103). Here, the read 3D model of the device may be, for example, a 3D model of a device that can be a work target according to the predetermined work area captured by the initial image acquired in step S101.

[0046] Fig. 7 is a diagram showing an example of 3D position and orientation estimation of a device. The image on the left side of Fig. 7 is an example of an image of a 3D model of the device loaded in step S103. Note that the data format of the loaded 3D model may be, for example, an STL / OBJ format that represents a three-dimensional shape by connecting triangular meshes, as shown on the left side of Fig. 7. Furthermore, the information processing device 10 determines the connection between meshes from the similarity of the slopes of the line segments that make up adjacent triangular meshes, and extracts the edge lines of the 3D model of the device as 3D line segments.

[0047] Returning to the explanation of FIG. 5, next, the information processing device 10 estimates, for example, the position and orientation in three-dimensional space of the device shown in the initial image acquired in step S101 (step S104). More specifically, as shown on the right side of FIG. 7, for example, the information processing device 10 extracts edges of the initial image acquired in step S101 using existing technology, and estimates the position and orientation of the device using the edges and the 3D line segments of the 3D model extracted in step S103. Note that the position and orientation of the device is also obtained in the form of a 3×3-dimensional rotation matrix and a 3-dimensional translation vector, and the example on the right side of FIG. 7 is an example in which a 3D model is superimposed on the initial image using the estimated rotation matrix and translation vector.

[0048] Returning to the explanation of FIG. 5, next, the information processing device 10 converts, for example, the relative coordinates of the device whose position and orientation have been estimated in step S104 into world coordinates (step S105). FIG. 8 is a diagram showing an example of coordinate conversion of the device. As shown in FIG. 8, the world coordinates of the device (object) can be calculated from the relationship between the relative coordinates of the camera device 100 (camera) and the world coordinates (world) using the following formula (1).

[0049]

number

[0050] In formula (1), X w are the world coordinates of the device whose position and orientation have been estimated in step S104. X0 is the relative coordinate of the device whose position and orientation have been estimated, and can be, for example, a four-dimensional coordinate with 1 added to the fourth dimension, that is, a homogeneous coordinate. W O is the matrix for converting the device relative coordinates into coordinates in the world coordinate system, and P W c is a matrix for converting the relative coordinates of the camera device 100 into coordinates in the world coordinate system. c O is a matrix for converting the relative coordinates of the device into coordinates in the relative coordinate system of the camera device. The matrix P is expressed, for example, by the following equation (2).

[0051]

number

[0052] In equation (2), r 11 ~r 33 The part is the rotation matrix, t x ~t z The part is a translation vector.

[0053] Returning to the description of FIG. 5, next, the information processing device 10 extracts, for example, three-dimensional rectangles of the components in the device whose positions and orientations were estimated in step S104 (step S106). FIG. 9 is a diagram showing an example of three-dimensional rectangle extraction of components in the device. As shown in FIG. 9, for example, CAD design data 8 stores information about the components in the device in a nested hierarchical structure for each device. Therefore, the information processing device 10 can extract three-dimensional rectangles 9 of each component in the device by acquiring the endpoints, i.e., relative coordinates, of components in a specific layer, such as the fourth layer. Furthermore, since the world coordinates of the device were obtained in step S105, the information processing device 10 calculates, for example, the world coordinates of the endpoints of the three-dimensional rectangles 9 of each component based on the world coordinates of the device.

[0054] 5, next, the information processing device 10 extracts a rectangular area surrounding the person, i.e., a bounding box, from the captured image using, for example, an existing object detection algorithm (step S107). Furthermore, since the processing from step S107 onwards is repeatedly performed for each frame of the video, which is training data, capturing a series of tasks, the captured image at the first time is the initial image acquired in step S101.

[0055] If a person area is not extracted from the captured image in step S107 and it is determined that no person is present in the captured image (step S108: No), the information processing device 10 reads, for example, the next frame of the video that is the training data (step S109), and then repeats the process from step S107 for the next frame.

[0056] On the other hand, if a person area is extracted from the captured image and it is determined that a person is present in the captured image (step S108: Yes), the information processing device 10 estimates two-dimensional skeletal information of the person using, for example, an existing skeletal estimation algorithm (step S110). Here, the existing skeletal estimation algorithm is, for example, a skeletal estimation algorithm that uses deep learning, such as HumanPoseEstimation, such as DeepPose or OpenPose.

[0057] Regarding estimation of two-dimensional skeletal information, for example, the information processing device 10 can acquire the two-dimensional skeletal information by inputting image data (each frame) into a trained machine learning model. FIG. 10 is a diagram illustrating an example of skeletal information. The two-dimensional skeletal information can use 18 pieces of definition information (numbered from 0 to 17) in which each joint identified by a known skeletal model is numbered. For example, the right shoulder joint (SHOULDER_RIGHT) is assigned number 7, the left elbow joint (ELBOW_LEFT) is assigned number 5, the left knee joint (KNEE_LEFT) is assigned number 11, and the right hip joint (HIP_RIGHT) is assigned number 14. Therefore, the information processing device 10 can acquire coordinate information of the 18 skeletal joints shown in FIG. 10 from the image data. For example, the information processing device 10 acquires "X coordinate = X7, Y coordinate = Y7, Z coordinate = Z7" as the position of the right shoulder joint (number 7). For example, the Z axis can be defined as the distance direction from the imaging device toward the object, the Y axis as the height direction perpendicular to the Z axis, and the X axis as the horizontal direction.

[0058] The information processing device 10 can also determine whole-body postures, such as standing, walking, crouching, sitting, and lying down, using a machine learning model that has been trained in advance on skeletal patterns. For example, the information processing device 10 can determine the closest whole-body posture by using a machine learning model that has been trained with a Multilayer Perceptron on skeletal information and some joints and angles between joints, such as the beauty and technique diagram in FIG.

[0059] Fig. 11 is a diagram showing an example of whole-body posture determination. As shown in Fig. 11, the information processing device 10 can detect the whole-body posture by acquiring (a) the angle of the joint between No. 10 "HIP_LEFT" and No. 11 "KNEE_LEFT," (b) the angle of the joint between No. 14 "HIP_RIGHT" and No. 15 "KNEE_RIGHT," (c) the angle of No. 11 "KNEE_LEFT," and (d) the angle of No. 15 "KNEE_RIGHT."

[0060] In addition, the information processing device 10 may estimate posture using a machine learning model such as a Multilayer Perceptron, which is generated by machine learning using some joints and angles between the joints as features and whole-body postures such as standing and crouching as correct labels.

[0061] Furthermore, the information processing device 10 may use, as a posture estimation algorithm, 3D Pose Estimation such as VNect, which estimates a three-dimensional posture from a single captured image. Furthermore, the information processing device 10 may estimate a posture from three-dimensional joint data, for example, using 3d-pose-baseline, which generates three-dimensional joint data from two-dimensional skeletal information.

[0062] Furthermore, the information processing device 10 may estimate the posture of a person by identifying the motion of each body part based on the orientation of each body part, such as the face, arms, elbows, etc., and the angle at which each body part is bent. Note that the algorithm for posture estimation and skeleton estimation is not limited to one type, and the posture and skeleton may be estimated in a composite manner using multiple algorithms.

[0063] 5, the information processing device 10 then estimates, for example, using existing technology, global three-dimensional skeletal coordinates in three-dimensional space of each part of the person's two-dimensional skeletal information relative to a reference position such as the waist (step S111). More specifically, first, the information processing device 10 estimates relative three-dimensional skeletal coordinates of each part of the person's two-dimensional skeletal information relative to a reference position such as the waist.

[0064] Fig. 12 is a diagram showing an example of relative 3D skeletal coordinate estimation. As shown in Fig. 12, relative 3D skeletal coordinate estimation involves (1) inputting an image of a person, (2) finding true values of each part using, for example, the position of the waist as a reference, and (3) estimating the relative 3D coordinates of each part from the waist.

[0065] Then, the information processing device 10 converts the estimated relative 3D skeleton coordinates into absolute 3D skeleton information relative to world coordinates using, for example, a homography transformation matrix. The homography transformation matrix is calculated by the information processing device 10 using an existing technique such as the DLT (Direct Linear Transformation) method. The homography transformation matrix is expressed by the following equation (3).

[0066]

number

[0067] In equation (3), u and v indicate input two-dimensional coordinates, and x, y, and 1 (z value) indicate transformed three-dimensional coordinates. Then, the information processing device 10 estimates the three-dimensional coordinates x and y of the feet from, for example, a homography transformation matrix. FIG. 13 is a diagram showing an example of the positional relationship of coordinate systems in 2D / 3D. FIG. 13 shows the positional relationship of each coordinate system. As shown in FIG. 13, the three-dimensional coordinates x and y of the feet are calculated, for example, from the two-dimensional coordinates (u ra , v ra ) and (u la , v la ) are converted into three-dimensional coordinates using a homography transformation matrix and calculated. The calculation of the three-dimensional coordinates of the right foot and the left foot is represented by, for example, the following equations (4) and (5), respectively.

[0068]

number

[0069]

number

[0070] In equations (4) and (5), H lg is the homography transformation matrix shown in Equation (3). For example, as shown in FIG. 13, the x and y coordinates of the waist can be determined as the midpoint between both legs. The z coordinate of the waist can be determined as, for example, l leg can be set as a fixed value.

[0071] Then, the information processing device 10 calculates similarity transformation parameters s from the local three-dimensional skeletal coordinates to the global three-dimensional skeletal coordinates using, for example, an existing technique such as Procrustes Analysis so that the transformed coordinates are closest to each other. lg , R lg , T lg Furthermore, the information processing device 10 estimates the estimated similarity transformation parameters s lg , R lg , T lg Each of the local three-dimensional skeleton coordinates is converted into a global three-dimensional skeleton coordinate using the following formula (6):

[0072]

number

[0073] 5, if there is a next frame of the video that is the training data (step S112: Yes), the information processing device 10, for example, reads the next frame (step S109), and repeats the process from step S107 for the next frame. Note that the estimated global three-dimensional skeleton coordinates may be stored in the storage unit 30, for example, for each frame.

[0074] On the other hand, if there is no next frame (step S112: No), the information processing device 10 identifies an object of interest from among the parts whose three-dimensional rectangles were extracted in step S106 (step S113). Identification of the object of interest may be performed based on a work schedule that indicates which work is to be performed on which part for each work period. The object of interest identified here is the first 3D ROI.

[0075] FIG. 14 is a diagram showing an example of a work schedule. The work schedule shown in FIG. 14 is data in which work content to be performed for each work time period and work components corresponding to the work content are linked and stored. To generate the work schedule shown in FIG. 14, for example, for each work time period, the distance between the positions of both hands in the global 3D skeleton coordinates and the center of the three-dimensional rectangle of each component is calculated, and the three-dimensional rectangle with the closest sum of distances is registered as the work component, thereby generating the work schedule. In other words, the work component registered as a three-dimensional rectangle for each work time period may be identified as the target object for that work time period. Note that in step S113, the work time period can be determined based on the elapsed time associated with each frame of the video, which is the training data, or the like.

[0076] Returning to the description of FIG. 5, next, the information processing device 10 corrects, for example, the first 3D ROI, which is the object of interest, with the second 3D ROI (step S114). FIG. 15 is a diagram showing an example of setting a 3D ROI. In FIG. 15, a three-dimensional rectangle 9 is the object of interest and the first 3D ROI. For example, as shown in FIG. 15, the information processing device 10 plots a sphere 15 of radius r, such as 5 cm (centimeters), in a three-dimensional space with respect to the positions of both hands in the global 3D skeletal coordinates for each task time. Then, the information processing device 10 calculates a union 16 of the three-dimensional space containing the sphere 15 for each task time, and sets the union 16 as a second 3D ROI indicating that the person uses the object of interest represented by the three-dimensional rectangle 9.

[0077] That is, for example, when the information processing device 10 determines whether or not the target object has been used, it determines whether or not the positions of both hands in the global 3D skeletal coordinates are contained in the union 16 set as the second 3D ROI, rather than whether or not they are contained in the three-dimensional rectangle 9. This results in correcting the three-dimensional rectangle 9 set as the first 3D ROI with the union 16 set as the second 3D ROI, making it possible to acquire a more appropriate ROI for the target object, which is the target object, from the captured image.

[0078] After step S114 is executed, the 3D ROI setting process shown in Fig. 5 ends. However, at this time, the second 3D ROI set as a more appropriate ROI may be represented by a connection of planar triangles, regarded as a convex polyhedron whose vertices indicate the surface shape, and stored in the storage unit 30 in STL format or the like. [Example]

[0079] Next, a second embodiment will be described in which a person's movement is determined using the second 3D ROI set in the first embodiment. The configurations of the information processing system according to the second embodiment and the information processing device 10 that is the subject of the movement are the same as those shown in Fig. 3 and Fig. 4 in the first embodiment. Fig. 16 is a flowchart showing the flow of a movement determination process using a 3D ROI according to the second embodiment.

[0080] 16, the information processing device 10 reads, for example, the second 3D ROI set in the first embodiment (step S201). The second 3D ROI is, for example, the second 3D ROI set by correcting the first 3D ROI, which is the target object, in step S114 of the 3D ROI setting process according to the first embodiment shown in FIG.

[0081] Next, the information processing device 10 acquires and reads from the storage unit 30, for example, captured images of a predetermined imaging range, such as a predetermined work area in a factory, captured by the camera device 100 (step S202). Note that in the second embodiment, the captured images captured by the camera device 100, more specifically, the monitoring video, are processed in real time, so the captured images are transmitted from the camera device 100 at any time and stored in the storage unit 30.

[0082] Next, the information processing device 10 extracts a rectangular area surrounding the person, that is, a bounding box, from the captured image read in step S202 using, for example, an existing object detection algorithm (step S203).

[0083] If a person area is not extracted from the captured image in step S203 and it is determined that no person is present in the captured image (step S204: No), the information processing device 10 returns to step S202, and reads, for example, the next frame of the monitoring video (step S202). Then, the process is repeated from step S203 for the next frame.

[0084] On the other hand, if a person area is extracted from the captured image and it is determined that a person is present in the captured image (step S204: Yes), the information processing device 10 estimates two-dimensional skeletal information of the person using, for example, an existing skeletal estimation algorithm (step S205). The two-dimensional skeletal information estimation process in step S205 is similar to the two-dimensional skeletal information estimation process in step S110 of the 3D ROI setting process according to the first embodiment shown in FIG.

[0085] Next, the information processing device 10 estimates, for example, global three-dimensional skeletal coordinates in three-dimensional space of each part of the person's two-dimensional skeletal information estimated in step S205, for example, relative to a reference position such as the waist, using an existing technology (step S206). The three-dimensional skeletal coordinate estimation process in step S206 is similar to the three-dimensional skeletal coordinate estimation process in step S111 of the 3D ROI setting process according to the first embodiment shown in FIG.

[0086] Next, the information processing device 10 determines the motion of the person in the captured image, for example, based on whether the positions of both hands in the global 3D skeletal coordinates of the person in the captured image are included in the second ROI read in step S201 (step S207). For example, if the positions of both hands in the global 3D skeletal coordinates are included in the second ROI, the information processing device 10 can determine that the person in the captured image is using, for example, an object such as a part corresponding to the second ROI. The motion determination in step S207 will be described in more detail.

[0087] Fig. 17 is a diagram showing an example of motion determination using a 3D ROI. Fig. 17 shows measurement of the positional relationship of 3D coordinate points to be determined, such as the coordinates of a person's right hand in a captured image, for each of the triangular meshes 17 of the loaded second ROI.

[0088] More specifically, the information processing device 10, for example, calculates a normal vector for each mesh 17, and calculates the angle formed between the vector of the right hand coordinate and the normal vector from the center of gravity position. Then, for example, if the angle formed with all meshes 17 of the loaded second ROI is 90 degrees or less, the information processing device 10 can determine that the right hand of the person in the captured image is included in the second ROI.

[0089] Returning to the explanation of FIG. 16, if there is a next frame of the monitoring video (step S208: Yes), the information processing device 10, for example, reads the next frame (step S202), and the process is repeated from step S203 for the next frame.

[0090] On the other hand, if there is no next frame (step S208: No), the motion determination process shown in FIG. 16 ends.

[0091] [effect] As described above, the information processing device 10 generates three-dimensional first skeletal information of the first person by performing skeletal recognition on the first person included in the first video, sets a first region of interest in three-dimensional space for each of a plurality of objects included in the first video based on pre-stored design data that defines the area of objects existing in three-dimensional space, identifies a first object being used by the first person from a plurality of objects included in the first video using the position information of the hand parts in the first skeletal information and the first region of interest, obtains a positional distribution of the hand parts in three-dimensional space based on the trajectory of movement of the hand parts in the first skeletal information, and sets a second region of interest in three-dimensional space indicating that the person is using the first object based on the positional distribution of the hand parts.

[0092] In this way, the information processing device 10 sets a 3D object ROI of the target object from the video based on design data that defines the area of the object existing in three-dimensional space, and acquires the position distribution of the hand parts based on the hand trajectory of the 3D skeletal information of the person included in the video to correct the 3D object ROI, thereby enabling the information processing device 10 to acquire a more appropriate ROI of the target object from the captured image.

[0093] In addition, the process of setting the second region of interest executed by the information processing device 10 includes a process of setting a region of a predetermined range centered on the position of the hand in three-dimensional space based on the position distribution of the hand parts, and setting a region in three-dimensional space that contains the region of the predetermined range as the second region of interest.

[0094] This allows the information processing device 10 to obtain a more appropriate ROI.

[0095] In addition, the process of setting the first region of interest executed by the information processing device 10 includes a process of setting a region in three-dimensional space as the first region of interest for each of multiple objects based on the endpoints of the objects included in the design data.

[0096] This allows the information processing device 10 to obtain a more appropriate ROI.

[0097] Furthermore, the information processing device 10 generates second three-dimensional skeletal information of the second person by performing skeletal recognition on the second person included in the second video, and determines whether the second person has used the first object based on the position information of the hand part of the second skeletal information and the second area of interest.

[0098] This allows the information processing device 10 to more accurately determine the movement of the person relative to the object.

[0099] [system] The information, including the processing procedures, control procedures, specific names, various data, and parameters shown in the above documents and drawings, may be changed as desired unless otherwise specified. Furthermore, the specific examples, distributions, and numerical values described in the embodiments are merely examples and may be changed as desired.

[0100] Furthermore, the specific form of distribution or integration of the components of each device is not limited to that shown in the drawings. That is, all or part of the components may be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc. Furthermore, all or any part of the processing functions of each device may be realized by a CPU (Central Processing Unit) and a program analyzed and executed by the CPU, or may be realized as hardware using wired logic.

[0101] [Hardware] Fig. 18 is a diagram illustrating an example of the hardware configuration of the information processing device 10. As shown in Fig. 18, the information processing device 10 includes a communication interface 10a, a hard disk drive (HDD) 10b, a memory 10c, and a processor 10d. The components shown in Fig. 18 are connected to each other via a bus or the like.

[0102] The communication interface 10a is a network interface card or the like, and performs communication with other information processing devices. The HDD 10b stores programs and data that operate the functions shown in FIG.

[0103] The processor 10d is a hardware circuit that operates a process that executes each function described in FIG. 4 and other figures by reading a program that executes the same processing as each processing unit shown in FIG. 4 from the HDD 10b or the like and expanding the program into the memory 10c. That is, this process executes the same functions as each processing unit of the information processing device 10. Specifically, the processor 10d reads a program having the same functions as the image acquisition unit, the object pose estimation unit, and the like from the HDD 10b or the like. Then, the processor 10d executes a process that executes the same processing as the image acquisition unit, the object pose estimation unit, and the like.

[0104] In this way, the information processing device 10 operates as an information processing device that executes operation control processing by reading and executing a program that executes the same processing as each processing unit shown in Fig. 4. The information processing device 10 can also realize functions similar to those of the above-described embodiment by reading a program from a recording medium using a medium reading device and executing the read program. Note that the program in these other embodiments is not limited to being executed by the information processing device 10. For example, this embodiment may also be applied in the same way to cases where another computer or server executes a program, or where these execute a program in cooperation with each other.

[0105] A program that executes the same processes as those of the processing units shown in Fig. 4 can be distributed via a network such as the Internet. This program can be recorded on a computer-readable recording medium such as a hard disk, a flexible disk (FD), a CD-ROM, a magneto-optical disk (MO), or a digital versatile disc (DVD), and can be executed by being read from the recording medium by a computer. [Explanation of symbols]

[0106] 1. Information Processing Systems 5. Skeletal information 6 ROI 7 indicators 8 Design Data 9 3D rectangle 10. Information processing equipment 10a communication interface 10b HDD 10c memory 10d processor 15 balls 16 Union 17 mesh 20 Communications Department 30 Storage section 40 Control Unit 50 Network 100 Camera Equipment

Claims

1. generating first three-dimensional skeletal information of a first person by performing skeletal recognition on the first person included in the first video; setting a first region of interest in the three-dimensional space for each of the plurality of objects included in the first image based on pre-stored design data that defines the region of the object existing in the three-dimensional space; identifying a first object being used by the first person from among the plurality of objects included in the first video using position information of a hand part of the first skeletal information and the first region of interest; acquiring a position distribution in the three-dimensional space where the hand part is present based on a trajectory of movement of the hand part of the first skeletal information; A second attention area in the three-dimensional space indicating that the person uses the first object is set based on the position distribution where the hand part is present. An information processing program that causes a computer to execute a process.

2. The process of setting the second region of interest includes: setting a region of a predetermined range centered on the position of the hand in the three-dimensional space based on the position distribution where the hand part is present; A region in the three-dimensional space that includes the region of the predetermined range is set as a second region of interest.

2. The information processing program according to claim 1, further comprising:

3. The process of setting the first region of interest includes: A region is set in the three-dimensional space as the first region of interest for each of the plurality of objects based on an end point of the object included in the design data.

2. The information processing program according to claim 1, further comprising:

4. generating second three-dimensional skeletal information of a second person by performing skeletal recognition on the second person included in the second video; Determine whether the second person has used the first object based on position information of a hand part of the second skeletal information and the second region of interest.

2. The information processing program according to claim 1, wherein the information processing program causes the computer to execute processing.

5. generating first three-dimensional skeletal information of a first person by performing skeletal recognition on the first person included in the first video; setting a first region of interest in the three-dimensional space for each of the plurality of objects included in the first image based on pre-stored design data that defines the region of the object existing in the three-dimensional space; identifying a first object being used by the first person from among the plurality of objects included in the first video using position information of a hand part of the first skeletal information and the first region of interest; acquiring a position distribution in the three-dimensional space where the hand part is present based on a trajectory of movement of the hand part of the first skeletal information; A second attention area in the three-dimensional space indicating that the person uses the first object is set based on the position distribution where the hand part is present. An information processing method characterized in that the processing is executed by a computer.

6. generating first three-dimensional skeletal information of a first person by performing skeletal recognition on the first person included in the first video; setting a first region of interest in the three-dimensional space for each of the plurality of objects included in the first image based on pre-stored design data that defines the region of the object existing in the three-dimensional space; identifying a first object being used by the first person from among the plurality of objects included in the first video using position information of a hand part of the first skeletal information and the first region of interest; acquiring a position distribution in the three-dimensional space where the hand part is present based on a trajectory of movement of the hand part of the first skeletal information; A second attention area in the three-dimensional space indicating that the person uses the first object is set based on the position distribution where the hand part is present. An information processing device comprising a control unit that executes processing.

Citation Information

Patent Citations

  • Operation determination method, operation determination device, and operation determination program

    JP2015106281A

  • Superimposed display method, superimposed display device, and superimposed display program

    JP2017091078A

  • Information processing system and component apparatus thereof, and method for monitoring real space

    JP2018074528A

  • Information processing device, information processing system, information processing method and program

    JP2018077637A

  • Determination device, determination method and program

    JP2019139570A