Identification device, identification method, program, and identification system
Patent Information
- Application Number
- JP2025028657
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2026-09-07
AI Technical Summary
【0021】 以上説明したように本発明によれば、物体の状態の識別精度を向上させることが可能な技術が提供される。
Smart Images

Figure 2026141901000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an identification device, an identification method, a program, and an identification system. [Background technology]
[0002] In recent years, technologies for identifying the state of objects have become known. Identifying the state of an object can utilize an identification model generated through training based on training data, along with the identification data itself. The sensors used to obtain the training data and the sensors used to obtain the identification data may be the same or different.
[0003] Regardless of whether the sensors used to obtain training data and the sensors used to obtain identification data are the same, the position and orientation of the training data sensors may differ from those of the identification data sensors. In particular, there may be no training data sensors with the same or similar position and orientation as the identification data sensors. In such cases, the identification model will have to identify the state of an object based on an unknown viewpoint, which may reduce the accuracy of the object's state identification.
[0004] Non-patent document 1 discloses a technique in which the same object is photographed from multiple viewpoints to obtain training data from these viewpoints, and training is performed using contrast learning based on the training data obtained from these multiple viewpoints. Such training can improve the accuracy of identifying the state of an object based on an unknown viewpoint. Furthermore, such training is performed by emphasizing common features of the three-dimensional skeletal data estimated from each of the training data obtained from multiple viewpoints. Therefore, such training can create an identification model that is robust to changes in viewpoint.
[0005] Furthermore, the technology described in Non-Patent Document 1 employs a Graph Convolutional Network (GCN) as the model structure. GCN performs convolutional learning using the skeletal structure of a human being. Therefore, GCN is more robust to changes in viewpoint compared to other architectures. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] Cunling Bian et.al, "View-Invariant Skeleton-based Action Recognition via Global-Local Contrastive Learning", [online], arXiv:2209.11634, [Retrieved January 27, 2025], Internet<https: / / arxiv.org / abs / 2209.11634> [Overview of the project] [Problems that the invention aims to solve]
[0007] In the technology described in Non-Patent Document 1, training data obtained from many viewpoints is necessary to significantly improve the accuracy of identifying the state of an object based on an unknown viewpoint. However, situations may arise where it is difficult to prepare training data obtained from many viewpoints. In such situations, it may become difficult to improve the accuracy of identifying the state of an object.
[0008] Therefore, it is desirable to provide technology that can improve the accuracy of identifying the state of an object. [Means for solving the problem]
[0009] To solve the above problems, according to one aspect of the present invention, an identification device is provided, comprising: an estimation unit that estimates the position and orientation of a first part of an object and the position and orientation of a second part of an object based on sensor data obtained by a sensor; a processing unit that obtains the converted position and orientation of the second part by converting the position and orientation of the second part based on the position and orientation of the first part to a position and orientation relative to the first part; and an identification unit that identifies the state of the object based on the converted position and orientation and an identification model for identifying the state of the object. The identification model may be an example of a learning model generated by machine learning.
[0010] The identification device may include an output unit that outputs information corresponding to the state of the object.
[0011] The identification device includes a determination unit that performs a determination based on the state of the object and obtains a determination result, and the output unit may output the determination result.
[0012] The aforementioned determination may include determining whether the state of the object is normal or not.
[0013] The processing unit may obtain the converted position and orientation by converting the position and orientation of the second part to a position and orientation relative to the first part, and then performing a process to set the norm of the vector representing the converted position to a predetermined value.
[0014] The sensor data includes a plurality of frames, the estimating unit estimates the position and orientation of the first part and the position and orientation of the second part for each frame included in the plurality of frames, the processing unit obtains the converted position and orientation of each frame by converting the position and orientation of the second part into a position and orientation based on the first part for each frame, and the identifying unit may identify the state of the object corresponding to the plurality of frames or each frame based on the converted position and orientation of each frame and the identification model.
[0015] The object may be a person or a robot.
[0016] The state may be an action.
[0017] The first part may be a hip joint.
[0018] Further, to solve the above problem, according to another aspect of the present invention, there is provided a computer-implemented identification method, comprising: estimating a position and orientation of a first part of an object and a position and orientation of a second part of the object based on sensor data obtained by a sensor; obtaining a converted position and orientation of the second part by converting the position and orientation of the second part into a position and orientation based on the first part based on the position and orientation of the first part; and identifying the state of the object based on the converted position and orientation and an identification model for identifying the state of the object.
[0019] Further, according to another aspect of the present invention to solve the above problem, there is provided a program that causes a computer to function as: an estimation unit that estimates the position and orientation of a first part of an object and the position and orientation of a second part of the object based on sensor data obtained by a sensor; a processing unit that obtains a converted position and orientation of the second part by converting the position and orientation of the second part into a position and orientation with the first part as a reference, based on the position and orientation of the first part; and an identification unit that identifies the state of the object based on the converted position and orientation and an identification model for identifying the state of the object.
[0020] Further, according to another aspect of the present invention to solve the above problem, there is provided an identification system comprising: an estimation unit that estimates the position and orientation of a first part of a first object and the position and orientation of a second part of the first object based on sensor data obtained by a sensor; a processing unit that obtains a converted position and orientation of the second part by converting the position and orientation of the second part into a position and orientation with the first part as a reference, based on the position and orientation of the first part; and an identification unit that identifies the state of the first object based on the converted position and orientation of the second part and an identification model for identifying the state of the first object, wherein the estimation unit estimates the position and orientation of a third part of a second object and the position and orientation of a fourth part of the second object based on training data, the processing unit obtains a converted position and orientation of the fourth part by converting the position and orientation of the fourth part into a position and orientation with the third part as a reference, based on the position and orientation of the third part, and the identification system comprises a learning unit that generates the identification model based on the converted position and orientation of the fourth part. [Effects of the Invention]
[0021] As described above, according to the present invention, there is provided a technique capable of improving the accuracy of identifying the state of an object. [Brief Description of Drawings]
[0022] [Figure 1] This figure shows an example of the functional configuration of an identification system 10 according to an embodiment of the present invention. [Figure 2] This flowchart shows an example of processing during the learning phase. [Figure 3] This flowchart shows an example of the processing at the identification stage. [Figure 4] This diagram illustrates the joint positions entered into the model in the comparative example. [Figure 5] This figure illustrates the joint positions input to the model in an embodiment of the present invention. [Figure 6] This is a flowchart showing the detailed processing steps for S130. [Figure 7] This figure shows the hardware configuration of an information processing device 900 as an example of an identification system 10 according to an embodiment of the present invention. [Figure 8] This figure shows an example of skeletal estimation based on sensor data from multiple viewpoints. [Modes for carrying out the invention]
[0023] Preferred embodiments of the present invention will be described in detail below with reference to the attached drawings. In this specification and drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant descriptions will be omitted.
[0024] (0. Overview) First, an overview of the embodiments of the present invention will be described.
[0025] As mentioned above, significantly improving the accuracy of identifying the state of an object from an unknown perspective requires training data obtained from many different viewpoints. However, situations can arise where it is difficult to obtain training data from many different viewpoints. For example, in a relatively confined environment (such as a factory work site), the area in which sensors can be installed may be limited, making such a situation likely to occur. In such situations, improving the accuracy of identifying the state of an object can become difficult.
[0026] Therefore, this specification will primarily describe techniques that can improve the accuracy of identifying the state of an object. According to such techniques, even in relatively confined environments where the range in which sensors can be installed is limited and only training data from a few viewpoints can be obtained, it is possible to improve the accuracy of identifying the state of an object. This makes it possible to maintain the accuracy of identifying the state of an object even if the sensors are installed at any position and orientation within a limited range.
[0027] The embodiments of the present invention have been described above.
[0028] (1. Details of the Embodiment) Next, we will describe the details of embodiments of the present invention.
[0029] (1-1. Explanation of Structure) First, an example of the configuration of the identification system 10 according to an embodiment of the present invention will be described. Figure 1 is a diagram showing an example of the functional configuration of the identification system 10 according to an embodiment of the present invention. As shown in Figure 1, the identification system 10 according to an embodiment of the present invention comprises a storage unit 105, a sensor 110, a model input data generation unit 120, a learning unit 130, an identification unit 140, a determination unit 150, and an output unit 160.
[0030] The identification system 10 is executed by a computer. The model input data generation unit 120 and the learning unit 130 may constitute a learning device. Furthermore, the model input data generation unit 120, the identification unit 140, the determination unit 150, and the output unit 160 may constitute an identification device. The identification system 10 may be implemented by a single computer or distributed across multiple computers.
[0031] The model input data generation unit 120, the learning unit 130, the identification unit 140, and the determination unit 150 include a processor such as a CPU (Central Processing Unit), and their functions can be realized by the processor loading a program stored in memory (not shown) into RAM (Random Access Memory) and executing it. In this case, a computer-readable recording medium on which the program is recorded may also be provided. Alternatively, the model input data generation unit 120, the learning unit 130, the identification unit 140, and the determination unit 150 may be configured with dedicated hardware, or with a combination of multiple hardware components.
[0032] (Storage unit 105) The storage unit 105 is composed of memory. The storage unit 105 stores training data used in the training phase. The training data is sensor data obtained by sensors (not shown). The sensor that obtains the training data and the sensor 110 that obtains the identification data may be the same sensor or different sensors. In any case, the position and orientation of the sensor that obtains the training data and the position and orientation of the sensor 110 that obtains the identification data are mainly assumed to be different.
[0033] The training data is time-series data, and the sensor that obtains the training data continuously acquires frames along the time series. For example, the sensor that obtains the training data may be a camera (e.g., an RGB (Red Green Blue) camera, a depth camera, etc.). However, as will be explained later, the sensor that obtains the training data is not limited to a camera. For example, the sensor that obtains the training data may be any sensor that can be used for estimating skeletal data.
[0034] Sensors that obtain training data are installed in the environment. For example, the environment in which sensors that obtain training data are installed may be a work area in a factory where people perform tasks. In this case, human actions can be identified as tasks performed by people. However, the environment in which sensors that obtain training data are installed is not limited to this example. Furthermore, the environment in which sensors that obtain training data are installed and the environment in which the sensor 110 that obtains identification data is installed may be the same environment or may be different environments.
[0035] Furthermore, the person whose actions are identified from the training data and the person whose actions are identified from the identification data may be the same person or different people. The person whose actions are identified from the identification data may correspond to the first object. Also, the person whose actions are identified from the training data may correspond to the second object.
[0036] For example, training data previously obtained by a sensor that acquires training data may be stored in the memory unit 105. The memory unit 105 also stores a model. The model may include a neural network. Based on such a model, an identification model for identifying actions is generated.
[0037] Furthermore, the memory unit 105 stores training data for actions identified from the training data. The training data for actions may be manually assigned to each frame included in the training data beforehand. For example, the training data for actions may be represented by a matrix in which vectors are arranged for each frame, with the actions "standby," "attach part A," "attach part B," and "attach part C" as elements, and the correct element being set to 1 and the non-correct element being set to 0.
[0038] Furthermore, the memory unit 105 may store a judgment model for making judgments based on actions. For example, the judgment may include determining whether an action is normal or not. As an example, the judgment may include determining whether the sequence of actions is normal or not. In this case, the judgment model may be a model that takes the sequence of actions identified by the identification unit 140 as input and determines whether the sequence of actions is normal or not.
[0039] For example, suppose the normal sequence of actions is "standby," "attach part A," "attach part B," and "attach part C." Suppose the person reverses the order of attaching part A and part B, resulting in the actions being "standby," "attach part B," "attach part A," and "attach part C." In this case, since the sequence of actions contains two parts that differ from the normal sequence, it may be judged that the sequence of actions is abnormal.
[0040] On the other hand, suppose the person does not make any mistakes in the order of attaching the parts, and the actions progress in the order of "standby," "attaching part A," "attaching part B," and "attaching part C." In such a case, since there is no part in the sequence of actions that differs from the normal sequence, the sequence of actions may be judged to be normal.
[0041] (Sensor 110) Sensor 110 obtains sensor data as identification data to be used in the identification stage. The identification data, like the training data, is time-series data, and sensor 110 continuously obtains frames along the time series. For example, sensor 110 may be a camera (e.g., an RGB camera, a depth camera, etc.). However, as will be explained later, sensor 110 is not limited to a camera. For example, sensor 110 may be any sensor that can be used for estimating skeletal data.
[0042] The sensor 110 is installed in the environment. For example, the environment in which the sensor 110 is installed may be a work site in a factory where people perform tasks. In this case, human actions can be identified as tasks performed by people. However, the environment in which the sensor 110 is installed is not limited to this example. For example, the identification data obtained by the sensor 110 may be output from the sensor 110 and acquired in real time by the model input data generation unit 120.
[0043] (Model input data generation unit 120) The model input data generation unit 120 generates input data to be input to the discrimination model based on training data during the training phase. On the other hand, during the discrimination phase, the model input data generation unit 120 generates input data to be input to the discrimination model based on discrimination data. To generate such input data, the model input data generation unit 120 includes a skeleton estimation unit 122 and a skeleton data processing unit 124.
[0044] (Skeletal estimation unit 122) The skeleton estimation unit 122 estimates the skeleton data of a person based on the training data during the learning phase. Furthermore, the skeleton estimation unit 122 estimates the skeleton data of a person based on the identification data during the identification phase. Note that both the training data and the identification data may contain multiple frames. Therefore, the skeleton estimation unit 122 may estimate the skeleton data of a person for each frame included in the multiple frames.
[0045] Skeletal data includes the position and orientation of each of several joints. The position of each of these joints is preferably three-dimensional data, but may also be two-dimensional data. Similarly, the orientation of each of these joints is preferably three-dimensional data, but may also be two-dimensional data. Note that the joints are examples of human body parts. Therefore, joints may be replaced with other human body parts.
[0046] The multiple joints included in the skeletal data include a joint that serves as the basis for the conversion process performed by the skeletal data processing unit 124 (hereinafter also referred to as the "reference joint"), and a joint that is the target of the conversion process (hereinafter also referred to as the "target joint").
[0047] The following explanation primarily assumes that the skeletal data contains 16 joints. However, the number of joints in the skeletal data may be other than 16. Furthermore, the structure (i.e., type of joint) of the joints in the skeletal data does not need to be limited. However, it is desirable that the structure and number of joints in the skeletal data be estimated to match the structure and number of joints input into the model.
[0048] Furthermore, the following explanation primarily assumes that the reference joint is the lumbar joint and the target joints are the 15 joints other than the lumbar joint. However, the reference joint does not necessarily have to be the lumbar joint. The reasons why it is preferable for the lumbar joint to be the reference joint will be explained later.
[0049] Furthermore, the reference joint whose position and orientation are estimated during the learning stage may correspond to the first region. Also, the target joint whose position and orientation are estimated during the learning stage may correspond to the second region. Furthermore, the reference joint whose position and orientation are estimated during the identification stage may correspond to the third region. Furthermore, the target joint whose position and orientation are estimated during the identification stage may correspond to the fourth region.
[0050] (Skeleton data processing unit 124) The skeletal data processing unit 124 converts the position of the target joint to a position relative to the reference joint (i.e., the position in the reference joint coordinate system) based on the position and orientation of the reference joint. Furthermore, the skeletal data processing unit 124 converts the orientation of the target joint to an orientation relative to the reference joint (i.e., the orientation in the reference joint coordinate system) based on the position and orientation of the reference joint. As a result, the skeletal data processing unit 124 obtains the converted position and orientation of the target joint.
[0051] Furthermore, the skeletal data processing unit 124 may obtain the converted position and orientation of the target joint by first converting the position and orientation of the target joint to a position and orientation relative to a reference joint, and then applying a process to set the norm of the vector representing the converted position to a predetermined value. Typically, the predetermined value may be 1, but the predetermined value may be a value other than 1.
[0052] As mentioned above, both the training data and the identification data may contain multiple frames. Therefore, the skeletal data processing unit 124 may perform a transformation of the position and posture of the target joint, and standardization of the transformed position, for each frame included in the multiple frames.
[0053] (Learning Section 130) The learning unit 130 generates a recognition model during the learning phase by training the model based on the transformed position and orientation of the target joint obtained by the skeletal data processing unit 124, the training data, and the model. More specifically, the learning unit 130 inputs the transformed position and orientation of the target joint into the model, and trains the model by updating the model's weight parameters based on the output values output from the model corresponding to the transformed position and orientation of the target joint, and the training data.
[0054] More specifically, the learning unit 130 calculates a loss based on the output values from the model and the training data. The learning unit 130 may also calculate the Softmax Cross Entropy error between the output values from the model and the training data as the loss. The learning unit 130 can update the model's weight parameters using backpropagation based on the loss calculated in this way. The learning algorithm may be gradient descent. The learning unit 130 stores the trained model as a discriminative model in the storage unit 105.
[0055] As mentioned above, training data may contain multiple frames. Therefore, the model may output an output value corresponding to each frame. In other words, there may be a one-to-one correspondence between frames and output values.
[0056] Alternatively, the model may output output values corresponding to multiple frames. That is, there may be a multiple-to-one correspondence between frames and output values. However, even if there is a multiple-to-one correspondence between frames and output values, since frames are continuously input to the model input data generation unit 120 in a time series, the model can output output values each time a frame is input to the model input data generation unit 120.
[0057] (Identification unit 140) The identification unit 140 identifies the person's actions during the identification stage. This identification of actions is performed online. Here, "online" may mean that processing is performed sequentially based on the input of frames from the sensor 110 to the model input data generation unit 120. Specifically, the identification unit 140 identifies the person's actions based on the transformed position and posture of the target joint obtained by the skeletal data processing unit 124 and the identification model stored in the memory unit 105.
[0058] More specifically, the identification unit 140 inputs the transformed position and posture of the target joint into the identification model and obtains an output value from the identification model corresponding to the transformed position and posture of the target joint. The identification unit 140 identifies an action based on the output value. For example, if the output value is represented by a vector containing elements for each action, the identification unit 140 may identify the action corresponding to the element with the largest value among the output values as a person's action.
[0059] As mentioned above, the identification data may contain multiple frames. Therefore, the identification model may output an output value corresponding to each frame. In other words, there may be a one-to-one correspondence between frames and output values.
[0060] Alternatively, the identification model may output output values corresponding to multiple frames. That is, there may be multiple-to-one correspondences between frames and output values. However, even if there is a multiple-to-one correspondence between frames and output values, since frames are continuously input to the model input data generation unit 120 in a time series, the identification model can output an output value each time a frame is input to the model input data generation unit 120.
[0061] (Judgment section 150) The determination unit 150 makes a determination based on the action identified by the identification unit 140 and obtains a determination result. For example, as described above, if a determination model for making an action-based determination is stored in the memory unit 105, the determination unit 150 may make a determination using such a determination model. Alternatively, the determination unit 150 may make an action-based determination using a rule-based method.
[0062] Here, the determination made by the determination unit 150 may include determining whether or not the behavior is normal. For example, determining whether or not the behavior is normal may include determining whether or not the sequence of actions is normal, as described above. However, as will be explained later, the determination made by the determination unit 150 is not limited to such examples.
[0063] (Output section 160) The output unit 160 is implemented by an output device and outputs information corresponding to the action identified by the identification unit 140. The output device may be installed in an environment where a person whose action is being identified is present (for example, a work site). This allows the person to immediately grasp the information corresponding to their action. Here, the form of the output device is not limited.
[0064] For example, the output device may include a display. In this case, the output unit 160 may display information corresponding to the action on the display. Alternatively, the output device may include a speaker. In this case, the output unit 160 may output information corresponding to the action through the speaker. Alternatively, the output device may include a lamp. In this case, the output unit 160 may output information corresponding to the action by lighting or flashing the lamp.
[0065] For example, the action-dependent information may be information indicating an action identified by the identification unit 140, such as "standby," "installing part A," "installing part B," and "installing part C." Alternatively, the action-dependent information may be a determination result obtained by the determination unit 150. For example, the output unit 160 may output a determination result regardless of whether the action is normal or not. Alternatively, the output unit 160 may output a determination result indicating that the action is abnormal based on the determination that the action is abnormal.
[0066] The above describes an example of the configuration of the identification system 10 according to an embodiment of the present invention.
[0067] (1-2. Operation Description) Next, an example of processing of the identification system 10 according to an embodiment of the present invention will be described. First, an example of processing of the identification system 10 in the learning stage will be described with reference to Figure 2 (and Figure 1 as appropriate). Processing in the learning stage is mainly performed by the model input data generation unit 120 and the learning unit 130. Next, an example of processing of the identification system 10 in the identification stage will be described with reference to Figure 3 (and Figure 1 as appropriate). Processing in the identification stage is mainly performed by the sensor 110, the model input data generation unit 120, the identification unit 140, the determination unit 150, and the output unit 160.
[0068] (Learning stage) Figure 2 is a flowchart illustrating an example of processing during the learning phase. As shown in Figure 2, during the learning phase, the skeleton estimation unit 122 acquires learning data from the memory unit 105 (S110). Then, based on the learning data, the skeleton estimation unit 122 estimates the position and posture of each of the multiple joints of the person (S120). The position and posture of each of the multiple joints of the person correspond to the person's skeleton data.
[0069] The skeletal data processing unit 124 converts the position of the target joint to a position relative to the reference joint (i.e., the position in the reference joint coordinate system) based on the position and orientation of the reference joint (S130). This allows the skeletal data processing unit 124 to obtain the converted position and orientation of the target joint. Details of S130 will be explained later with reference to Figures 4 to 6.
[0070] The learning unit 130 generates a classification model by training the model based on the transformed position and orientation of the target joint obtained by the skeletal data processing unit 124, the training data, and the model (S140). More specifically, the learning unit 130 inputs the transformed position and orientation of the target joint into the model, and trains the model by updating the model's weight parameters based on the output values output from the model corresponding to the transformed position and orientation of the target joint, and the training data.
[0071] The learning unit 130 determines whether or not to terminate the learning phase (S150). If the learning unit 130 determines not to terminate the learning phase (NO in S150), it proceeds to S110. On the other hand, if the learning unit 130 determines to terminate the learning phase (YES in S150), it stores the trained model as an identification model in the storage unit 105 (S160) and terminates the learning phase.
[0072] Furthermore, the termination conditions for the learning phase are not particularly limited and can be any conditions that indicate that learning has been performed to a certain extent. Specifically, the termination conditions for the learning phase may include the condition that the loss is less than a threshold. Alternatively, the termination conditions for the learning phase may include the condition that the change in loss is less than a threshold (the condition that the loss has converged). Alternatively, the termination conditions for the learning phase may include the condition that the weight parameters have been updated a predetermined number of times.
[0073] The above describes an example of the processing of the identification system 10 during the learning phase.
[0074] (Identification stage) Figure 3 is a flowchart illustrating an example of processing during the identification stage. In the example shown in Figure 3, the data acquisition cycle, which is the period during which frames are acquired from sensor 110, and the inference cycle, which is the period during which behavior identification is performed, are the same. However, the data acquisition cycle and the inference cycle may be different. That is, behavior identification may be performed in an inference cycle that is asynchronous to the data acquisition cycle.
[0075] Thus, the processing order in the identification stage does not necessarily have to be as shown in Figure 3, and may be appropriately modified to the extent that information corresponding to the identified behavior can be notified to the person in real time. Notifying the person in real time means that the notification should be given at a time that allows the person to notice the abnormality in their behavior while they are performing the action. For example, the person may be notified of the abnormal behavior 5 seconds after the abnormal behavior occurs.
[0076] As shown in Figure 3, during the identification stage, the skeleton estimation unit 122 acquires the sensor data output from the sensor 110 as identification data (S210). Then, based on the identification data, the skeleton estimation unit 122 estimates the position and posture of each of the multiple joints of the person (S120).
[0077] The skeletal data processing unit 124 converts the position of the target joint to a position relative to the reference joint (i.e., the position in the reference joint coordinate system) based on the position and orientation of the reference joint (S130). This allows the skeletal data processing unit 124 to obtain the converted position and orientation of the target joint. Details of S130 will be explained later with reference to Figures 4 to 6.
[0078] The identification unit 140 identifies the person's actions based on the transformed position and posture of the target joint obtained by the skeletal data processing unit 124 and the identification model stored in the memory unit 105 (S250).
[0079] More specifically, the identification unit 140 inputs the transformed position and posture of the target joint into the identification model and obtains an output value from the identification model corresponding to the transformed position and posture of the target joint. The identification unit 140 identifies an action based on the output value. For example, if the output value is represented by a vector containing elements for each action, the identification unit 140 may identify the action corresponding to the element with the largest value among the output values as a person's action.
[0080] The determination unit 150 makes a determination based on the action identified by the identification unit 140 and obtains a determination result. For example, the determination unit 150 may determine whether the action is normal or not (S260). For example, the determination unit 150 may determine whether the sequence of actions is normal or not. However, as will be explained later, the determination made by the determination unit 150 is not limited to such examples.
[0081] The output unit 160 outputs information corresponding to the behavior identified by the identification unit 140. For example, the output unit 160 may output the determination result obtained by the determination unit 150 as an example of information corresponding to the behavior (S270). For example, the output unit 160 may output the determination result regardless of whether the behavior is normal or not. Alternatively, the output unit 160 may output a determination result indicating that the behavior is abnormal based on the determination that the behavior is abnormal.
[0082] The information corresponding to the action output by the output unit 160 is understood by the person. By understanding the information corresponding to the action, the person can become aware of errors in their own actions (for example, errors in work).
[0083] The identification unit 140 determines whether or not to terminate the identification stage (S280). If the identification unit 140 determines not to terminate the identification stage (NO in S280), it proceeds to S210. On the other hand, if the identification unit 140 determines to terminate the identification stage (YES in S280), it terminates the identification stage.
[0084] The termination conditions for the identification stage are not particularly limited. For example, the termination conditions for the identification stage may include a condition that a predetermined termination operation has been input by a user of the identification system 10. Alternatively, the termination conditions for the identification stage may include a condition that the current time has reached a predetermined termination time.
[0085] The above describes an example of the processing of the identification system 10 during the identification stage.
[0086] (Details of S130) Next, we will explain the details of S130 with reference to Figures 4 to 6. In the following explanation, position is represented by three-dimensional coordinates in the reference coordinate system, and orientation is represented by quaternions representing rotations in the reference coordinate system. Orientation may be represented by methods other than quaternions. For example, orientation may be represented by a rotation matrix or Euler angles.
[0087] Figure 4 is a diagram illustrating the joint positions input to the model in the comparative example. Figure 5 is a diagram illustrating the joint positions input to the model in the embodiment of the present invention. First, with reference to Figures 4 and 5, we will briefly explain the differences in the joint positions input to the model between the comparative example and the embodiment of the present invention.
[0088] In the examples shown in Figures 4 and 5, cameras C1 and C2 are installed as examples of sensors. The coordinate system xyz relative to the position and orientation of camera C1 is the camera coordinate system {C1}. The coordinate system xyz relative to the position and orientation of camera C2 is the camera coordinate system {C2}. The skeletal data f1 of the person whose actions are identified includes the position and orientation of 16 joints. As an example of the 16 joints, the joint code "f11" is shown.
[0089] Generally, when skeletal estimation is performed from video images obtained by a camera, the position and orientation of the joints are often represented in the camera coordinate system of the camera that obtained the video images.
[0090] As shown in Figure 4, the position of the head joint estimated from the video image obtained by camera C1 is represented by a vector Vh1 extending from the origin of the camera coordinate system {C1} to the position of the head joint. Similarly, the position of the head joint estimated from the video image obtained by camera C2 is represented by a vector Vh2 extending from the origin of the camera coordinate system {C2} to the position of the head joint.
[0091] Furthermore, as shown in Figure 4, the position of the hip joint estimated from the video image obtained by camera C1 is represented by a vector Vp1 extending from the origin of the camera coordinate system {C1} to the position of the hip joint. Similarly, the position of the hip joint estimated from the video image obtained by camera C2 is represented by a vector Vp2 extending from the origin of the camera coordinate system {C2} to the position of the hip joint.
[0092] As can be seen from the example shown in Figure 4, the characteristics of the vectors representing joint positions relative to the different camera coordinate systems {C1} and {C2} (e.g., the length of the vector and the direction of the vector) will differ even when representing the position of the same joint (e.g., the head joint or the hip joint) of the same person.
[0093] On the other hand, in embodiments of the present invention, the joint positions and orientations relative to the different camera coordinate systems {C1} and {C2} are converted to positions and orientations relative to a single coordinate system. This allows the joint positions and orientations relative to the different camera coordinate systems {C1} and {C2} to be normalized to the same position and orientation.
[0094] For example, as shown in Figure 5, consider transforming vectors Vh1 and Vh2 into vectors with {PL} as the origin, which is a coordinate system xyz based on the position and posture of the lumbar joint, and setting the norm of the transformed vectors to 1. Then, both vectors are transformed into unit vectors Vph that extend from the origin of {PL}, which is a coordinate system xyz based on the position and posture of the lumbar joint, towards the origin of {HD}, which is a coordinate system xyz based on the position and posture of the head joint.
[0095] While the process of setting the norm to 1 is not mandatory, it is expected that doing so will reduce the influence of individual differences in physical characteristics such as height or build, and enable the identification of behaviors that are not dependent on a person's physical characteristics. Typically, the norm is set to 1, but it may be set to a predetermined value other than 1.
[0096] This normalization ensures that highly common features (i.e., joint positions and postures) are obtained regardless of whether the features are estimated from video footage obtained by camera C1 or camera C2. Furthermore, the recognition accuracy of a model trained on these features is independent of the camera's viewpoint. However, in reality, viewpoint-specific noise exists, so it is not always possible to obtain completely common features.
[0097] Figure 6 is a flowchart detailing the processing steps of S130. The detailed processing steps of S130 will be explained with reference to Figure 6 (and to Figures 1 to 5 as appropriate).
[0098] Here, when there is a second joint that moves in conjunction with the movement of the first joint, the first joint is also called the "parent joint," and the second joint is also called the "child joint." Furthermore, a parent joint that does not have a parent joint of its own is also called the "root joint."
[0099] The reference joint can be arbitrarily selected by the user during training. However, in general motion capture, the estimation error of child joints is greatly influenced by the estimation error of parent joints, and the estimation error of the root joint is considered to be the smallest. Therefore, considering the stabilization of the position and posture of the reference joint, it is desirable to select the root joint as the reference joint.
[0100] Since the lumbar joint is an example of a root joint, the example shown in Figure 6 assumes that the reference joint is the lumbar joint. Furthermore, it is assumed that the target joints are 15 joints other than the lumbar joint. Each of the 15 target joints is assigned a unique number i (1 ≤ i ≤ 15), and these 15 target joints are also referred to as "joint number i." The explanation will proceed assuming that joint number 1 is the head and knee joint.
[0101] As shown in Figure 6, the skeletal data processing unit 124 acquires the position and orientation of the reference joint estimated by the skeletal estimation unit 122, as well as the position and orientation of joints 1 through 15 (S131).
[0102] In the following description, it is assumed that the position and orientation of the reference joint and the position and orientation of the first joint are estimated from a moving image obtained by the camera C1. Let the position of the reference joint be p C1_PL , and let the position of the first joint be p C1_HD . Further, let the orientation of the reference joint be q C1_PL , and let the orientation of the first joint be q C1_HD . The position is represented by three-dimensional data, and the orientation is represented by a quaternion.
[0103] Next, the skeleton data processing unit 124 creates a homogeneous transformation matrix using the position and orientation of the reference joint (S132). In the following description, let the homogeneous transformation matrix be H PL_C1 . The homogeneous transformation matrix H PL_C1 can be represented by a 4-dimensional × 4-dimensional matrix. The homogeneous transformation matrix H PL_C1 can be derived based on the positional and orientational relationship between {PL}, which is a coordinate system xyz with the position and orientation of the reference joint as the reference, and the camera coordinate system {C1}. That is, the homogeneous transformation matrix H PL_C1 can be derived based on the position p of the reference joint C1_PL and the orientation q of the reference joint C1_PL .
[0104] More specifically, the skeleton data processing unit 124 converts the orientation q of the reference joint C1_PL into a 3-dimensional × 3-dimensional rotation matrix. Let this rotation matrix be Q. The skeleton data processing unit 124 derives the transpose matrix Q of the rotation matrix Q T . Further, let the position p of the reference joint C1_PL be a vector P. At this time, the skeleton data processing unit 124 can derive the homogeneous transformation matrix H PL_C1 as shown in the following formula (1).
[0105]
Mathematical Expression
[0106] Next, the skeleton data processing unit 124 sets i to 1, and the homogeneous transformation matrix H PL_C1The position of joint i is transformed to its position in the reference joint coordinate system using (S133). General matrix multiplication can be applied to the position transformation.
[0107] More specifically, the skeletal data processing unit 124 processes the homogeneous transformation matrix H PL_C1 The calculation result is obtained by multiplying the i-th joint's position by a four-dimensional vector obtained by adding 1 to the fourth element of the three-dimensional vector representing the position of the i-th joint. Then, the skeletal data processing unit 124 extracts the first three elements of the calculated four-dimensional vector and arranges the three extracted elements to form a three-dimensional vector, which represents the position p in the reference joint coordinate system. PL_HD It is calculated as follows.
[0108] Next, the skeletal data processing unit 124 converts the pose of joint i to its pose in the reference joint coordinate system (S134). The pose conversion involves converting the pose of the reference joint q C1_PL And the position of joint q C1_HD The following may be applied. More specifically, the skeletal data processing unit 124 determines the posture q of the reference joint. C1_PL The posture q is the reverse rotation of the same object. C1_PL And the position of joint i, p C1_HD The result of combining and according to the quaternion synthesis formula is the pose q of joint i in the reference joint coordinate system. PL_HD It is calculated as follows.
[0109] Next, the skeletal data processing unit 124 processes the vectors representing the positions of the joints after conversion so that their norm becomes 1 (S135).
[0110] Next, the skeletal data processing unit 124 determines whether processing has been completed for all joints (joints 1 through 15) (S137). If there are any joints that have not been processed (NO in S137), the skeletal data processing unit 124 increments i by 1 (S136) and proceeds to S133. On the other hand, if processing has been completed for all joints (YES in S137), the skeletal data processing unit 124 terminates processing.
[0111] As explained with reference to Figure 6, by performing a coordinate transformation on the position and orientation of each joint, viewpoint-independent features (i.e., joint position and orientation) can be extracted. Furthermore, in the embodiment of the present invention, this coordinate transformation can be performed every frame. That is, a coordinate transformation can be performed each time the skeleton estimation unit 122 performs skeleton estimation. At this time, the position and orientation of the reference joint are also estimated every frame, and the homogeneous transformation matrix is updated every frame. Therefore, according to the embodiment of the present invention, unlike the comparative example which identifies actions using features based on a static coordinate system fixed to the environment, it is possible to identify actions using features based on a dynamic coordinate system not fixed to the environment.
[0112] (2. Hardware Configuration Example) Next, an example of the hardware configuration of the identification system 10 according to an embodiment of the present invention will be described. In the following, an example of the hardware configuration of the information processing device 900 will be described as an example of the hardware configuration of the identification system 10 according to an embodiment of the present invention. Note that the example of the hardware configuration of the information processing device 900 described below is merely one example of the hardware configuration of the identification system 10. Therefore, the hardware configuration of the identification system 10 may be modified by removing unnecessary components from the hardware configuration of the information processing device 900 described below, or new components may be added.
[0113] Figure 7 shows the hardware configuration of an information processing device 900 as an example of an identification system 10 according to an embodiment of the present invention. The information processing device 900 includes a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, a RAM (Random Access Memory) 903, a host bus 904, a bridge 905, an external bus 906, an interface 907, an input device 908, an output device 909, a storage device 910, and a communication device 911.
[0114] The CPU 901 functions as both an arithmetic processing unit and a control unit, controlling the overall operation of the information processing unit 900 according to various programs. The CPU 901 may also be a microprocessor. The ROM 902 stores programs and arithmetic parameters used by the CPU 901. The RAM 903 temporarily stores programs used in the execution of the CPU 901 and parameters that change as needed during its execution. These are interconnected by a host bus 904, which consists of a CPU bus and other components.
[0115] The host bus 904 is connected to an external bus 906, such as a PCI (Peripheral Component Interconnect / Interface) bus, via a bridge 905. It is not always necessary to configure the host bus 904, bridge 905, and external bus 906 separately; these functions may be implemented on a single bus.
[0116] The input device 908 consists of input means for the user to input information, such as a mouse, keyboard, touch panel, buttons, microphone, switches, and levers, and an input control circuit that generates input signals based on the user's input and outputs them to the CPU 901. The user operating the information processing device 900 can input various types of data to the information processing device 900 or instruct it to perform processing operations by operating this input device 908.
[0117] The output device 909 includes, for example, display devices such as CRT (Cathode Ray Tube) display devices, liquid crystal display (LCD) devices, OLED (Organic Light Emitting Diode) devices, lamps, and audio output devices such as speakers.
[0118] The storage device 910 is a device for storing data. The storage device 910 may include a storage medium, a recording device for recording data on the storage medium, a reading device for reading data from the storage medium, and a deletion device for deleting data recorded on the storage medium. The storage device 910 is composed of, for example, an HDD (Hard Disk Drive). This storage device 910 drives the hard disk and stores programs executed by the CPU 901 and various data.
[0119] The communication device 911 is a communication interface composed of, for example, a communication device for connecting to a network. The communication device 911 may support either wireless or wired communication.
[0120] The above describes an example of the hardware configuration of the identification system 10 according to an embodiment of the present invention.
[0121] (3. Explanation of effects) According to an embodiment of the present invention, the positions and orientations of other joints are transformed to the positions and orientations relative to the coordinate system of the reference joint using a homogeneous transformation matrix created from the position and orientation of a reference joint. Therefore, features with high commonality can be obtained regardless of the position and orientation of the sensor. By inputting similarly normalized features into a discrimination model trained using these normalized features, viewpoint-independent behavioral discrimination can be achieved.
[0122] Furthermore, according to embodiments of the present invention, by transforming the positions and postures of other joints based on a reference joint estimated each frame, it is possible to perform behavioral identification using features based on a dynamic coordinate system that is not fixed to the environment. This makes it possible to perform behavioral recognition that does not depend on the absolute position of a person.
[0123] (4. Various variations) Although preferred embodiments of the present invention have been described in detail above with reference to the attached drawings, the present invention is not limited to these examples. It is clear to any person with ordinary skill in the art to which the present invention belongs that various modifications or alterations can be conceived within the scope of the technical idea described in the claims, and these are also understood to fall within the technical scope of the present invention.
[0124] (Sensor variations) The above primarily assumes that the training data and classification data are video images obtained by a camera. However, the training data and classification data may also be still images obtained by a camera. Furthermore, a camera is just one example of a sensor, and video and still images are examples of sensor data. Therefore, the training data and classification data may also be sensor data obtained by sensors other than a camera.
[0125] For example, the training data and the identification data may each be sound data obtained by a microphone. Alternatively, the training data and the identification data may each be a three-dimensional point cloud obtained by LiDAR (Light Detection and Ranging), acceleration obtained by an accelerometer, angular velocity obtained by a gyroscope, or electromyography (EMG) obtained by an EMG sensor.
[0126] (Variations in identification processing) The above primarily described an example where a discrimination model identifies human behavior. However, a human being is just one example of an action subject. Therefore, the action subject whose behavior is identified may be something other than a human (for example, a robot or an animal). For example, the robot may be an industrial robot.
[0127] Furthermore, the actions of an agent are just one example of an object's state. Therefore, the discriminant model may also discriminate against object states other than those of an agent. For example, the state of an object could be the state of a vehicle. Specifically, by estimating multiple parts of a vehicle, designating one part as a reference joint, and considering the remaining parts as target joints, the state of the vehicle can be accurately identified by performing the same normalization process as described above. The state of the vehicle could be, for example, whether the vehicle is turning left, turning right, or driving on a curved road.
[0128] (Variations in the judgment process) The above description illustrates an example in which the determination performed by the determination unit 150 includes determining whether the sequence of actions is normal or not. However, the determination performed by the determination unit 150 is not limited to this example. For example, the determination performed by the determination unit 150 may include determining the action duration, which is the time the same action continued, based on the action identification result by the identification unit 140.
[0129] In this case, the output unit 160 may output the duration of the action itself. Alternatively, the determination unit 150 may determine whether the duration of the action is longer than a threshold, and the output unit 160 may output the result of that determination. Alternatively, the output unit 160 may output that the duration of the action is abnormal if it is longer than a threshold. With such a configuration, it can be expected that the action speed (e.g., work speed) of the person who checks the outputted information will increase. Furthermore, the determinations made by the determination unit 150 are not limited to these examples, and embodiments of the present invention can be easily applied to applications that perform various determinations.
[0130] (Use of data from multiple viewpoints) The above mainly described an example where the normalization described above is applied to skeletal data obtained from single-view sensor data, and the normalized data is used for training. However, the normalization described above may also be applied to skeletal data obtained from sensor data from multiple views, and the normalized data may be used for training. This may enable more accurate behavioral identification.
[0131] When motion capture is performed on single-viewpoint sensor data, viewpoint-specific noise due to occlusion may be included in the measured values. Therefore, viewpoint-specific noise will also be included in the normalized data. By using normalized skeletal data obtained from sensor data of multiple viewpoints for training, it is possible to create a model that can perform identification while taking noise into account. In such cases, sensor data from multiple viewpoints can be collected during the training data collection phase, and single-viewpoint sensor data can be used for identification data to perform action identification.
[0132] (Variations in skeletal estimation) The above mainly described an example in which the skeleton estimation unit 122 performs skeleton estimation based on sensor data from a single viewpoint. However, the skeleton estimation unit 122 may also perform skeleton estimation based on sensor data from multiple viewpoints. Such an example will be explained with reference to Figure 8.
[0133] Figure 8 shows an example of skeletal estimation based on sensor data from multiple viewpoints. As shown in Figure 8, there are skeletal estimation methods that integrate sensor data obtained from many cameras (cameras C1 to C3) to achieve highly accurate skeletal estimation. Embodiments of the present invention can also be applied to such skeletal estimation methods.
[0134] Generally, in skeletal estimation methods based on sensor data from multiple viewpoints, the reference coordinate system related to the output is determined by calibration. In the example shown in Figure 8, it is assumed that the reference coordinate system related to the output is determined to be {W}. However, even in such a case, as described above, it is possible to convert the position and orientation of the target joint to the coordinate system of the reference joint. Furthermore, even if calibration is performed again and the origin of the reference coordinate system related to the output is changed in such a case, it is possible to perform action identification without degrading accuracy. [Explanation of Symbols]
[0135] 10 Identification System 105 Storage section 110 Sensor 120 Model Input Data Generation Unit 122 Skeleton Estimation Section 124 Skeleton Data Processing Unit 130 Learning Department 140 Identification Unit 150 Judgment section 160 Output section
Claims
1. An estimation unit that estimates the position and orientation of a first part of an object and the position and orientation of a second part of the object based on sensor data obtained by the sensor, A processing unit that obtains the converted position and orientation of the second part by converting the position and orientation of the second part to a position and orientation relative to the first part based on the position and orientation of the first part, An identification unit that identifies the state of the object based on the converted position and orientation and an identification model for identifying the state of the object, An identification device equipped with the following features.
2. The identification device is It includes an output unit that outputs information corresponding to the state of the object, The identification device according to claim 1.
3. The identification device is The system includes a determination unit that performs a determination based on the state of the object and obtains a determination result, The output unit outputs the determination result. The identification device according to claim 2.
4. The determination includes determining whether the state of the object is normal or not. The identification device according to claim 3.
5. The processing unit converts the position and orientation of the second part into a position and orientation relative to the first part, and obtains the converted position and orientation by performing a process to set the norm of the vector representing the converted position to a predetermined value. The identification device according to claim 1.
6. The aforementioned sensor data includes multiple frames, The estimation unit estimates the position and orientation of the first part and the position and orientation of the second part for each frame included in the plurality of frames. The processing unit obtains the converted position and orientation for each frame by converting the position and orientation of the second part relative to the position and orientation of the first part for each frame. The identification unit identifies the state of the object corresponding to the plurality of frames or each frame based on the converted position and orientation of each frame and the identification model. The identification device according to claim 1.
7. The object is a person or a robot. The identification device according to claim 1.
8. The aforementioned state is an action. The identification device according to claim 7.
9. The first part mentioned above is the lumbar joint. The identification device according to claim 7.
10. Based on sensor data obtained by the sensor, the position and orientation of a first part of the object and the position and orientation of a second part of the object are estimated. Based on the position and orientation of the first part, the position and orientation of the second part are converted to a position and orientation relative to the first part, thereby obtaining the converted position and orientation of the second part. Based on the converted position and orientation and an identification model for identifying the state of the object, the state of the object is identified. A computer-based identification method, including [a specific method].
11. Computers, An estimation unit that estimates the position and orientation of a first part of an object and the position and orientation of a second part of the object based on sensor data obtained by the sensor, A processing unit that obtains the converted position and orientation of the second part by converting the position and orientation of the second part to a position and orientation relative to the first part based on the position and orientation of the first part, An identification unit that identifies the state of the object based on the converted position and orientation and an identification model for identifying the state of the object, A program that makes it function as such.
12. An estimation unit that estimates the position and orientation of a first part of a first object and the position and orientation of a second part of the first object based on sensor data obtained by the sensor, A processing unit that obtains the converted position and orientation of the second part by converting the position and orientation of the second part to a position and orientation relative to the first part based on the position and orientation of the first part, An identification unit that identifies the state of the first object based on the transformed position and orientation of the second part and an identification model for identifying the state of the first object, An identification system comprising, The estimation unit estimates the position and orientation of the third part of the second object and the position and orientation of the fourth part of the second object based on the training data. The processing unit obtains the converted position and orientation of the fourth part by converting the position and orientation of the fourth part to a position and orientation relative to the third part, based on the position and orientation of the third part. The aforementioned identification system is The system includes a learning unit that generates the identification model based on the transformed position and orientation of the fourth part, Identification system.