Estimation device, learning device, method, and program

The estimation device improves behavior recognition accuracy by using video information extraction and object affordance-based model updates to differentiate between symmetric or similar actions in video sequences.

JP7754316B2Active Publication Date: 2025-10-15NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024530263
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-07-01
Publication Date
2025-10-15
Estimated Expiration
2042-07-01

AI Technical Summary

Technical Problem

Existing behavior recognition technologies struggle to accurately differentiate between symmetric or similar actions in video sequences due to the symmetry or similarity in the time series of RGB information or posture characteristics, leading to misrecognition.

Method used

An estimation device that extracts video information as vectors and estimates behavior based on model parameters updated with information related to the affordances of objects associated with the subject's actions, using a video information extraction unit and a behavior estimation unit.

Benefits of technology

Enhances the accuracy of behavior estimation regardless of video type by incorporating object affordance information, reducing misrecognition of symmetric or similar actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007754316000003
    Figure 0007754316000003
  • Figure 0007754316000004
    Figure 0007754316000004
  • Figure 0007754316000005
    Figure 0007754316000005
Patent Text Reader

Abstract

An estimation device according to an embodiment of the present invention has: a video information extraction unit that extracts video information in which a video composed of time-series frame images is vectorially represented; and an action estimation unit that, on the basis of the video information extracted by the video information extraction unit and a parameter of a model, estimates information pertaining to an action of a subject recorded in the video. The parameter is updated on the basis of information pertaining to the affordance of an object that is estimated on the basis of the video information extracted by the video information extraction unit and is associated with the action of the subject, and information pertaining to the action that is estimated by the action estimation unit.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD Embodiments of the present invention relate to an estimation device, a learning device, a method, and a program. [Background technology]

[0002] There is a technology called behavior recognition that recognizes the behavior of subjects in video. This technology is expected to be applied to various use cases, such as automatic recording of worker activities or on-site situation analysis.

[0003] For example, behavior recognition can recognize the behavior of workers on-site in a factory, making it possible to understand the worker's work status and detect dangerous behavior. Furthermore, behavior recognition can recognize the behavior of customers in a retail store, for example, making it possible to obtain customer information that cannot be obtained from purchase history. Furthermore, behavior recognition can recognize the behavior of nurses or patients in a hospital, making it possible to automate the nurses' record-keeping work. The applications of behavior recognition are endless.

[0004] Behavior recognition is formulated as a classification problem of estimating which of predefined behaviors the behavior of a subject recorded in a video corresponds to. Commonly proposed behavior recognition technologies include, for example, a technology that recognizes behavior using a sequence of RGB information in a video of a subject, as described in Non-Patent Document 1, and a technology that recognizes behavior using a sequence of postures of a subject, as described in Non-Patent Document 2.

[0005] As described in Non-Patent Documents 1 and 2, it is common to provide a model with a sequence of RGB information of a frame image or information on the posture of a subject, and correct information on labels that represent actions paired with this sequence, i.e., correct action labels, and allow the model to learn. However, simply learning pairs of RGB information or subject posture information and correct action labels makes it difficult to accurately recognize actions in which the changes in input information are symmetric or similar over time.

[0006] For example, when a video recording the action of "a person pushing a door open" is compared with a video recording the action of "a person pulling a door closed," it can be said that the characteristics of the subject's appearance expressed by RGB information or the characteristics of the person's posture, such as their arms or standing posture, have symmetry in the time series direction.

[0007] For example, when a video recording the action of "a person staples paper with a stapler" is compared with a video recording the action of "a person cutting paper with scissors," it can be said that the characteristics of human posture have similarities in the time series direction, such as holding paper in one hand and opening and closing the thumb and index finger holding scissors in the other hand. [Prior art documents] [Non-patent literature]

[0008] [Non-Patent Document 1] Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming He. SlowFast Networks for Video Recognition. in Proceedings of the IEEE / CVF International Conference on Computer Vision (ICCV). pp. 6202-6211. 2019. [Non-patent document 2] Lei Shi, Yifan Zhang, Jian Cheng, and Hanqing Lu. Skeleton-Based Action Recognition With Multi-Stream Adaptive Graph Convolutional Networks. in IEEE Transactions on Image Processing, vol. 29, pp. 9532-9545, 2020. Summary of the Invention [Problem to be solved by the invention]

[0009] In the above action of "a person pushing the door open," the door transitions from a closed state to an open state due to the human action. Also, in the above action of "a person pulling the door closed," the door transitions from an open state to a closed state due to the human action.

[0010] Furthermore, in the above action of "a human cutting a piece of paper with scissors," the state of a single piece of paper transitions to a state in which it is separated into multiple pieces of paper, and in the above action of "a human stapler binding a piece of paper together," the state of a single piece of paper transitions to a state in which the multiple pieces of paper are combined into one piece of paper, due to the human action. In this way, information about the behavior can be thought of as being present in the surrounding environment, such as the objects the actor is acting on.

[0011] The "door open state" given in the example of the above behavior of "a person pushing the door open" is a "state in which it is impossible for a person to push it open." In other words, the behavior of "a person pushing the door open" can be considered to be a transition from the "door closed state" to the "door closed state," i.e., a "state in which it is possible for a person to push it open."

[0012] In other words, a change in the state of an object that an agent is acting on has significance in terms of a change in the affordance of the action available to humans. If we can teach a system to learn such changes in the affordance of objects for humans, we can reduce the misrecognition of actions that are symmetric or similar in the time series, such as "a human pushing a door open" and "a human pulling a door closed," and we can expect to establish more accurate action recognition.

[0013] This invention has been made in light of the above circumstances, and its purpose is to provide an estimation device, learning device, method, and program that are capable of estimating the behavior of an actor with higher accuracy regardless of the type of video. [Means for solving the problem]

[0014] An estimation device according to one aspect of the present invention includes a video information extraction unit that extracts video information in which a video composed of a time-series of frame images is vector-represented, and a behavior estimation unit that estimates information related to the behavior of a subject recorded in the video based on the video information extracted by the video information extraction unit and model parameters, wherein the parameters are updated based on information related to the affordances of objects associated with the behavior of the subject, which is estimated based on the video information extracted by the video information extraction unit, and information related to the behavior estimated by the behavior estimation unit.

[0015] An estimation method according to one aspect of the present invention is a method performed by an estimation device, in which the video information extraction unit of the estimation device extracts video information in which a video composed of a time-series of frame images is expressed as a vector, and the behavior estimation unit of the estimation device estimates information related to the behavior of a subject recorded in the video based on the video information extracted by the video information extraction unit and model parameters, and the parameters are updated based on information related to the affordances of an object associated with the behavior of the subject, which is estimated based on the video information extracted by the video information extraction unit, and the information related to the behavior estimated by the behavior estimation unit. [Effects of the Invention]

[0016] According to the present invention, the behavior of the performer can be estimated with higher accuracy regardless of the type of video. [Brief explanation of the drawings]

[0017] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of a behavior recognition device according to the first embodiment. [Figure 2] FIG. 2 is a flowchart showing an example of a procedure of a learning process performed by the behavior recognition device according to the first embodiment. [Figure 3] FIG. 3 is a flowchart showing an example of a procedure of an inference (estimation) process performed by the behavior recognition device according to the first embodiment. [Figure 4] FIG. 4 is a flowchart illustrating an example of a processing procedure performed by the video information extraction unit during the learning process by the behavior recognition device according to the first embodiment. [Figure 5] FIG. 5 is a flowchart illustrating an example of a procedure of processing by the behavior estimation unit during learning processing by the behavior recognition device according to the first embodiment. [Figure 6] FIG. 6 is a flowchart illustrating an example of a processing procedure performed by the affordance estimation unit during the learning process by the behavior recognition device according to the first embodiment. [Figure 7]FIG. 7 is a flowchart illustrating an example of a procedure of processing by the parameter update unit during learning processing by the behavior recognition device according to the first embodiment. [Figure 8] FIG. 8 is a flowchart illustrating an example of a processing procedure by the behavior estimation unit during the inference processing by the behavior recognition device according to the first embodiment. [Figure 9] FIG. 9 is a block diagram illustrating an example of the configuration of a behavior recognition device according to the second embodiment. [Figure 10] FIG. 10 is a flowchart illustrating an example of a processing procedure performed by the affordance estimation unit during the learning process by the behavior recognition device according to the second embodiment. [Figure 11] FIG. 11 is a flowchart illustrating an example of a procedure of processing by a parameter update unit during learning processing by the behavior recognition device according to the second embodiment. [Figure 12] FIG. 12 is a block diagram showing an example of the hardware configuration of the behavior recognition device. DETAILED DESCRIPTION OF THE INVENTION

[0018] An embodiment of the present invention will be described below with reference to the drawings. (First embodiment) First, a first embodiment will be described. Fig. 1 is a block diagram showing an example of the configuration of a behavior recognition device according to the first embodiment. The behavior recognition device 100 shown in Fig. 1 is a device that recognizes the behavior of an actor recorded in a first-person perspective video or a third-person perspective video (time-series frame images). A first-person perspective video is, for example, a video obtained by capturing a visual field of a person from the head or chest of the person. A third-person perspective video is, for example, a video obtained by capturing a bird's-eye view of the entire person from a third-person perspective.

[0019] The behavior recognition device 100 may be configured as a computer equipped with a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a RAM (Random Access Memory), and a ROM (Read Only Memory) storing a program for executing the learning process and inference process in the behavior recognition device, which will be described later.

[0020] The behavior recognition device 100 includes a video information extraction unit 10, a behavior estimation unit 20, an affordance estimation unit 30, a storage unit 40, and a parameter update unit 50. The configuration of each unit involved in the learning process and the inference process by the behavior recognition device 100 will be described in detail below, with the learning process and the inference process being separated.

[0021] <<<Learning process>>> FIG. 2 is a flowchart illustrating an example of a procedure of a learning process performed by the behavior recognition device according to the first embodiment. In the learning process, the video information extraction unit 10 acquires video, i.e., a time series of frame images (sometimes referred to as a series of frame images) (symbol a in FIG. 1), and extracts the video information by converting the video into vector-represented video information, i.e., information indicating the features of the video, and transmits this video information (symbol b in FIG. 1) to the behavior estimation unit 20 and the affordance estimation unit 30 (S10).

[0022] The behavior estimation unit 20 acquires the video information transmitted from the video information extraction unit 10, and uses this video information to generate (estimate) a behavior label distribution (symbol c in Figure 1) indicating the behavior of the subject recorded in the video, and transmits this behavior label distribution to the parameter update unit 50 (S20). For example, in a classification problem where one of three types of actions, "eating," "drinking," and "sitting," is to be recognized, this behavior label distribution is expressed in the form of a probability distribution in which the sum of the real values ​​of each element is 1, such as [0.8, 0.15, 0.05], for video information showing the subject "eating." During the learning process, the parameters of a neural network (described later) are updated so that the probability distribution of the generated behavior label distribution becomes or approaches [1.0, 0.0, 0.0]. During the inference process (described later), when the probability distribution of the generated behavior label distribution is [0.8, 0.15, 0.05] in the above example, the behavior label corresponding to the element with the largest real value, i.e., the behavior label "eat" corresponding to "0.8", may be output from the behavior estimation unit 20.

[0023] The affordance estimation unit 30 acquires the video information transmitted from the video information extraction unit 10, generates (estimates) affordance information (symbol d in Figure 1) corresponding to the whole or part of the frame image related to the extraction of this video information, and transmits this affordance information to the parameter update unit 50 (S30).

[0024] The storage unit 40 stores parameters of the neural networks used by the video information extraction unit 10, the behavior estimation unit 20, and the affordance estimation unit 30.

[0025] The parameter update unit 50 acquires (1) a correct action label distribution (symbol e in FIG. 1) which is correct information on the action label distribution, (2) an action label distribution transmitted from the action estimation unit 20, (3) correct affordance information (symbol e in FIG. 1) which is correct information on the affordance information, and (4) affordance information transmitted from the affordance estimation unit 30. The above-mentioned correct action label distribution is information in which the action element corresponding to the input video in the above-mentioned probability distribution is expressed as "1" and other elements as "0". Then, the parameter update unit 50 calculates the difference between the correct behavior label distribution and the behavior label distribution from the behavior estimation unit 20, and the difference between the correct affordance information and the affordance information from the affordance estimation unit 30, and updates the parameters of the neural network so that these differences become smaller (S40).

[0026] <<<Inference processing>>> FIG. 3 is a flowchart illustrating an example of a procedure of an inference process performed by the behavior recognition device according to the first embodiment. The inference process is similar to the learning process, but there are different processes for the behavior estimation unit 20, the affordance estimation unit 30, and the parameter update unit 50.

[0027] In the inference process, the same process as that during the learning process by the video information extraction unit 10 is performed (S10), and the behavior estimation unit 20 acquires the video information transmitted from the video information extraction unit 10, uses this video information to generate (estimate) a behavior label distribution indicating the behavior of the subject, determines a behavior label (symbol c1 in Figure 1) from this behavior label distribution, and outputs this behavior label to the outside (S25).

[0028] On the other hand, unlike the learning process, in the inference process, the affordance estimation unit 30 does not acquire video information transmitted from the video information extraction unit 10, and does not generate affordance information corresponding to each frame image.

[0029] Unlike the learning process, in the inference process, the parameter update unit 50 does not acquire (1) the correct action label distribution corresponding to the video, (2) the action label distribution transmitted from the action estimation unit 20, (3) the correct affordance information corresponding to the frame image, and (4) the affordance information transmitted from the affordance estimation unit 30. Therefore, the parameter update unit 50 does not calculate the difference between the correct behavior label distribution and the behavior label distribution from the behavior estimation unit 20, and the difference between the correct affordance information and the affordance information from the affordance estimation unit 30, and does not update the parameters.

[0030] Hereinafter, the operations of the units involved in the learning process and the inference process by the behavior recognition device 100 according to the first embodiment will be specifically described, with the learning process and the inference process being separated.

[0031] <<<Learning process (operation of each part)>>> FIG. 4 is a flowchart illustrating an example of a processing procedure performed by the video information extraction unit during the learning process by the behavior recognition device according to the first embodiment. In the learning process, the video information extractor 10 first acquires a video, that is, a sequence of frame images (step S101). Next, the video information extractor 10 converts the acquired video into video information that is a one-dimensional vector (step S102). This video information may be any information that expresses a video as a vector. For example, the video information may be a vector that can be input to and output from SlowFast, which is described in Non-Patent Document 1, where a series of video frame images are input. Next, the video information extraction unit 10 outputs the video information resulting from the above conversion and transmits it to the behavior estimation unit 20 and the affordance estimation unit 30 (step S103).

[0032] FIG. 5 is a flowchart illustrating an example of a procedure of processing by the behavior estimation unit during learning processing by the behavior recognition device according to the first embodiment. The behavior estimation unit 20 first acquires the video information transmitted from the video information extraction unit 10 (step S201). Next, the activity estimation unit 20 generates an activity label distribution of the video from the acquired video information (step S202). The action involved in this generation may be any action that converts video information into an action label distribution.

[0033] In the case of single-label classification, the activity estimation unit 20 converts the video information into a probability distribution, i.e., an activity label distribution, whose dimensionality is the number of types of predefined activity labels, using, for example, a linear layer and a softmax function. In the case of multi-label classification, the behavior estimation unit 20 converts the video information into a distribution (behavior label distribution) whose dimensionality is the number of types of predefined behavior labels, using a linear layer and a sigmoid function, or a linear layer and a tanh function (hyperbolic tangent function). Alternatively, for example, the linear layer may be a network such as a multilayer perceptron. Next, the activity estimation unit 20 outputs the generated activity label distribution of the video to the parameter update unit 50 (step S203).

[0034] FIG. 6 is a flowchart illustrating an example of a processing procedure performed by the affordance estimation unit during the learning process by the behavior recognition device according to the first embodiment. First, the affordance estimation unit 30 acquires the video information transmitted from the video information extraction unit 10 (step S301). Next, the affordance estimation unit 30 generates an adjective distribution corresponding to all or part of the frame images used to acquire the video information from the acquired video information (step S302). This adjective distribution is information in which the state of an object related to the behavior of the subject recorded in the video is expressed as a real-valued vector.

[0035] The operation for this generation may be any operation that converts video information into an adjective distribution corresponding to the target frame image. The affordance estimation unit 30 converts the video information into a distribution whose dimensionality is the number of types of predefined adjective labels, that is, an adjective distribution, using, for example, a linear layer and a softmax function. Alternatively, for example, the linear layer may be a network such as a multilayer perceptron. Next, the affordance estimation unit 30 outputs the generated adjective distribution corresponding to the frame images of the video to the parameter update unit 50 (step S303).

[0036] As described above, the memory unit 40 stores parameters of the neural networks used by the video information extraction unit 10, the behavior estimation unit 20, and the affordance estimation unit 30.

[0037] FIG. 7 is a flowchart illustrating an example of a procedure of processing by the parameter update unit during learning processing by the behavior recognition device according to the first embodiment. The parameter update unit 50 first acquires (1) a correct action label distribution in which actions corresponding to the input video are expressed as "1" and others as "0," i.e., correct information of the action label distribution; (2) an action label distribution generated by the action estimation unit 20; (3) a correct adjective distribution in which the state of an object related to the action captured in the frame image is expressed as a real-valued vector; and (4) an adjective distribution in which the state of an object related to the action captured in the frame image is expressed as a real-valued vector, generated by the affordance estimation unit 30 (S351).

[0038] The correct adjective distribution for the nth (1≦N) frame of a video with a frame length of N is obtained as a distribution in which the element corresponding to the initial state is "1-(n-1 / N-1)", the element corresponding to the final state is "n-1 / N-1", and all other elements are 0. The transition from the initial state to the final state means that, for example, when a subject takes the action of "opening a door," the object "door" is initially in a "closed state (initial state)," and the action of "opening the door" ultimately causes the object "door" to transition to an "open state (final state)."

[0039] Next, the parameter update unit 50 updates the parameters of the neural network in the storage unit 40 based on the constraints (S352). Specifically, the above constraints refer to constraints that are imposed when updating each parameter of the neural network so that the shape of the correct behavioral label distribution matches the shape of the behavioral label distribution from the behavior estimation unit 20, and so that the shape of the correct adjective distribution matches the shape of the adjective distribution from the affordance estimation unit 30. Any learning method may be used to update the parameters as long as it is set to satisfy these constraints.

[0040] For example, the parameter update unit 50 may update the correct behavior label distribution as p(x act ), and the behavior label distribution from the behavior estimation unit 20 is q(x act ), and the correct adjective distribution is t(x adj ), and the adjective distribution from the affordance estimation unit 30 is u(x adj ), the loss can be calculated according to the following equation (1), and the parameters stored in the storage unit 40 can be updated so as to reduce this loss. Here, γ in equation (1) represents a hyperparameter that adjusts the influence of the affordance estimation unit 30, and img in equation (1) represents a frame image of the video.

[0041]

number

[0042] <<<Inference process (operation of each part)>>> The inference process according to the first embodiment is different from the learning process according to the first embodiment in that some of the processes of the behavior estimation unit 20, the affordance estimation unit 30, and the parameter update unit 50 are different.

[0043] FIG. 8 is a flowchart illustrating an example of a processing procedure by the behavior estimation unit during the inference processing by the behavior recognition device according to the first embodiment. In the inference process, the behavior estimation unit 20 first acquires the video information transmitted from the video information extraction unit 10 (step S401). Next, the activity estimation unit 20 generates an activity label distribution of the video from the acquired video information (step S402). As in the learning process, the operation related to this generation may be any operation that converts the video information into an activity label distribution.

[0044] Next, the activity estimation unit 20 determines an activity label for the video from the generated activity label distribution (step S403). In the case of single-label classification, the activity estimation unit 20 can, for example, select the activity label corresponding to the element with the largest value in the activity label distribution. In the case of multi-label classification, for example, the threshold θ l is defined in advance, and the behavior estimation unit 20 determines the threshold θ l Action labels for videos corresponding to elements having values ​​equal to or greater than 1 can be selected. Next, the activity estimation unit 20 outputs the generated activity label of the video to the outside (step S404).

[0045] In the inference process, unlike the learning process, the affordance estimation unit 30 does not acquire video information transmitted from the video information extraction unit 10, and does not generate an adjective distribution corresponding to each frame image.

[0046] In the inference process, unlike the learning process, the parameter update unit 50 does not acquire (1) the correct action label distribution corresponding to the video, (2) the action label distribution transmitted from the action estimation unit 20, (3) the correct adjective distribution corresponding to the frame image, and (4) the adjective distribution transmitted from the affordance estimation unit 30. Therefore, the parameter update unit 50 does not calculate the difference between these two types of distributions, i.e., (1) the difference between the correct behavior label distribution and the behavior label distribution from the behavior estimation unit 20, and (2) the difference between the correct adjective distribution and the adjective distribution from the affordance estimation unit 30, and does not update the parameters.

[0047] In the first embodiment described above, by learning the changes in the affordance of the object that the actor is acting on along with the actor's actions, the actor's actions can be estimated with higher accuracy regardless of the type of video, such as first-person perspective or third-person perspective.

[0048] (Second embodiment) Next, a description will be given of a second embodiment of the present invention, with reference to Fig. 9, which is a block diagram showing an example of the configuration of a behavior recognition device according to the second embodiment. The behavior recognition device 100a according to the second embodiment includes a video information extraction unit 10, a behavior estimation unit 20, an affordance estimation unit 30a, a storage unit 40, and a parameter update unit 50a. That is, compared to the first embodiment, the second embodiment includes an affordance estimation unit 30a and a parameter update unit 50a instead of the affordance estimation unit 30 and the parameter update unit 50.

[0049] Hereinafter, the operations of each unit involved in the learning process and the inference process by the behavior recognition device 100a according to the second embodiment will be specifically described, dividing them into the learning process and the inference process. Note that the operations of parts not specifically described are the same as those of the first embodiment.

[0050] <<<Learning process (operation of each part)>>> FIG. 10 is a flowchart illustrating an example of a processing procedure performed by the affordance estimation unit during the learning process by the behavior recognition device according to the second embodiment. The affordance estimation unit 30a first acquires the video information transmitted from the video information extraction unit 10 (step S501). Next, the affordance estimation unit 30a generates from the acquired video information (1) an object label distribution (sometimes referred to as a noun label distribution) indicating the objects corresponding to the action labels, and (2) an adjective distribution corresponding to all or part of the frame images used when acquiring the video information (step S502). For example, when the predefined objects are three types, "stapler," "scissors," and "pencil," the object label distribution is output as a probability distribution where the sum of the real values ​​of each element is 1, such as [0.25, 0.7, 0.05], for video information recording the subject cutting paper with "scissors."

[0051] The operation involved in this generation may be any operation that converts video information into (1) an object label distribution corresponding to an action label, and (2) an adjective distribution corresponding to a target frame image.

[0052] For example, the affordance estimation unit 30a uses a linear layer and a softmax function to generate a distribution of video information whose dimensionality is the number of predefined objects, i.e., an object label distribution, and uses another linear layer and a softmax function to generate a distribution whose dimensionality is the number of predefined types of adjective labels, i.e., an adjective distribution. Alternatively, for example, each of the linear layers may be a network such as a multilayer perceptron.

[0053] Next, the affordance estimation unit 30a outputs the generated (1) object label distribution corresponding to the action labels of the video and (2) adjective distribution corresponding to the frame images to the parameter update unit 50a (step S503).

[0054] As described above, the storage unit 40 also stores parameters of the neural networks used by the video information extraction unit 10, the behavior estimation unit 20, and the affordance estimation unit 30a.

[0055] FIG. 11 is a flowchart illustrating an example of a procedure of processing by a parameter update unit during learning processing by the behavior recognition device according to the second embodiment. The parameter update unit 50a first acquires (1) a correct action label distribution in which actions corresponding to the input video are expressed as "1" and others as "0", (2) an action label distribution generated by the action estimation unit 20, (3) a correct object label distribution in which objects corresponding to the action labels of the input video are expressed as "1" and others as "0", i.e., correct information of the object label distribution, (4) an object label distribution generated by the affordance estimation unit 30a, (5) a correct adjective distribution in which the state of an object related to the action captured in the frame image is expressed as a real-valued vector, and (6) an adjective distribution in which the state of an object related to the action captured in the frame image is expressed as a real-valued vector, generated by the affordance estimation unit 30a (S551).

[0056] Next, the parameter update unit 50a updates the parameters of the neural network in the storage unit 40 based on the constraints (S552). Specifically, the above constraints are constraints imposed when updating each parameter of the neural network so that the correct behavioral label distribution matches the shape of the behavioral label distribution from the behavior estimation unit 20, the correct object label distribution matches the shape of the object label distribution from the affordance estimation unit 30a, and the correct adjective distribution matches the shape of the adjective distribution from the affordance estimation unit 30a. Any learning method may be used to update the parameters as long as it is set to satisfy these constraints.

[0057] For example, the parameter update unit 50a may update the correct behavior label distribution as p(x act ), and the behavior label distribution from the behavior estimation unit 20 is q(x act ), and the correct object label distribution is r(x noun ), and the object label distribution from the affordance estimation unit 30a is defined as s(x noun ), and the correct adjective distribution is t(x adj ), and the adjective distribution from the affordance estimation unit 30a is u(x adj ), the loss is calculated according to the following equation (2), and the parameters stored in the storage unit 40 can be updated so as to reduce this loss. In the formula (2), γ represents a hyperparameter that adjusts the influence of the affordance estimation unit 30a, and img represents a frame image of the video.

[0058]

number

[0059] <<<Inference process (operation of each part)>>> The inference process according to the second embodiment is different from the learning process according to the second embodiment in the processing of the affordance estimation unit 30a and the parameter update unit 50a.

[0060] Specifically, in the inference process, unlike the learning process, the affordance estimation unit 30a does not acquire video information transmitted from the video information extraction unit 10, and does not generate an object label distribution corresponding to the action labels of the video and an adjective distribution corresponding to each frame image.

[0061] In the inference process, unlike the learning process, the parameter update unit 50a does not acquire (1) the correct action label distribution corresponding to the video, (2) the action label distribution transmitted from the action estimation unit 20, (3) the correct object label distribution corresponding to the correct action label, (4) the object label distribution transmitted from the affordance estimation unit 30a, (5) the correct adjective distribution corresponding to the frame image, and (6) the adjective distribution transmitted from the affordance estimation unit 30a.

[0062] Therefore, the parameter update unit 50a does not calculate the differences between these three types of distributions, i.e., (1) the difference between the correct behavior label distribution and the behavior label distribution from the behavior estimation unit 20, (2) the difference between the correct object label distribution and the object label distribution from the affordance estimation unit 30a, and (3) the difference between the correct adjective distribution and the adjective distribution from the affordance estimation unit 30a, and does not update the parameters.

[0063] In the second embodiment described above, in addition to the configuration described in the first embodiment, an object label distribution is generated as a type of affordance information, and parameters are updated taking this object label distribution into consideration, so that the behavior of the performer can be estimated with even higher accuracy.

[0064] FIG. 12 is a block diagram showing an example of the hardware configuration of a behavior recognition device according to one embodiment of the present invention. 12, the behavior recognition device 100 according to the above embodiment is configured, for example, by a server computer or a personal computer, and has a hardware processor 111A such as a CPU. A program memory 111B, a data memory 112, an input / output interface 113, and a communication interface 114 are connected to this hardware processor 111A via a bus 115. The same applies to the behavior recognition device 100a shown in FIG.

[0065] The communication interface 114 includes, for example, one or more wireless communication interface units, and enables transmission and reception of information to and from the communication network NW. As the wireless interface, for example, an interface that adopts a low-power wireless data communication standard such as a wireless LAN (Local Area Network) is used.

[0066] To the input / output interface 113, an input device 200 and an output device 300 that are attached to the behavior recognition device 100 and used by a user or the like are connected. The input / output interface 113 takes in operation data input by a user or the like through an input device 200 such as a keyboard, a touch panel, a touchpad, or a mouse, and outputs output data to an output device 300 including a display device using a liquid crystal or organic EL (Electro Luminescence) for display. Note that the input device 200 and the output device 300 may be devices built into the behavior recognition device 100, or may be input devices and output devices of other information terminals that can communicate with the behavior recognition device 100 via a network NW.

[0067] The program memory 111B is a non-transitory tangible storage medium that is a combination of a non-volatile memory that can be written to and read from at any time, such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive), and a non-volatile memory such as a ROM, and stores programs necessary to execute various control processes, etc., according to one embodiment.

[0068] The data memory 112 is a tangible storage medium, for example, a combination of the above-mentioned nonvolatile memory and a volatile memory such as RAM, and is used to store various data acquired and created during various processing steps.

[0069] The behavior recognition device 100 according to one embodiment of the present invention can be configured as a data processing device having the units shown in FIG. 1 as software-based processing function units.

[0070] The storage device used as a working memory by each unit of the behavior recognition device 100 and the storage device used as the storage unit 40 can be configured by using the data memory 112 shown in Fig. 12. However, these configured storage areas are not essential components within the behavior recognition device 100, and may be areas provided in an external storage medium such as a USB (Universal Serial Bus) memory, or a storage device such as a database server located in the cloud.

[0071] The processing function units in each of the above units can be realized by reading and executing a program stored in the program memory 111B by the hardware processor 111A. Note that some or all of these processing function units may be realized in various other forms, including integrated circuits such as an application specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).

[0072] The methods described in each embodiment may be stored as a program (software means) that can be executed by a computer on a recording medium such as a magnetic disk (e.g., a floppy disk, a hard disk, etc.), an optical disk (e.g., a CD-ROM, a DVD, an MO, etc.), or a semiconductor memory (e.g., a ROM, a RAM, a flash memory, etc.), or may be transmitted and distributed via a communication medium. The program stored on the medium also includes a configuration program that configures the software means (including not only execution programs but also tables and data structures) that the computer executes. The computer that realizes this device reads the program stored on the recording medium and, in some cases, configures the software means using the configuration program, and executes the above-described processing by having the operation controlled by this software means. The term "recording medium" as used herein is not limited to a storage medium for distribution, but also includes a storage medium such as a magnetic disk or semiconductor memory installed inside the computer or in a device connected via a network.

[0073] The present invention is not limited to the above-described embodiments, and various modifications can be made in the implementation stage without departing from the spirit of the invention. Furthermore, the embodiments may be implemented in appropriate combinations, in which case the combined effects can be obtained. Furthermore, the above-described embodiments include various inventions, and various inventions can be extracted by combining selected elements from the disclosed elements. For example, if the problem can be solved and the desired effect can be obtained even if some elements are deleted from all elements shown in the embodiments, the configuration from which these elements are deleted can be extracted as an invention. [Explanation of symbols]

[0074] 100, 100a...Behavior recognition device 10...Video information extraction unit 20…Behavior estimation section 30, 30a...Affordance estimation section 40...Storage section 50, 50a...Parameter update section

Claims

1. a video information extraction unit that extracts video information in which a video composed of time-series frame images is expressed as a vector; a behavior estimation unit that estimates information related to the behavior of a subject recorded in the video based on the video information extracted by the video information extraction unit and model parameters; Equipped with the parameters are updated based on information on affordances of an object related to the behavior of the subject, which is estimated based on the video information extracted by the video information extraction unit, and information on the behavior estimated by the behavior estimation unit. Estimation device.

2. an affordance estimation unit that estimates information related to an affordance of an object associated with the subject's behavior based on the video information extracted by the video information extraction unit of the estimation device according to claim 1; a parameter update unit that updates parameters of the model based on a comparison result between the information on the behavior estimated by the behavior estimation unit of the estimation device according to claim 1 and correct answer information on the information on the behavior, and a comparison result between the information on the affordance estimated by the affordance estimation unit and correct answer information on the information on the affordance; A learning device comprising:

3. The affordance estimation unit an adjective distribution in which a state of an object related to the behavior of the subject is expressed as a vector based on the video information extracted by the video information extraction unit, as information related to the affordance; The learning device according to claim 2 .

4. The affordance estimation unit an adjective distribution in which a state of an object related to the behavior of the subject is expressed as a vector based on the video information extracted by the video information extraction unit, and information related to an object indicated by information related to the behavior of the subject based on the video information extracted by the video information extraction unit, are estimated as the information related to the affordance; The learning device according to claim 2 .

5. The parameter update unit updating parameters of the model based on a comparison result between the information related to the behavior estimated by the behavior estimation unit and correct answer information about the information related to the behavior, and a comparison result between the adjective distribution estimated by the affordance estimation unit and correct answer information about the adjective distribution; The learning device according to claim 3 .

6. The parameter update unit updating parameters of the model based on a comparison result between the information related to the behavior estimated by the behavior estimation unit and correct answer information for the information related to the behavior, a comparison result between the adjective distribution estimated by the affordance estimation unit and correct answer information for the adjective distribution, and a comparison result between the information related to the object estimated by the affordance estimation unit and correct answer information for the information related to the object; The learning device according to claim 4 .

7. A method performed by an estimation device, comprising: extracting video information in which a video composed of a time series of frame images is expressed as a vector by a video information extraction unit of the estimation device; an action estimation unit of the estimation device estimates information related to the action of the subject recorded in the video based on the video information extracted by the video information extraction unit and model parameters; the parameters are updated based on information on affordances of an object related to the behavior of the subject, which is estimated based on the video information extracted by the video information extraction unit, and information on the behavior estimated by the behavior estimation unit. Estimation method.

8. A program that causes a processor to function as each of the components of the estimation device according to claim 1 or the learning device according to any one of claims 2 to 6.

Citation Information

Patent Citations

  • Scene recognizing device

    JP2000293685A

  • Action / intention presumption system, action / intention presumption method, action / intention pesumption program and computer-readable recording medium with program recorded thereon

    JP2005242759A

  • Action estimation device and method

    JP2010170212A

  • Behavior estimation device, behavior estimation method, program and behavior estimation system

    JP2022072825A