Behavior estimation device, behavior estimation method, and recording medium

The behavior estimation device enhances action prediction accuracy by integrating person, object, and peripheral features, addressing the limitations of existing methods by considering contextual relevance.

JP7865106B2Active Publication Date: 2026-05-26NEC CORP

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NEC CORP
Filing Date
2022-06-09
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing behavior estimation methods fail to accurately predict human actions due to the lack of consideration for the relevance between individuals and surrounding objects, leading to reduced estimation accuracy.

Method used

A behavior estimation device that integrates person, object, and peripheral features through a series of aggregation processes, including feature extraction, integration, and estimation using neural networks to enhance accuracy.

Benefits of technology

Improves the accuracy of predicting human actions by considering the relevance of objects and surroundings, maintaining estimation precision even when relevant objects are absent or undetectable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007865106000001
    Figure 0007865106000001
  • Figure 0007865106000002
    Figure 0007865106000002
  • Figure 0007865106000003
    Figure 0007865106000003
Patent Text Reader

Abstract

To provide an action estimation device and the like which can improve estimation accuracy when estimating the action of a person.SOLUTION: In an action estimation device, person feature extraction means extracts a feature of a person detected from a plurality of time-series images, object feature extraction means extracts a feature of an object detected from the plurality of time-series images, periphery feature extraction means extracts a feature of the periphery of the person in the plurality of time-series images, feature aggregation means performs aggregation processing for aggregating the feature of the person, the feature of the object and the feature of the periphery of the person, and action estimation processing means performs processing of estimating the action of the person included in the plurality of images on the basis of information including the processing result of the aggregation processing.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a technology that can be used for estimating human behavior.

Background Art

[0002] Techniques for estimating the behavior of a person in an image have been conventionally known.

[0003] Specifically, for example, in Patent Document 1, a posture feature of a person shown in an image generated by an imaging device is extracted, a peripheral feature indicating the shape, position, or type of an object around the person shown in the image is extracted, the peripheral feature is filtered based on the posture feature and the importance of the peripheral feature set in association with the posture feature, and a behavior class of the person shown in the image is estimated based on the posture feature and the filtered peripheral feature.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, according to the method disclosed in Patent Document 1, for example, there is a problem that the behavior of the person may not be correctly estimated because the relevance between the person in the image and the objects around the person is not considered. Further, according to the method disclosed in Patent Document 1, for example, there is a problem that the behavior of the person may not be correctly estimated when an object around the person cannot be detected.

[0006] That is, according to the method disclosed in Patent Document 1, there is a problem that the estimation accuracy when estimating the behavior of a person is reduced.

[0007] One objective of this disclosure is to provide a behavior estimation device that can improve the estimation accuracy when estimating a person's behavior. [Means for solving the problem]

[0008] In one aspect of this disclosure, the behavior estimation device comprises: a person feature extraction means for extracting person features detected from a plurality of time-series images; an object feature extraction means for extracting object features detected from the plurality of images; a peripheral feature extraction means for extracting peripheral features of the person in the plurality of images; a feature aggregation means for performing aggregation processing to aggregate the person's features, the object's features, and the peripheral features of the person; and a behavior estimation processing means for performing processing to estimate the person's behavior included in the plurality of images based on information including the processing results of the aggregation processing. The feature aggregation means performs the following aggregation processes: first, by performing a first process to integrate the features of the person with the features of the person's surroundings, to obtain a first processing result; second, by performing a second process to integrate the features of an object with the first processing result, to obtain a second processing result; and third, by performing a third process to integrate the features of the person's surroundings with the second processing result, to obtain a third processing result.

[0009] In other aspects of this disclosure, the behavior estimation device includes: a person feature extraction means for extracting person features detected from a plurality of time-series images; an object feature extraction means for extracting object features detected from the plurality of images; a peripheral feature extraction means for extracting peripheral features of the person in the plurality of images; a feature integration means for performing integration processing to integrate the person features, the object features, and the peripheral features of the person; an aggregation processing means for performing aggregation processing to aggregate the person features and the processing results of the integration processing; and a behavior estimation processing means for estimating the behavior of the person included in the plurality of images based on information including the processing results of the aggregation processing.

[0010] In yet another aspect of this disclosure, the behavior estimation method extracts features of a person detected from a series of images, extracts features of an object detected from the images, extracts features of the person's surroundings in the images, performs an aggregation process to combine the features of the person, the features of the object, and the features of the person's surroundings, and estimates the behavior of the person included in the images based on the information including the results of the aggregation process. A method for estimating behavior, wherein the aggregation process involves obtaining a first processing result by performing a first processing that integrates the characteristics of the person with the characteristics of the person's surroundings; obtaining a second processing result by performing a second processing that integrates the characteristics of the object with the first processing result; and obtaining a third processing result by performing a third processing that integrates the characteristics of the person's surroundings with the second processing result.

[0011] In yet another aspect of this disclosure, the behavior estimation method extracts features of a person detected from a series of images, extracts features of an object detected from the images, extracts features of the person's surroundings in the images, performs an integration process to combine the features of the person, the features of the object, and the features of the person's surroundings, performs an aggregation process to aggregate the features of the person and the results of the integration process, and estimates the behavior of the person included in the images based on the information including the results of the aggregation process.

[0012] In yet another aspect of this disclosure, the recording medium extracts the characteristics of a person detected from a series of images, extracts the characteristics of an object detected from the series of images, extracts the characteristics of the person's surroundings in the series of images, performs an aggregation process to aggregate the characteristics of the person, the characteristics of the object, and the characteristics of the person's surroundings, and causes a computer to perform a process to estimate the actions of the person included in the series of images based on the information including the processing results of the aggregation process. The computer is instructed to perform the following processes as part of the aggregation process: first, perform a first process to integrate the features of the person with the features of the person's surroundings to obtain a first processing result; second, perform a second process to integrate the features of the object with the first processing result to obtain a second processing result; and third, perform a third process to integrate the features of the person's surroundings with the second processing result to obtain a third processing result. Record the program.

[0013] In yet another aspect of this disclosure, the storage medium records a program that causes a computer to perform a process to extract features of a person detected from a series of images, extract features of an object detected from the series of images, extract features of the person's surroundings in the series of images, perform an integration process to combine the features of the person, the features of the object, and the features of the person's surroundings, perform an aggregation process to aggregate the features of the person and the results of the integration process, and estimate the actions of the person included in the series of images based on information including the results of the aggregation process. [Effects of the Invention]

[0014] This disclosure makes it possible to improve the accuracy of estimations when predicting a person's actions. [Brief explanation of the drawing]

[0015] [Figure 1]A diagram showing the outline of the behavior estimation device according to the first embodiment. [Figure 2] A block diagram showing the hardware configuration of the behavior estimation device according to the first embodiment. [Figure 3] A block diagram showing the functional configuration of the behavior estimation device according to the first embodiment. [Figure 4] A diagram showing an example of the configuration of the peripheral feature extraction unit included in the behavior estimation device according to the first embodiment. [Figure 5] A diagram showing an example of the configuration of the feature aggregation unit included in the behavior estimation device according to the first embodiment. [Figure 6] A diagram showing an example of the configuration of the aggregation processing unit included in the behavior estimation device according to the first embodiment. [Figure 7] A flowchart for explaining the processing performed in the behavior estimation device according to the first embodiment. [Figure 8] A diagram showing an example of the configuration of the feature aggregation unit included in the behavior estimation device according to a modification of the first embodiment. [Figure 9] A diagram showing a configuration example when a plurality of feature integration units are provided in the feature aggregation unit of FIG. 8. [Figure 10] A diagram showing an example of the configuration of the feature aggregation unit included in the behavior estimation device according to a modification of the first embodiment. [Figure 11] A block diagram showing the functional configuration of the behavior estimation device according to the second embodiment. [Figure 12] A flowchart for explaining the processing performed in the behavior estimation device according to the second embodiment.

Mode for Carrying Out the Invention

[0016] Hereinafter, preferred embodiments of the present disclosure will be described with reference to the drawings.

[0017] <First Embodiment> [Schematic Configuration] Figure 1 is a schematic diagram of the behavior estimation device according to the first embodiment. The behavior estimation device 100 is composed of a device such as a personal computer. The behavior estimation device 100 also performs behavior estimation processing to estimate the behavior of a person contained in video footage captured by a camera or the like. The behavior estimation device 100 also outputs the estimation results obtained from the aforementioned behavior estimation processing to an external device.

[0018] [Hardware configuration] Figure 2 is a block diagram showing the hardware configuration of the behavior estimation device according to the first embodiment. As shown in Figure 2, the behavior estimation device 100 includes an interface (IF) 111, a processor 112, a memory 113, a recording medium 114, and a database (DB) 115.

[0019] IF111 performs data input and output with external devices. Video captured by cameras, etc., is input to the behavior estimation device 100 via IF111. The estimation results obtained by the behavior estimation device 100 are output to external devices via IF111 as needed.

[0020] The processor 112 is a computer such as a CPU (Central Processing Unit) and controls the entire behavior estimation device 100 by executing a pre-prepared program. Specifically, the processor 112 performs processing such as behavior estimation processing.

[0021] Memory 113 consists of ROM (Read Only Memory), RAM (Random Access Memory), and other components. Memory 113 is also used as working memory while the processor 112 is executing various processes.

[0022] The recording medium 114 is a non-volatile, non-temporary recording medium such as a disk-shaped recording medium or semiconductor memory, and is configured to be detachable from the behavior estimation device 100. The recording medium 114 stores various programs that the processor 112 executes. When the behavior estimation device 100 performs various processes, the programs stored in the recording medium 114 are loaded into the memory 113 and executed by the processor 112.

[0023] DB115 stores, for example, information input through IF111 and processing results obtained from the processor 112.

[0024] [Functional Configuration] Figure 3 is a block diagram showing the functional configuration of the behavior estimation device according to the first embodiment. As shown in Figure 3, the behavior estimation device 100 includes an image acquisition unit 11, a person region detection unit 12, a person feature extraction unit 13, an object region detection unit 14, an object feature extraction unit 15, a surrounding feature extraction unit 16, a feature aggregation unit 17, and a behavior estimation processing unit 18.

[0025] The video acquisition unit 11 acquires video captured by a camera or the like, and outputs the acquired video to the person area detection unit 12, the object area detection unit 14, and the surrounding feature extraction unit 16.

[0026] In the following explanation, unless otherwise specified, it will be assumed that the video acquired by the video acquisition unit 11 includes one or more people and one or more objects.

[0027] The person region detection unit 12 performs processing to detect people in each of the multiple time-series images contained in the video obtained by the video acquisition unit 11. The person region detection unit 12 also generates person detection information, which is information that can identify the region in the image corresponding to the person detected by the above processing as a person region, and outputs the generated person detection information to the person feature extraction unit 13. A person region may be, for example, a rectangular region that individually surrounds one or more people in the image.

[0028] The person feature extraction unit 13 is configured to perform processing using a neural network (hereinafter abbreviated as NN), such as a CNN (Convolutional Neural Network). The person feature extraction unit 13 identifies person regions in the image based on the person detection information obtained by the person region detection unit 12, and performs processing to extract the features of the person included in the identified person region. Specifically, the person feature extraction unit 13 extracts person features, which are feature quantities that satisfy predetermined conditions from among a plurality of feature quantities calculated based on the pixel values ​​of each pixel included in the person region, as the features of the person included in that person region. Furthermore, if the person detection information includes multiple person regions, the person feature extraction unit 13 performs the above processing for each of the multiple person regions. Finally, the person feature extraction unit 13 generates person feature information, which is information relating to the person features calculated in the person region in the image, and outputs the generated person feature information to the feature aggregation unit 17.

[0029] The object region detection unit 14 performs processing to detect predetermined objects other than people in each of the multiple time-series images contained in the video obtained by the video acquisition unit 11. The object region detection unit 14 also generates object detection information, which is information that can identify the region in the image corresponding to the predetermined object detected by the above processing as an object region, and outputs the generated object detection information to the object feature extraction unit 15. An object region may be, for example, a rectangular region that individually surrounds one or more objects in the image. Furthermore, if, for example, video footage of a construction site is obtained by the video acquisition unit 11, the object region detection unit 14 may perform processing to detect objects such as an excavator and a shovel as predetermined objects.

[0030] The object feature extraction unit 15 is configured to perform processing using a neural network (NN), such as a CNN. The object feature extraction unit 15 identifies object regions in the image based on object detection information obtained by the object region detection unit 14, and performs processing to extract the features of objects contained in the identified object region. Specifically, the object feature extraction unit 15 extracts object features, which are features that satisfy predetermined conditions from among a plurality of feature quantities calculated based on the pixel values ​​of each pixel contained in the object region, as features of the objects contained in the object region. Furthermore, if the object detection information contains multiple object regions, the object feature extraction unit 15 performs the above processing for each of the multiple object regions. Finally, the object feature extraction unit 15 generates object feature information, which is information relating to the object features calculated in the object regions in the image, and outputs the generated object feature information to the feature aggregation unit 17.

[0031] The peripheral feature extraction unit 16 extracts features from each of the multiple time-series images contained in the video obtained by the video acquisition unit 11, and outputs the information representing the extracted features as peripheral feature information to the feature aggregation unit 17. The peripheral feature extraction unit 16 also includes, for example, a feature extraction unit 161 and a region division unit 162, as shown in Figure 4. Figure 4 is a diagram showing an example of the configuration of the peripheral feature extraction unit included in the behavior estimation device according to the first embodiment.

[0032] The feature extraction unit 161 is configured to perform processing using a neural network (NN), such as a CNN. The feature extraction unit 161 extracts peripheral features from the entire image obtained by the video acquisition unit 11, which are features that satisfy predetermined conditions among a plurality of feature quantities calculated based on pixel values, as features of the image. In other words, the feature extraction unit 161 extracts the following as image features included in the video output from the video acquisition unit 11: features of people included in the image, features of objects included in the image, and features of the background area other than the people and objects. The aforementioned background features may be rephrased as features of the area surrounding the person detected by the person region detection unit 12. The aforementioned background features include, for example, features of objects located at a certain distance or more from the person detected by the person region detection unit 12, and features of objects other than predetermined objects detected by the object region detection unit 14.

[0033] The region division unit 162 performs a process of dividing the region corresponding to each of the multiple features extracted by the feature extraction unit 161 into rectangular regions. The region division unit 162 also generates peripheral feature information, which is information related to the peripheral feature quantities calculated in each rectangular region divided by the region division described above, and outputs the generated peripheral feature information to the feature aggregation unit 17.

[0034] In this embodiment, the peripheral feature extraction unit 16 is not limited to a configuration in which the region division unit 162 is provided after the feature extraction unit 161, but may also be configured in which the feature extraction unit 161 is provided after the region division unit 162. In such a case, for example, the region division unit 162 may perform the process of dividing the entire image contained in the video output from the video acquisition unit 11 into a plurality of rectangular regions, and the feature extraction unit 161 may perform the process of extracting a region representing the features of the image from the plurality of rectangular regions, and the process of generating peripheral feature information, which is information relating to the peripheral feature quantities calculated in the extracted region.

[0035] The feature aggregation unit 17 performs processing to aggregate a sequence of person features corresponding to multiple person features contained in n (n≧2) person feature information obtained in time series from the person feature extraction unit 13, a sequence of object features corresponding to multiple object features contained in n object feature information obtained in time series from the object feature extraction unit 15, and a sequence of peripheral features corresponding to multiple peripheral features contained in n peripheral feature information obtained in time series from the peripheral feature extraction unit 16. The feature aggregation unit 17 also outputs a feature vector obtained according to the processing result of the above processing to the behavior estimation processing unit 18. The above feature vector can be rephrased as information including the processing result of the aggregation processing performed by the feature aggregation unit 17. Furthermore, the feature aggregation unit 17 includes, for example, time-series feature integration units 171A, 171B, and 171C, and an aggregation processing unit 172, as shown in Figure 5. Figure 5 is a diagram showing an example of the configuration of the feature aggregation unit included in the behavior estimation device according to the first embodiment.

[0036] The time-series feature integration unit 171A performs processing to integrate multiple person features included in the person feature sequence into one. Specifically, the time-series feature integration unit 171A integrates multiple person features into one by, for example, calculating the average value of multiple person features included in the person feature sequence. Alternatively, the time-series feature integration unit 171A integrates multiple person features into one by, for example, performing weighted addition using a neural network (NN) with a Source-Target Attention mechanism. When processing using an NN with a Source-Target Attention mechanism is performed in the time-series feature integration unit 171A, for example, the person features to be weighted should be input to the query of the mechanism, and the same person feature sequence should be input to the key and value of the mechanism. Furthermore, the time-series feature integration unit 171A performs processing to integrate multiple person features into one for each person that can be identified based on the person feature information or the person feature sequence. Therefore, if there are k (k≧1) people who can be identified based on the person feature information or the person feature sequence, the time-series feature integration unit 171A performs a process to integrate multiple person features into one, thereby obtaining k integrated person features.

[0037] The time-series feature integration unit 171B performs processing to integrate multiple object features included in the object feature sequence into one. Specifically, the time-series feature integration unit 171B integrates multiple object features into one by, for example, calculating the average value of multiple object features included in the object feature sequence. Alternatively, the time-series feature integration unit 171B integrates multiple object features into one by, for example, performing weighted addition using a NN with a Source-Target Attention mechanism. When processing using a NN with a Source-Target Attention mechanism is performed in the time-series feature integration unit 171B, for example, the object features to be weighted should be input to the query of the mechanism, and the same object feature sequence should be input to the key and value of the mechanism. Furthermore, the time-series feature integration unit 171B performs processing to integrate multiple object features into one for each object that can be identified based on the object feature information or object feature sequence. Therefore, if there are m (m≧1) objects that can be identified based on the object feature information or object feature sequence, the time-series feature integration unit 171B performs processing to integrate multiple object features into one, thereby obtaining m integrated object features.

[0038] The time-series feature integration unit 171C performs processing to integrate multiple peripheral features included in the peripheral feature sequence into one. Specifically, the time-series feature integration unit 171C integrates multiple peripheral features into one by, for example, calculating the average value of multiple peripheral features included in the peripheral feature sequence. Alternatively, the time-series feature integration unit 171C integrates multiple peripheral features into one by, for example, performing weighted summation using multiple peripheral features included in the peripheral feature sequence and a neural network (NN) having a Source-Target Attention mechanism. When processing using an NN with a Source-Target Attention mechanism is performed in the time-series feature integration unit 171C, for example, the peripheral features to be weighted should be input to the mechanism's query, and the same peripheral feature sequence should be input to the mechanism's key and value. Furthermore, the time-series feature integration unit 171C performs processing to integrate multiple peripheral features into one for each rectangular region that can be identified based on the peripheral feature information or peripheral feature sequence. Therefore, if there are p (p≧1) rectangular regions that can be identified based on the surrounding feature information or the sequence of surrounding features, the time-series feature integration unit 171C performs a process to integrate multiple surrounding features into one, thereby obtaining p integrated surrounding features.

[0039] The aggregation processing unit 172 performs processing to aggregate k person features obtained by the time-series feature integration unit 171A, m object features obtained by the time-series feature integration unit 171B, and p surrounding features obtained by the time-series feature integration unit 171C. The aggregation processing unit 172 also acquires a feature vector corresponding to the processing result of the above-mentioned processing and outputs the acquired feature vector to the behavior estimation processing unit 18. In this embodiment, multiple aggregation processing units 172 may be connected in series, as long as the feature vector generated according to the processing result of the above-mentioned processing is output to the behavior estimation processing unit 18. Furthermore, the aggregation processing unit 172 has, for example, feature integration units 172A, 172B, and 172C, as shown in Figure 6. Figure 6 is a diagram showing an example of the configuration of the aggregation processing unit included in the behavior estimation device according to the first embodiment.

[0040] The feature integration unit 172A obtains a first processing result by performing a first process to integrate the features of a person with the features of the person surrounding that person, and outputs the obtained first processing result to the feature integration unit 172B. The feature integration unit 172A also has a pre-trained neural network (NN) that integrates person features and surrounding features to obtain scene features, which are features associated with the surrounding features for the person features. Furthermore, the feature integration unit 172A obtains k scene features corresponding to each of the k person features by inputting k person features and p surrounding features into an NN having a Source-Target Attention mechanism and performing weighted addition. Specifically, the feature integration unit 172A obtains k scene features corresponding to each of the k person features by inputting one person feature to be weighted into the query of the Source-Target Attention mechanism, and inputting a feature sequence having p surrounding features into the key and value of the mechanism and performing weighted addition k times. In other words, scene features are obtained as features that represent the characteristics of the shooting scene (for example, the background at the time of shooting) for each of the k people included in the video captured by a camera or the like.

[0041] The feature integration unit 172B obtains a second processing result by performing a second processing that integrates object features with the first processing result output from the feature integration unit 172A, and outputs the obtained second processing result to the feature integration unit 172C. The feature integration unit 172B also has a pre-trained neural network (NN) that integrates scene features and object features to obtain situational features, which are features that associate the object features with the scene features. Furthermore, the feature integration unit 172B inputs the k scene features and m object features obtained by the feature integration unit 172A into an NN having a Source-Target Attention mechanism and performs weighted addition to obtain k situational features corresponding to each of the k scene features. Specifically, the feature integration unit 172A inputs one scene feature to be weighted into the Source-Target Attention mechanism's query, and inputs a feature sequence having m object features into the mechanism's key and value, and repeats this weighting addition process k times to obtain k situational features corresponding to each of the k scene features. In other words, the situational features are obtained as features representing the characteristics of the shooting situation (for example, the distance between the person and the object at the time of shooting) for each of the k people included in the video captured by the camera or the like.

[0042] Furthermore, the feature integration unit 172B acquires relevance information, which is information indicating the degree of relevance of m objects to the actions of each of k people captured by a camera or the like.

[0043] Specifically, the feature integration unit 172B acquires, for example, the entropy value of the weights obtained as a result of processing the Softmax layer in the Source-Target Attention mechanism as relevance information. For example, if the aforementioned entropy value is relatively large, it can be estimated that the relevance of the object to the actions of a person captured by a camera or the like is relatively high. Conversely, for example, if the aforementioned entropy value is relatively small, it can be estimated that the relevance of the object to the actions of a person captured by a camera or the like is relatively low.

[0044] On the other hand, the feature integration unit 172B may be trained to give the largest weight to a dummy feature, which is set as a feature that mimics a desired object, for example, when there is no object in the image that corresponds to a person's action. When such training is performed, the feature integration unit 172B can obtain the weight value set for the dummy feature as relevance information when it inputs a feature sequence obtained by adding a dummy feature (vector sequence) to m object features into the key and value of the Source-Target Attention mechanism and performs weight addition. For example, if the weight value set for the dummy feature is relatively small, it can be estimated that the relevance of the object to the person's action captured by the camera, etc., is relatively low. Also, for example, if the weight value set for the dummy feature is relatively large, it can be estimated that the relevance of the object to the person's action captured by the camera, etc., is relatively high. In this embodiment, instead of training using dummy features, the feature integration unit 172B may perform training using training data that has labels indicating the relationship between a person's action and an object. Even when such learning is performed, the feature integration unit 172B can obtain correlation information similar to the magnitude of the weights set for the dummy features.

[0045] The feature integration unit 172C obtains a third processing result by performing a third processing on the second processing result output from the feature integration unit 172B, which integrates the features surrounding the person, and outputs the information including the obtained third processing result to the action estimation processing unit 18. The feature integration unit 172C also has a pre-trained neural network (NN) that integrates situational features and surrounding features to obtain integrated features, which are features that associate the surrounding features with the situational features. Furthermore, the feature integration unit 172C inputs the k situational features and relevance information obtained by the feature integration unit 172B, along with p surrounding features, into an NN having a Source-Target Attention mechanism and performs weighted addition to obtain k integrated features corresponding to each of the k situational features. Specifically, the feature integration unit 172C inputs one situation feature to be weighted into the Source-Target Attention mechanism's query, and inputs a feature sequence having p peripheral features into the key and value of the mechanism. This process is repeated k times to obtain k integrated features corresponding to each of the k situation features. The feature integration unit 172C also sets the weights in the aforementioned weighting addition based on relevance information. Specifically, for example, the feature integration unit 172C sets the weights in the aforementioned weighting addition by dividing the entropy value identified from the relevance information by the maximum entropy value. Alternatively, for example, the feature integration unit 172C sets the weights in the aforementioned weighting addition by subtracting the weight value for a dummy feature identified from the relevance information from "1". The feature integration unit 172C also obtains feature vectors corresponding to the k integrated features and outputs these obtained feature vectors to the action estimation processing unit 18.

[0046] Here, according to the processing of the feature integration unit 172C as described above, for example, if there is an object in the image that is highly relevant to the person's actions, the weight in the weight addition is set to a relatively small value. In such cases, the feature integration unit 172C obtains a feature vector that includes relatively small integrated features.

[0047] Furthermore, according to the processing of the feature integration unit 172C as described above, for example, if there are no objects in the image that are highly relevant to the person's actions, or if no objects highly relevant to the person's actions can be detected, the weights in the weighting addition are set to relatively large values. In these cases, the feature integration unit 172C obtains a feature vector that includes relatively large integrated features.

[0048] In other words, the processing of the feature integration unit 172C described above makes it possible to obtain a feature vector having integrated features that can be used to estimate the actions of each of the k people included in the video captured by a camera or the like.

[0049] The behavior estimation processing unit 18 performs behavior estimation processing to estimate the actions of people included in the video output from the video acquisition unit 11, based on the output information obtained by inputting the feature vectors obtained by the feature aggregation unit 17 into the estimation model 18A. The behavior estimation processing unit 18 also outputs the estimation results obtained by the aforementioned behavior estimation processing to an external device.

[0050] The estimation model 18A is configured as a model having a neural network, such as a CNN. Furthermore, the estimation model 18A is configured as a trained model that has been trained using machine learning with training data that associates feature vectors obtained from multiple time-series images with action labels that represent the actions of a person included in those images as one of a predetermined set of actions.Therefore, the estimation model 18A can obtain an action score as output information, which is a value indicating the probability for each class when the feature vectors obtained by the feature aggregation unit 17 are classified into one of a set of classes corresponding to each of a predetermined set of actions.In addition, when the estimation model 18A has the above configuration, the action estimation processing unit 18 can obtain as the estimation result of the actions of a person included in the video output from the video acquisition unit 11 an action corresponding to the class with the largest value among the multiple values ​​included in the aforementioned action score.

[0051] In this embodiment, the parameters of the neural network (NN) used in the processing of the person feature extraction unit 13, the object feature extraction unit 15, and / or the surrounding feature extraction unit 16 may be adjusted based on the parameters of the estimated model 18A obtained by the behavior estimation processing unit 18. Also in this embodiment, the parameters of the NN used in the processing of the feature aggregation unit 17 may be adjusted based on the parameters of the estimated model 18A obtained by the behavior estimation processing unit 18.

[0052] [Processing flow] Next, the processing flow performed in the behavior estimation device according to the first embodiment will be described. Figure 7 is a flowchart illustrating the processing performed in the behavior estimation device according to the first embodiment.

[0053] First, the behavior estimation device 100 acquires video footage including one or more people and one or more objects (step S11).

[0054] Next, the behavior estimation device 100 detects people and objects from a series of images included in the video acquired in step S11 (step S12).

[0055] Next, the behavior estimation device 100 extracts the characteristics of the person and object detected in step S12 (step S13). The behavior estimation device 100 also extracts the characteristics of the surrounding area of ​​the person detected in step S12 (step S13).

[0056] Next, the behavior estimation device 100 aggregates the features extracted in step S13 (step S14) using relevance information that indicates the degree of relevance of the object detected in step S12 to the person's behavior detected in step S12.

[0057] Finally, the behavior estimation device 100 estimates the behavior of the person detected in step S12 based on the features aggregated in step S14 (step S15).

[0058] As described above, according to this embodiment, feature quantities used to estimate the actions of a person in a video can be calculated using the characteristics of a person in the video, the characteristics of an object in the video, and the characteristics of the area surrounding the person in the video. Furthermore, as described above, according to this embodiment, the calculation results of the feature quantities used to estimate the actions of a person in a video can be varied according to the degree of relevance of the object to the person's actions. Therefore, according to this embodiment, the estimation accuracy when estimating a person's actions can be improved. Moreover, according to this embodiment, for example, even in cases where there are no objects in the image that are highly relevant to the person's actions, or where objects that are highly relevant to the person's actions cannot be detected, the estimation accuracy when estimating a person's actions can be kept to a minimum.

[0059] [Differentiation] The following describes modifications of the above embodiment. For simplicity, detailed explanations of the parts to which the previously described processes can be applied will be omitted as appropriate.

[0060] (Variation 1) According to this embodiment, the behavior estimation processing unit 18 may have a first estimation model, a second estimation model, and a third estimation model instead of the estimation model 18A. The first and second estimation models should be trained to estimate the behavior of a person included in multiple time-series images using a neural network (NN) with different weights than the estimation model 18A. The third estimation model should be trained to estimate the behavior of a person included in multiple time-series images using a neural network (NN) with similar weights to the estimation model 18A.

[0061] Furthermore, in the aforementioned case, the behavior estimation processing unit 18 may obtain a first estimation result by inputting a first feature vector including scene features to the first estimation model, obtain a second estimation result by inputting a second feature vector including situation features to the second estimation model, and obtain a third estimation result by inputting a third feature vector including integrated features to the third estimation model. Also, in the aforementioned case, the behavior estimation processing unit 18 may estimate the actions of a person included in the video output from the video acquisition unit 11 based on the first estimation result, the second estimation result, and the third estimation result. Specifically, the behavior estimation processing unit 18 may, when the multiple estimation results in the first estimation result, the second estimation result, and the third estimation result match, set the action corresponding to those multiple estimation results as the final estimation result. Furthermore, in the aforementioned cases, for example, the loss calculated based on scene features may be applied to the loss function in the first estimation model, the loss calculated based on situational features may be applied to the loss function in the second estimation model, and the loss calculated based on integrated features may be applied to the loss function in the third estimation model. As the aforementioned losses, for example, cross-entropy loss can be used.

[0062] (Modification 2) Figure 8 shows an example of the configuration of a feature aggregation unit included in a behavior estimation device according to a modification of the first embodiment. According to this embodiment, instead of the feature aggregation unit 17, a feature aggregation unit 27 as shown in Figure 8 may be provided in the behavior estimation device 100.

[0063] The feature aggregation unit 27 includes time-series feature integration units 271A, 271B, 271C, and 271D, a feature integration unit 272, and an aggregation processing unit 273.

[0064] The time-series feature integration unit 271A has the capability to perform the same processing as the time-series feature integration unit 171A, and outputs the person feature obtained by integrating multiple person features included in the person feature sequence to the aggregation processing unit 273.

[0065] The time-series feature integration unit 271B has the capability to perform the same processing as the time-series feature integration unit 171A, and outputs the person feature obtained by integrating multiple person features included in the person feature sequence to the feature integration unit 272.

[0066] The time-series feature integration unit 271C has the capability to perform the same processing as the time-series feature integration unit 171B, and outputs the object feature obtained by integrating multiple object feature quantities included in the object feature sequence to the feature integration unit 272.

[0067] The time-series feature integration unit 271D has the capability to perform the same processing as the time-series feature integration unit 171C, and outputs the peripheral feature obtained by integrating multiple peripheral features included in the peripheral feature sequence to the feature integration unit 272.

[0068] The feature integration unit 272 obtains integrated features by performing the same processing as feature integration units 172A, 172B, and 172C based on the person features obtained by the time-series feature integration unit 271B, the object features obtained by the time-series feature integration unit 271C, and the surrounding features obtained by the time-series feature integration unit 271D, and outputs the obtained integrated features to the aggregation processing unit 273. In other words, the feature integration unit 272 is configured to perform processing other than the processing related to obtaining feature vectors corresponding to integrated features among the processing performed by the aggregation processing unit 172. Furthermore, the feature integration unit 272 is configured to be able to obtain relevance information similar to that of the aggregation processing unit 172 and to perform processing using said relevance information.

[0069] The aggregation processing unit 273 obtains a feature vector based on the person features output from the time-series feature integration unit 271A and the integrated features output from the feature integration unit 272, and outputs the obtained feature vector to the action estimation processing unit 18. The aggregation processing unit 273 may, for example, perform a process to concatenate the person features and integrated features in the feature dimension direction, or a process to calculate the sum of the person features and integrated features, in order to obtain the feature vector.

[0070] Furthermore, according to this modified example, the connection between the three time-series feature integration units 271B to 271D and the aggregation processing unit 273 is not limited to a single feature integration unit 272; for example, y (y≧2) feature integration units 272 may be connected in series, as shown in Figure 9. Figure 9 shows an example configuration when multiple feature integration units are provided in the feature aggregation unit of Figure 8.

[0071] According to the configuration illustrated in Figure 9, the first of the y feature integration units 272 receives the person feature quantity obtained by the time-series feature integration unit 271B, the object feature quantity obtained by the time-series feature integration unit 271C, and the surrounding feature quantity obtained by the time-series feature integration unit 271D as input. Also according to the configuration illustrated in Figure 9, the z(2≦z≦y)th feature integration unit 272 receives the integrated feature quantity obtained by the (z-1)th feature integration unit 272, the object feature quantity obtained by the time-series feature integration unit 271C, and the surrounding feature quantity obtained by the time-series feature integration unit 271D as input. Furthermore, according to the configuration illustrated in Figure 9, the integrated feature quantity obtained by the yth feature integration unit 272 is output to the aggregation processing unit 273.

[0072] (Variation 3) Figure 10 shows an example of the configuration of a feature aggregation unit included in a behavior estimation device according to a modification of the first embodiment. According to this embodiment, instead of the feature aggregation unit 17, a feature aggregation unit 37 as shown in Figure 10 may be provided in the behavior estimation device 100.

[0073] The feature aggregation unit 37 includes a time-series feature integration unit 371, an integration processing unit 372, and an aggregation processing unit 373.

[0074] The time-series feature integration unit 371 has the capability to perform the same processing as the time-series feature integration unit 171A, and outputs the person feature obtained by integrating multiple person features included in the person feature sequence to the aggregation processing unit 373.

[0075] The integration processing unit 372 acquires integrated features based on the person feature sequence, the object feature sequence, and the surrounding feature sequence, and outputs the acquired integrated features to the aggregation processing unit 373. The integration processing unit 372 also includes a feature integration unit 372A and a time-series feature integration unit 372B.

[0076] The feature integration unit 372A obtains integrated features by performing the same processing as feature integration units 172A, 172B, and 172C based on the person feature sequence, object feature sequence, and surrounding feature sequence, and outputs the obtained integrated features to the time-series feature integration unit 372B. In other words, the feature integration unit 372A is configured to be able to obtain relevance information similar to that of the aggregation processing unit 172, and to perform processing using said relevance information.

[0077] The time-series feature integration unit 372B performs processing to integrate multiple integrated features obtained in a time series from the feature integration unit 372A into one. Specifically, the time-series feature integration unit 372B integrates multiple integrated features into a single person feature by performing weighted addition using, for example, multiple integrated features and a neural network (NN) having a Source-Target Attention mechanism. The time-series feature integration unit 372B also outputs the integrated features, which have been integrated for each person identifiable based on the person feature information or the person feature sequence, to the aggregation processing unit 373. The time-series feature integration unit 372B may also feed back the integrated features so that they are incorporated into the processing of the feature integration unit 372A. Specifically, the time-series feature integration unit 372B may feed back the integrated features so that they are included in the person feature sequence, or so that they are used in place of scene features, or so that they are used in place of situation features. Furthermore, the processing of the time-series feature integration unit 372B may be performed using, for example, a neural network (NN) having an LSTM (Long Short-Term Memory).

[0078] The aggregation processing unit 373 obtains a feature vector based on the person features output from the time-series feature integration unit 371 and the integrated features output from the integration processing unit 372, and outputs the obtained feature vector to the action estimation processing unit 18. The aggregation processing unit 373 may, for example, perform a process to concatenate the person features and integrated features in the feature dimension direction, or a process to calculate the sum of the person features and integrated features, in order to obtain the feature vector.

[0079] <Second Embodiment> Figure 11 is a block diagram showing the functional configuration of the behavior estimation device according to the second embodiment.

[0080] The behavior estimation device 500 according to this embodiment has the same hardware configuration as the behavior estimation device 100. The behavior estimation device 500 also includes a person feature extraction means 511, an object feature extraction means 512, a surrounding feature extraction means 513, a feature aggregation means 514, and a behavior estimation processing means 515.

[0081] Figure 12 is a flowchart illustrating the processes performed in the behavior estimation device according to the second embodiment.

[0082] The person feature extraction means 511 extracts the features of a person detected from multiple images in a time series (step S51).

[0083] The object feature extraction means 512 extracts the features of the detected object from a series of images (step S52).

[0084] The peripheral feature extraction means 513 extracts peripheral features of a person in multiple time-series images (step S53).

[0085] The feature aggregation means 514 performs aggregation processing to aggregate the features of a person, the features of an object, and the features of the person's surroundings (step S54).

[0086] The behavior estimation processing means 515 performs processing to estimate the behavior of a person included in a series of images based on the information including the processing results of the aggregation process (step S55).

[0087] According to this embodiment, the estimation accuracy when estimating a person's actions can be improved.

[0088] Some or all of the above embodiments may also be described as follows, but are not limited to the following:

[0089] (Note 1) A person feature extraction means for extracting person features detected from multiple images in a time series, Object feature extraction means for extracting features of objects detected from the aforementioned plurality of images, A peripheral feature extraction means for extracting peripheral features of the person in the aforementioned multiple images, A feature aggregation means for performing aggregation processing to aggregate the characteristics of the person, the characteristics of the object, and the characteristics of the person's surroundings. Action estimation processing means for performing processing to estimate the actions of the person included in the plurality of images based on information including the processing results of the aggregation process, A behavior estimation device having the following features.

[0090] (Note 2) The feature aggregation means is an action estimation device as described in Appendix 1, which performs the aggregation process by obtaining a first processing result by performing a first processing that integrates the features of the person with the features of the person's surroundings, obtaining a second processing result by performing a second processing that integrates the features of an object with the first processing result, and obtaining a third processing result by performing a third processing that integrates the features of the person's surroundings with the second processing result.

[0091] (Note 3) The feature aggregation means acquires relevance information, which is information indicating the degree of relevance of the object to the person's actions, based on predetermined parameters used in the second process, and performs the third process using the acquired relevance information, as described in Appendix 2.

[0092] (Note 4) The behavior estimation processing means is a behavior estimation device as described in Appendix 2, which estimates the behavior of the person included in the plurality of images based on a first estimation result obtained by estimating the person's behavior from information including the first processing result, a second estimation result obtained by estimating the person's behavior from information including the second processing result, and a third estimation result obtained by estimating the person's behavior from information including the third processing result, instead of information including the processing result of the aggregation process.

[0093] (Note 5) A person feature extraction means for extracting person features detected from multiple images in a time series, Object feature extraction means for extracting features of objects detected from the aforementioned plurality of images, A peripheral feature extraction means for extracting peripheral features of the person in the aforementioned multiple images, A feature integration means for performing integration processing to integrate the characteristics of the person, the characteristics of the object, and the characteristics of the person's surroundings, Aggregation processing means for performing aggregation processing to aggregate the characteristics of the person and the processing results of the integration process, A behavior estimation processing means for estimating the actions of people included in the plurality of images based on information including the processing results of the aggregation process, A behavior estimation device having the following features.

[0094] (Note 6) Extract the features of a person detected from multiple images in a time series, Features of the object detected from the aforementioned multiple images are extracted, Extract the surrounding features of the person in the aforementioned multiple images, An aggregation process is performed to combine the characteristics of the person, the characteristics of the object, and the characteristics of the person's surroundings. A method for estimating the actions of a person included in a plurality of images, based on information including the processing results of the aggregation process.

[0095] (Note 7) Extract the features of a person detected from multiple images in a time series, Features of the object detected from the aforementioned multiple images are extracted, Extract the surrounding features of the person in the aforementioned multiple images, An integration process is performed to combine the characteristics of the person, the characteristics of the object, and the characteristics of the person's surroundings. An aggregation process is performed to combine the characteristics of the aforementioned person and the processing results of the aforementioned integration process. A method for estimating the actions of people included in a plurality of images, based on information including the processing results of the aggregation process.

[0096] (Note 8) Extract the features of a person detected from multiple images in a time series, Features of the object detected from the aforementioned multiple images are extracted, Extract the surrounding features of the person in the aforementioned multiple images, An aggregation process is performed to combine the characteristics of the person, the characteristics of the object, and the characteristics of the person's surroundings. A recording medium that stores a program that causes a computer to perform a process to estimate the actions of the people included in the plurality of images, based on information including the processing results of the aggregation process.

[0097] (Note 9) Extract the features of a person detected from multiple images in a time series, Features of the object detected from the aforementioned multiple images are extracted, Extract the surrounding features of the person in the aforementioned multiple images, An integration process is performed to combine the characteristics of the person, the characteristics of the object, and the characteristics of the person's surroundings. An aggregation process is performed to combine the characteristics of the aforementioned person and the processing results of the aforementioned integration process. A recording medium that stores a program that causes a computer to perform a process to estimate the actions of people included in the multiple images based on information including the processing results of the aggregation process.

[0098] Although the present disclosure has been described above with reference to embodiments and examples, the present disclosure is not limited to the above embodiments and examples. Various modifications to the structure and details of the present disclosure can be understood by those skilled in the art within the scope of the present disclosure. [Explanation of Symbols]

[0099] 13. Person Feature Extraction Unit 15. Object Feature Extraction Unit 16. Peripheral Feature Extraction Unit 17 Feature aggregation section 18 Action Estimation Processing Unit 100 Behavior estimation device

Claims

1. A person feature extraction means for extracting person features detected from multiple images in a time series, Object feature extraction means for extracting features of objects detected from the aforementioned plurality of images, A peripheral feature extraction means for extracting peripheral features of the person in the aforementioned multiple images, A feature aggregation means for performing aggregation processing to aggregate the characteristics of the person, the characteristics of the object, and the characteristics of the person's surroundings. Action estimation processing means for performing processing to estimate the actions of the person included in the plurality of images based on information including the processing results of the aggregation process, It has, The feature aggregation means is a behavior estimation device that performs the aggregation process by obtaining a first processing result by performing a first processing that integrates the features of the person with the features of the person's surroundings, obtaining a second processing result by performing a second processing that integrates the features of the object with the first processing result, and obtaining a third processing result by performing a third processing that integrates the features of the person's surroundings with the second processing result.

2. The behavior estimation device according to Claim 1, wherein the feature aggregation means acquires relevance information, which is information indicating the degree of relevance of the object to the person's behavior, based on predetermined parameters used in the second process, and performs the third process using the acquired relevance information.

3. The behavior estimation device according to claim 1, wherein the behavior estimation processing means estimates the behavior of the person included in the plurality of images based on a first estimation result obtained by estimating the behavior of the person from information including the first processing result, a second estimation result obtained by estimating the behavior of the person from information including the second processing result, and a third estimation result obtained by estimating the behavior of the person from information including the third processing result, instead of information including the processing result of the aggregation process.

4. A person feature extraction means for extracting person features detected from multiple time-series images, Object feature extraction means for extracting features of objects detected from the aforementioned plurality of images, A peripheral feature extraction means for extracting peripheral features of the person in the aforementioned multiple images, A feature integration means for performing integration processing to integrate the characteristics of the person, the characteristics of the object, and the characteristics of the person's surroundings, Aggregation processing means for performing aggregation processing to aggregate the characteristics of the person and the processing results of the integration process, A behavior estimation processing means for estimating the actions of people included in the plurality of images based on information including the processing results of the aggregation process, A behavior estimation device having the following features.

5. Extract the characteristics of a person detected from multiple images in a time series, Extract the features of the object detected from the aforementioned multiple images, Extract the surrounding features of the person in the aforementioned multiple images, An aggregation process is performed to combine the characteristics of the person, the characteristics of the object, and the characteristics of the person's surroundings. A behavior estimation method for estimating the actions of a person included in a plurality of images based on information including the processing results of the aggregation process, An action estimation method that performs the following processing as the aggregation process: first processing to obtain a first processing result by integrating the characteristics of the person with the characteristics of the person's surroundings; second processing to obtain a second processing result by integrating the characteristics of the object with the first processing result; and third processing to obtain a third processing result by integrating the characteristics of the person's surroundings with the second processing result.

6. Extract the characteristics of a person detected from multiple images in a time series, Extract the features of the object detected from the aforementioned multiple images, Extract the surrounding features of the person in the aforementioned multiple images, An integration process is performed to combine the characteristics of the person, the characteristics of the object, and the characteristics of the person's surroundings. An aggregation process is performed to combine the characteristics of the aforementioned person and the processing results of the aforementioned integration process. A behavior estimation method for estimating the actions of people included in a plurality of images based on information including the processing results of the aggregation process.

7. Extract the characteristics of a person detected from multiple images in a time series, Extract the features of the object detected from the aforementioned multiple images, Extract the surrounding features of the person in the aforementioned multiple images, An aggregation process is performed to combine the characteristics of the person, the characteristics of the object, and the characteristics of the person's surroundings. Based on the information including the processing results of the aggregation process, the computer is made to perform a process to estimate the actions of the person included in the plurality of images. A recording medium that records a program causing the computer to perform the following processes as the aggregation process: first, obtaining a first processing result by performing a first process to integrate the characteristics of the person with the characteristics of the person's surroundings; second, obtaining a second processing result by performing a second process to integrate the characteristics of the object with the first processing result; and third, obtaining a third processing result by performing a third process to integrate the characteristics of the person's surroundings with the second processing result.

8. Extract the characteristics of a person detected from multiple images in a time series, Extract the features of the object detected from the aforementioned multiple images, Extract the surrounding features of the person in the aforementioned multiple images, An integration process is performed to combine the characteristics of the person, the characteristics of the object, and the characteristics of the person's surroundings. An aggregation process is performed to combine the characteristics of the aforementioned person and the processing results of the aforementioned integration process. A recording medium that stores a program that causes a computer to perform a process to estimate the actions of people included in the multiple images, based on information including the processing results of the aggregation process.