Behavior recognition device, method, and program

By incorporating environmental information with video features, the technology enhances behavior recognition accuracy by leveraging convolutional and recurrent neural networks to accurately estimate actions based on both performer movements and their surroundings.

JP7768246B2Active Publication Date: 2025-11-12NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023564376
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-02
Publication Date
2025-11-12
Estimated Expiration
2041-12-02

AI Technical Summary

Technical Problem

Existing behavior recognition technologies are prone to misrecognition as they rely solely on information about the movements and postures of the performer, leading to inaccurate action estimation.

Method used

The technology integrates environmental information with video feature information to estimate candidate actions by detecting and fusing actor and environmental information, using a combination of convolutional neural networks and recurrent neural networks to enhance accuracy.

Benefits of technology

This approach effectively narrows down candidates for the subject's actions, enabling precise behavior estimation by considering both the performer's movements and their surrounding environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007768246000001
    Figure 0007768246000001
  • Figure 0007768246000002
    Figure 0007768246000002
  • Figure 0007768246000003
    Figure 0007768246000003
Patent Text Reader

Abstract

An aspect of this invention includes: acquiring video data, in which a range including at least a surrounding region of an acting entity is image-captured; detecting environment information of the surrounding region, regarding which a degree of relation with an action of the acting entity satisfies a predetermined condition, from the video data that is acquired; estimating candidates of an action of the acting entity on the basis of the environment information that is detected and video feature information extracted from the video data; and deciding an action of the acting entity upon from the candidates of action that are estimated.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] One aspect of the present invention relates to an apparatus, method, and program for recognizing behavior used to estimate, for example, a person's behavior. [Background technology]

[0002] For example, video-based behavior recognition technology is expected to be applied in a variety of situations. For example, in manufacturing sites like factories or work sites like logistics warehouses, recognizing the work processes of workers can visualize the overall work situation and determine whether the work being performed by a worker is physically demanding or dangerous, thereby preventing industrial accidents. Furthermore, by equipping surveillance cameras with human behavior recognition capabilities, it is possible to go beyond passive surveillance that simply captures the entire surveillance area and enable active surveillance, such as focusing on suspicious individuals. Other potential applications include recording the in-store behavior of customers visiting convenience stores and supermarkets, which cannot be obtained using POS (Point of Sale) data. The list goes on and on.

[0003] Incidentally, examples of behavior recognition technologies that have been proposed include a technology that recognizes behavior using the skeleton and movements of a person in a third-person video, as described in Non-Patent Document 1, and Non-patent document 2 As described in [2], a technique for simultaneously estimating the behavior and gaze direction of a person in a first-person video is known. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Lei Shi, Yifan Zhang, Jian Cheng, and Hanqing Lu. “Skeleton-Based Action Recognition With Multi-Stream Adaptive Graph Convolutional Networks.” in IEEE Transactions on Image Processing, vol. 29, pp. 9532-9545, 2020.

[0005] [Non-patent document 2] Yifei Huang, Minjie Cai, Zhenqiang Li, Feng Lu, and Yoichi Sato. “Mutual Context Network for Jointly Estimating Egocentric Gaze and Action.” in IEEE Transactions on Image Processing, vol. 29, pp. 7795-7806, 2020. Summary of the Invention [Problem to be solved by the invention]

[0006] However, the techniques described in Non-Patent Documents 1 and 2 have the problem that they are prone to misrecognition because they recognize actions using only information about the movements and postures of the performer.

[0007] The present invention has been made in light of the above circumstances, and aims to provide a technology that effectively narrows down candidates for the subject's actions, thereby enabling the subject's actions to be estimated with high accuracy. [Means for solving the problem]

[0008] In order to solve the above problems, a first aspect of the behavior recognition device or method according to the present invention includes a processing unit or a processing step for acquiring video data capturing an area including an actor and its surrounding area, a processing unit or a processing step for detecting actor information representing the movement or posture of the actor from the acquired video data, and a processing unit or a processing step for detecting actor information representing the movement or posture of the actor from the acquired video data. The degree of involvement with the action of the performer in the surrounding area satisfies a predetermined condition. a processing unit or a predetermined process for detecting environmental information; The candidate actions of the actor are estimated based on fusion information of the environmental information and the video feature information extracted from the video data, and the actor information is further taken into account in the estimated candidate actions. The subject of the action The aforementioned Take action decision The present invention is configured to include a processing unit or process for:

[0009] A second aspect of the behavior recognition device or method of the present invention comprises a processing unit or processing step that acquires video data capturing an area that includes at least the surrounding area of ​​the actor; a processing unit or processing step that detects, from the acquired video data, environmental information of the surrounding area whose degree of relationship with the actor's behavior satisfies predetermined conditions; and a processing unit or processing step that estimates candidate behaviors of the actor based on the detected environmental information and video feature information extracted from the video data, and determines the actor's behavior from the estimated candidate behaviors. [Effects of the Invention]

[0010] That is, according to the first and second aspects of the present invention, it is possible to provide a technology that effectively narrows down candidates for the behavior of the performer and enables the behavior of the performer to be estimated with high accuracy. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a block diagram showing an example of the hardware configuration of a behavior recognition device according to an embodiment of the present invention, together with the configuration of its peripheral parts. [Figure 2] FIG. 2 is a block diagram showing an example of the software configuration of the behavior recognition device according to an embodiment of the present invention. [Figure 3]FIG. 3 is a flowchart showing an example of the processing procedure and processing content of the video feature information extraction processing executed by the video feature information extraction processing unit included in the behavior recognition device shown in FIG. [Figure 4] FIG. 4 is a flowchart showing an example of the processing procedure and processing content of the environment segment extraction processing executed by the environment segment extraction processing unit included in the behavior recognition device shown in FIG. [Figure 5] FIG. 5 is a flowchart showing an example of the processing procedure and processing content of the behavior area detection processing executed by the behavior area detection processing unit included in the behavior recognition device shown in FIG. [Figure 6] FIG. 6 is a flowchart showing an example of the processing procedure and processing content of the environmental information generation processing executed by the environmental information generation processing unit included in the behavior recognition device shown in FIG. [Figure 7] FIG. 7 is a flowchart showing an example of the processing procedure and processing content of the information fusion processing executed by the information fusion processing unit included in the behavior recognition device shown in FIG. [Figure 8] FIG. 8 is a flowchart showing an example of the processing procedure and processing content of the behavior estimation processing executed in the learning phase by the behavior estimation processing unit included in the behavior recognition device shown in FIG. [Figure 9] FIG. 9 is a flowchart showing an example of the processing procedure and processing content of the behavior estimation processing executed in the inference phase by the behavior estimation processing unit included in the behavior recognition device shown in FIG. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0013] [One embodiment] (Configuration example) FIG. 1 is a block diagram showing an example of the hardware configuration of a behavior recognition device according to an embodiment of the present invention together with the configuration of its peripheral parts, and FIG. 2 is a block diagram showing an example of the software configuration of the behavior recognition device.

[0014] The behavior recognition device AS is composed of an information processing device such as a server computer or a personal computer, and a camera CM and a terminal MT are connected to the behavior recognition device AS via a signal cable or a network.

[0015] The camera CM includes a first camera mounted on a ceiling or wall capable of capturing an image of the area to be monitored, and a second camera attached to the head of a person when the subject of the action to be recognized is a person. The first camera captures an area including the entire subject of the action to be recognized in the area to be monitored from a third-person perspective and transmits the time-series video data. The second camera captures an area including the entire subject of the action to be recognized from a first-person perspective and transmits the time-series video data.

[0016] The camera CM may be only one of the first and second cameras. The time-series video data may be transmitted live from the camera CM, or may be stored in a video database and then read out.

[0017] The terminal MT is used by, for example, a system administrator or an administrator who manages the behavior of actors, and is composed of an information processing terminal such as a personal computer. When the behavior recognition device AS is set to the learning phase, the terminal MT is used to input video data captured within an area including a candidate actor to be recognized and its surrounding area, and information representing the distribution of correct behavior labels corresponding to the video data, to the behavior recognition device AS. The terminal MT is also used to receive the behavior labels of the person to be recognized estimated by the behavior recognition device AS during inference.

[0018] The input process of the information representing the correct behavior label distribution and the reception process of the behavior labels estimated by the behavior recognition device AS may be distributed among a plurality of terminals MT.

[0019] The behavior recognition device AS has a control unit 1 that uses hardware processors such as a central processing unit (CPU) and a graphics processing unit (GPU), and is connected to the control unit 1 via a bus 5 with a storage unit having a program storage unit 2 and a data storage unit 3, and an input / output interface (hereinafter, interface will be abbreviated as I / F) unit 4.

[0020] The input / output I / F section 4 has a communication interface function, and transmits and receives video data and various input / output data to and from the camera CM and the terminal MT via a signal cable or a network.

[0021] The program storage unit 2 is configured by combining, for example, a non-volatile memory such as a solid-state drive (SSD) as a storage medium that can be written to and read from at any time, and a non-volatile memory such as a read-only memory (ROM), and stores middleware such as an operating system (OS), as well as application programs required to execute various control processes according to one embodiment. Hereinafter, the OS and each application program will be collectively referred to as the program.

[0022] The data storage unit 3 is, for example, a combination of a non-volatile memory such as an SSD that can be written to and read from at any time as a storage medium, and a volatile memory such as a RAM (Random Access Memory), and is equipped with a video data storage unit 31 and a parameter storage unit 32 as the main storage units required to implement one embodiment.

[0023] The video data storage unit 31 is used to temporarily store the time-series video data transmitted from the camera CM for behavior recognition processing.

[0024] The parameter storage unit 32 saves parameters to be set in a learning model used in each processing unit, which will be described later, included in the control unit 1. The parameter storage unit 32 also stores pair data of nouns and verbs acquired from text data.

[0025] The control unit 1 includes, as processing functions necessary to implement one embodiment, a video data acquisition processing unit 11, a video feature information extraction processing unit 12, an environment segment extraction processing unit 13, an environment information generation processing unit 14, a behavior area detection processing unit 15, an information fusion processing unit 16, a behavior estimation processing unit 17, and a parameter update processing unit 18. Of these, the environment segment extraction processing unit 13, the environment information generation processing unit 14, and the behavior area detection processing unit 15 function as environment information detection processing units.

[0026] Furthermore, among the above-mentioned processing units 11 to 18, the video feature information extraction processing unit 12, the environment segment extraction processing unit 13, the behavior area detection processing unit 15, and the behavior estimation processing unit 17 all perform their respective processing using a learning model configured, for example, by a convolutional neural network (CNN).

[0027] The processing units 11 to 18 are all realized by causing a hardware processor in the control unit 1 to execute application programs stored in the program storage unit 2.

[0028] The video data acquisition processing unit 11 acquires time-series third-person perspective and first-person perspective video data output from the terminal MT in the learning phase and from the camera CM in the inference phase via the input / output I / F unit 4, and performs processing to temporarily store each of the acquired video data in the video data storage unit 31.

[0029] The image feature information extraction processing unit 12 extracts image features from each frame image of the image data stored in the image data storage unit 31, and outputs the extracted image feature information to the information fusion processing unit 16.

[0030] The environment segment extraction processing unit 13 extracts the area of ​​the object (environment segment) depicted in each frame image of the video data stored in the video data storage unit 31 and the name of the object (environment segment label), and outputs the extracted pair of the environment segment and the environment segment label to the environment information generation processing unit 14.

[0031] The environment information generation processing unit 14 receives pairs of the environment segments and environment segment labels from the environment segment extraction processing unit 13 for each frame image, and performs processing to generate environment information corresponding to the environment segment label for each environment segment. Note that if the subject of the action is a person, the environment information is information that represents actions that the person can perform. An example of the environment information will be described in the operation example.

[0032] The action area detection processing unit 15 detects, for each frame image of the video data stored in the video data storage unit 31, an area that the actor focuses on when acting from that image area as an "action area." The action area detection processing unit 15 also detects areas other than the action area as "peripheral areas." Then, it performs processing to output information representing the detected "action area" and "peripheral areas" to the information fusion processing unit 16 as an action area mask.

[0033] For each frame image, the information fusion processing unit 16 receives video feature information, environmental information, and behavior area masks from the video feature information extraction processing unit 12, environmental information generation processing unit 14, and behavior area detection processing unit 15, respectively. Then, for each environment segment, the information fusion processing unit 16 classifies the environmental information into "attention environment information" and "surrounding environment information" based on the behavior area mask, and performs processing to fuse the classified "attention environment information" and "surrounding environment information" with the video feature information to generate fusion information corresponding to the frame image. The generated fusion information is passed from the information fusion processing unit 16 to the behavior estimation processing unit 17.

[0034] The action estimation processing unit 17 receives the fusion information from the information fusion processing unit 16 and estimates the action label distribution of the agent based on this fusion information. Then, in the learning phase, the estimated action label distribution is output to the parameter update processing unit 18. On the other hand, in the inference phase, the action label of the agent is determined from the estimated action label distribution, and the determined action label is output from the input / output I / F unit 4 to the terminal MT.

[0035] In the learning phase, the parameter update processing unit 18 acquires a correct behavior label distribution corresponding to video data from the terminal MT, and also acquires the estimated behavior label distribution from the behavior estimation processing unit 17. Then, the parameter update processing unit 18 calculates the difference between the acquired correct behavior label distribution and the estimated behavior label distribution, and performs processing to update each parameter stored in the parameter storage unit 32 in the direction that reduces this difference.

[0036] (Example of operation) Next, an example of the operation of the behavior recognition device AS configured as above will be described. Note that the following operation will be described taking as an example the case where the subject of the action to be recognized is a person.

[0037] (1) Learning Phase (1-1) Acquisition of video data and extraction of video feature information FIG. 3 is a flowchart showing an example of the video feature information extraction process executed by the control unit 1 of the behavior recognition apparatus AS.

[0038] In the learning phase, the control unit 1 of the behavior recognition device AS, under the control of the video data acquisition processing unit 11, acquires first-person perspective video data and third-person perspective video data showing a person to be recognized and their surrounding circumstances, which are stored in advance in the terminal MT or an associated database, via the input / output I / F unit 4. The acquired video data is then temporarily stored in the video data storage unit 31. At the same time, under the control of the parameter update processing unit 18, the control unit 1 of the behavior recognition device AS acquires a correct behavior label distribution corresponding to the video data from the terminal MT or the database.

[0039] When the video data is acquired, the control unit 1 of the behavior recognition device AS reads the video data frame by frame from the video data storage unit 31 in step S11 under the control of the video feature information extraction processing unit 12. Then, in steps S12 and S13, the control unit 1 repeatedly executes a process of extracting features from each of the read frame images.

[0040] For example, video feature information extraction processing unit 12 inputs each of the frame images into a learning model configured by a convolutional neural network, and obtains a one-dimensional vector indicating the feature amount for each frame image from this learning model. Then, in step S14, video feature information extraction processing unit 12 outputs the one-dimensional vector as video feature information to information fusion processing unit 16. Note that the video feature information may be any information that expresses the vector of the video.

[0041] (1-2) Extraction of environmental segments and generation of environmental information 4 and 6 are flowcharts showing an example of the processing procedure and processing contents of the environment segment extraction processing and the environment information generation processing, respectively, executed by the control unit 1 of the behavior recognition device AS.

[0042] The control unit 1 of the behavior recognition device AS first reads video data frame by frame from the video data storage unit 31 in step S21 under the control of the environment segment extraction processing unit 13. Then, in steps S22 and S23, it repeatedly executes a process of extracting from each of the read frame images the area of ​​an object (environment segment) depicted in the frame image and the name of the object (environment segment label). Then, in step S24, the environment segment extraction processing unit 13 outputs the extracted pair of the environment segment and the environment segment label to the environment information generation processing unit 14.

[0043] An environment segment can be any information that indicates a partial area in each frame image. For example, if the coordinates of the top left corner of a partial area are (xo, yo), its width is we, and its height is he, it is expressed as (xo, yo, wo, ho). An environment segment label is the name of an object that exists in the partial area.

[0044] The process of extracting the environmental segments and environmental segment labels is performed using, for example, Faster-RCNN, a type of convolutional neural network (CNN). In this example, Faster-RCNN takes a frame image as input, determines whether a predefined object exists in the frame image, and if so, outputs information indicating the area coordinates and the object name.

[0045] (1-3) Generation of environmental information The control unit 1 of the behavior recognition device AS then executes a process of generating environmental information as follows under the control of the environmental information generation processing unit 14.

[0046] FIG. 6 is a flowchart showing the procedure and content of the environmental information generation process executed by the control unit 1 of the behavior recognition device AS.

[0047] In step S41, the environment information generation processing unit 14 first receives environment segments and environment segment labels from the environment segment extraction processing unit 13. Then, in steps S42 to S44, the environment information generation processing unit 14 repeatedly executes a process of generating, for each frame image, environment information corresponding to the environment segment label of the object present in the environment segment for each environment segment extracted from the frame image. Then, in step S45, the environment information generation processing unit 14 outputs, for each frame image, a pair of the generated environment information and the corresponding environment segment to the information fusion processing unit 16.

[0048] The environmental information may be any information that expresses actions that a person can perform. For example, verb-noun pairs may be extracted in advance from a large amount of text data, and a verb frequency distribution expressed as a one-dimensional vector that expresses how many times each verb appears for a noun corresponding to an environmental segment label may be used as the environmental information. Alternatively, a verb probability distribution may be generated by dividing each element of the verb frequency distribution by the total frequency, and this verb probability distribution may be used as the environmental information. Alternatively, a weighted verb frequency distribution or a weighted verb probability distribution may be obtained by multiplying the verb frequency distribution or each element of the verb probability distribution by the inverse document frequency value of the corresponding verb, and this may be used as the environmental information.

[0049] (1-4) Detection of activity area Further, the control unit 1 of the behavior recognition device AS, under the control of the behavior area detection processing unit 15, executes the process of detecting the behavior area and the surrounding area from the frame image as follows.

[0050] FIG. 5 is a flowchart showing an example of the processing procedure and processing contents of the behavior area detection processing executed by the control unit 1 of the behavior recognition device AS.

[0051] That is, the action area detection processing unit 15 first reads video data frame by frame from the video data storage unit 31 in step S31. Next, in steps S32 to S34, areas in each frame image where the degree of involvement with a person's action is greater than a preset threshold are detected as action areas, and other areas are designated as peripheral areas. Then, a two-dimensional vector of the same size as the frame image is generated, with the detected action area designated as "1" and the peripheral area designated as "0". The above process is repeatedly performed for each frame image.

[0052] Finally, in step S35, the behavior area detection processing unit 15 outputs a two-dimensional vector represented by the generated behavior area and the surrounding area to the information fusion processing unit 16 as a behavior area mask.

[0053] The method for detecting the activity area may be any method that can divide each pixel constituting the frame image into an activity area and a non-activity area. For example, a detection method based on the area of ​​a person's hands can be considered.

[0054] This detection method, for example, first uses hand segmentation technology to extract the area of ​​a person's hand to obtain the pixels where the hand is drawn. Next, if the coordinates of the pixel where the target person's hand is drawn are (xh, yh) and the action area threshold is θh, the pixels included in the range (xh -θh, yh -θh, 2 × θh +1, 2 × θh +1) based on that pixel are detected as the action area, and the other pixels as the surrounding area. Note that each element of the range information indicates the x and y coordinates of the top left corner of the range, as well as the width and height. This detection method is described in detail in, for example, the following literature.

[0055] Takehiko Ohkawa, Takuma Yagi, Atsushi Hashimoto, Yoshitaka Ushiku, and Yoichi Sato. Foreground-Aware Stylization and Consensus Pseudo-Labeling for Domain Adaptation of First-Person Hand Segmentation. in IEEE Access, vol. 9, pp. 94644-94655, 2021.

[0056] Another possible detection method is a segmentation method based on the gaze direction of a person. This method first uses gaze prediction technology to extract the gaze area of ​​a person and obtains the pixels on which the target person's gaze is focused. Next, if the coordinates of the pixel representing the gaze of the person to be recognized are (xe, ye) and the behavior area threshold is θe, the pixels included in the range (xe - θe, ye - θe, 2 × θe + 1, 2 × θe + 1) based on that pixel are detected as the behavior area, and the other pixels as the peripheral area. Note that each element of the range information indicates the x and y coordinates of the top left corner of the range, as well as the width and height. This detection method is described in detail in, for example, the following literature.

[0057] Yifei Huang, Minjie Cai, Zhenqiang Li, and Yoichi Sato. Predicting Gaze in Egocentric Video by Learning Task-dependent Attention Transition. in Proceedings of the European Conference on Computer Vision (ECCV), pp. 754-769, 2018. The above-described behavioral area detection process can also be realized using a learning model configured by a convolutional neural network.

[0058] (1-5) Information fusion Next, the control unit 1 of the behavior recognition device AS, under the control of the information fusion processing unit 16, executes a process of fusing the environmental information of interest and the surrounding environmental information with the video feature information to generate fusion information as follows.

[0059] FIG. 7 is a flowchart showing an example of the processing procedure and processing contents of the information fusion processing executed by the control unit 1 of the behavior recognition device AS.

[0060] That is, first, in step S51, the information fusion processing unit 16 receives, for each frame image, video feature information, environmental information, and an action area mask from the video feature information extraction processing unit 12, the environmental information generation processing unit 14, and the action area detection processing unit 15. Next, in a processing loop of steps S52 to S56, the information fusion processing unit 16 repeatedly executes the process of generating fusion information for each frame image by fusing the video feature information, environmental information of the region of interest, and environmental information of the surrounding regions.

[0061] That is, in the processing loop of steps S52 to S56, the information fusion processing unit 16 first repeatedly determines, for each environment segment, whether or not the environment segment overlaps with the behavior area mask, i.e., whether the environment segment is a "region of interest" or a "peripheral area," in steps S53 to S54. Next, in step S55, based on the determination result, the information fusion processing unit 16 fuses environmental information in the "region of interest" to generate attention environment information, and fuses environmental information in the "peripheral area" to generate peripheral environment information.

[0062] The attention environment information and the surrounding environment information may be any information that can represent multiple pieces of environmental information belonging to the "action area" and the "surrounding area" as a single piece of environmental information. For example, attention environment information represented by a single vector is generated by adding up each element of the environmental information belonging to the "action area." Similarly, surrounding environment information represented by a single vector is generated by adding up each element of the environmental information belonging to the "surrounding area." Alternatively, for example, weighted environmental information may be generated by multiplying each element of the environmental information by a confidence level expressed as a decimal point in the range of 0 to 1 obtained by the environment segment extraction processing unit 13, and the attention environment information and the surrounding environment information are generated by performing the addition process in the same manner as above.

[0063] Then, in step S56, the information fusion processing unit 16 connects the generated focused environment information and surrounding environment information with the corresponding video feature information for each frame image to generate fusion information represented by a new one-dimensional vector corresponding to the frame image. Finally, in step S57, the information fusion processing unit 16 outputs the fusion information generated for each frame image to the behavior estimation processing unit 17.

[0064] (1-6) Estimation of the subject's actions Next, the control unit 1 of the behavior recognition device AS, under the control of the behavior estimation processing unit 17, estimates the behavior label of the person who is the subject of the action based on the fusion information as follows.

[0065] FIG. 8 is a flowchart showing an example of the processing procedure and processing content of the behavior estimation processing executed by the control unit 1 of the behavior recognition device AS in the learning phase.

[0066] That is, first, in step S61, the behavior estimation processing unit 17 receives fusion information corresponding to each frame image from the information fusion processing unit 16. Next, in step S62, based on all the received fusion information, behavior expression information is generated that expresses the behavior of the person who is the subject of the action as a one-dimensional vector.

[0067] This process may be any process that converts a series of fusion information into behavioral expressions. For example, this can be realized by inputting the fusion information into a recurrent neural network (RNN) such as a Long Short Term Memory (LSTM) or a Gated Recurrent Unit (GRU), and using the information output from the neural network as behavioral expression information.

[0068] Next, in step S63, the behavior estimation processing unit 17 estimates an behavior label distribution based on the generated behavior expression information. This behavior label distribution estimation process may be any process that converts behavior expressions into behavior label distributions. For example, in the case of single-label classification, the behavior expressions are converted into a probability distribution (behavior label distribution) of the number of dimensions of predefined behavior recognition labels using a linear layer and a softmax function. In addition, in the case of multi-label classification, the behavior expressions are converted into a distribution (behavior label distribution) of the number of dimensions of predefined behavior labels using a linear layer and a sigmoid function, or a linear layer and a tanh function.

[0069] The behavior estimation processing unit 17 outputs information representing the estimated behavior label distribution to the parameter update processing unit 18 in step S64.

[0070] (1-7) Parameter update The control unit 1 of the behavior recognition device AS, under the control of the parameter update processing unit 18, executes processing to update each parameter stored in the parameter storage unit 32 as follows.

[0071] The parameters to be updated are parameters set in each learning model used by the video feature information extraction processing unit 12, the environment segment extraction processing unit 13, the behavior area detection processing unit 15, and the behavior estimation processing unit 17.

[0072] That is, the parameter update processing unit 18 first receives the correct behavior label distribution corresponding to the video input from the terminal MT and the behavior label estimated by the behavior estimation processing unit 17. The correct behavior label expresses behavior as "1" and other actions as "0".

[0073] The parameter update processing unit 18 then updates the parameters of each learning model stored in the parameter storage unit 32 based on preset constraint conditions. The constraint conditions define conditions for updating each parameter so that the shape of the correct behavior label distribution matches the shape of the estimated behavior label distribution, for example.

[0074] Any learning method may be used for the parameter update process as long as it satisfies the above constraints. For example, when the correct behavioral label distribution is p(x) and the estimated behavioral label distribution is q(x), it is possible to use a method in which the cross entropy is calculated according to the following formula and the parameters are updated to reduce this cross entropy. -Σ x p(x) log(q(x)).

[0075] The behavior recognition device AS repeatedly performs the above-described learning process, for example, each time it acquires video data including multiple people and their surroundings that are expected to be recognized and the corresponding correct behavior labels, and terminates the learning process when the cross entropy becomes less than a threshold value.

[0076] (2) Inference Phase After the learning process in the learning phase is completed, when the inference phase is set, the control unit 1 of the behavior recognition device AS executes the inference process as follows. In the inference process, detailed explanations of the same processing parts as those in the learning process described above will be omitted, and only the different parts will be explained.

[0077] (2-1) Acquisition of video data and extraction of video feature information In the inference phase, the control unit 1 of the behavior recognition device AS, under the control of the video data acquisition processing unit 11, acquires first-person perspective video data and third-person perspective video data, which are captured by the camera CM and show the person to be recognized and the surrounding situation, via the input / output I / F unit 4.

[0078] When the video data is acquired, the control unit 1 of the behavior recognition device AS reads the video data frame by frame from the video data storage unit 31 under the control of the video feature information extraction processing unit 12, using a learning model. Then, the control unit 1 repeatedly executes a process of extracting features each consisting of a one-dimensional vector from each of the read frame images. Note that, in this case as well, the video feature information may be any information that expresses the vectors of the video.

[0079] (2-2) Extraction of environmental segments and their labels and generation of environmental information The control unit 1 of the behavior recognition device AS reads video data frame by frame from the video data storage unit 31 using a learning model under the control of the environment segment extraction processing unit 13. Then, from each of the read frame images, partial areas in which objects are depicted in the frame image are extracted as environment segments, and environment segment labels indicating the names of the objects present in the partial areas are extracted. This process is repeated. The environment segment extraction processing unit 13 then outputs pairs of the extracted environment segments and environment segment labels to the environment information generation processing unit 14.

[0080] (2-3) Generation of environmental information The control unit 1 of the behavior recognition device AS then executes a process of generating environmental information using the learning model under the control of the environmental information generation processing unit 14.

[0081] That is, the environment information generation processing unit 14 first receives the environment segments and the environment segment labels from the environment segment extraction processing unit 13. Then, for each frame image, the environment information generation processing unit 14 repeats the process of generating environment information corresponding to the environment segment labels of the objects present in the environment segment for each environment segment extracted from the frame image. The environment information is information that represents the actions that can be performed by people present in the environment segment. Then, for each frame image, the environment information generation processing unit 14 outputs the generated pair of the environment information and the corresponding environment segment to the information fusion processing unit 16.

[0082] (2-4) Detection of activity area Further, under the control of the behavior area detection processing unit 15, the control unit 1 of the behavior recognition device AS executes the process of detecting a behavior area and a peripheral area from a frame image using a learning model as follows.

[0083] That is, the behavior area detection processing unit 15 first reads the video data from the video data storage unit 31 for each frame. Next, it detects from each frame image an area whose degree of involvement with the person's behavior is greater than a preset threshold as the behavior area, and the remaining area as the peripheral area. It then generates a two-dimensional vector of the same size as the frame image, with the detected behavior area set to "1" and the peripheral area set to "0". The above process is repeatedly performed for each frame image. The behavior area detection processing unit 15 then outputs the generated two-dimensional vector represented by the behavior area and peripheral area to the information fusion processing unit 16 as a behavior area mask.

[0084] (2-5) Information fusion Next, the control unit 1 of the behavior recognition device AS, under the control of the information fusion processing unit 16, executes a process of fusing the environmental information of interest and the surrounding environmental information with the video feature information to generate fusion information as follows.

[0085] That is, the information fusion processing unit 16 first receives, for each frame image, video feature information, environmental information, and behavior area mask from the video feature information extraction processing unit 12, the environmental information generation processing unit 14, and the behavior area detection processing unit 15. Next, the information fusion processing unit 16 repeatedly determines, for each environment segment, whether the environment segment overlaps with the behavior area mask, i.e., whether it is an "attention area" or a "peripheral area," and generates attention environment information by fusing the environmental information in the "attention area," and generates peripheral environment information by fusing the environmental information in the "peripheral area."

[0086] Then, for each frame image, the information fusion processing unit 16 connects the generated focused environment information and surrounding environment information with the corresponding video feature information to generate fusion information represented by a new one-dimensional vector corresponding to the frame image, and outputs the fusion information generated for each frame image to the behavior estimation processing unit 17.

[0087] (2-6) Behavior estimation The control unit 1 of the behavior recognition device AS, under the control of the behavior estimation processing unit 17, performs processing to estimate behavior labels from the fusion information using a learning model as follows.

[0088] FIG. 9 is a flowchart showing an example of the processing procedure and processing content of the behavior estimation processing executed by the control unit 1 of the behavior recognition device AS in the inference phase.

[0089] That is, the behavior estimation processing unit 17 first receives fusion information corresponding to each frame image from the information fusion processing unit 16 in step S71. Next, in step S72, based on all the received fusion information, behavior expression information is generated, which expresses the behavior of the person who is the subject of the action as a one-dimensional vector. This processing may be any processing that converts a series of fusion information into a behavior expression. For example, the behavior expression is the final output obtained by inputting the fusion information into an LSTM or GRU.

[0090] Subsequently, in step S73, the behavior estimation processing unit 17 estimates the behavior label distribution based on the generated behavior expression information. This behavior label distribution estimation process may be any process that converts behavior expressions into the behavior label distribution.

[0091] For example, in the case of single-label classification, the action representation is converted into a probability distribution (action label distribution) of the dimension number of predefined action recognition labels using a linear layer and a softmax function. In the case of multi-label classification, the action representation is converted into a distribution (action label distribution) of the dimension number of predefined action labels using a linear layer and a sigmoid function, or a linear layer and a tanh function.

[0092] In step S74, the activity estimation processing unit 17 determines an activity label from the estimated activity label distribution. For example, in the case of single-label classification, the activity label corresponding to the element with the largest value in the activity label distribution is selected. On the other hand, in the case of multi-label classification, a threshold θl for label determination is defined in advance, and an activity label corresponding to an element with a value equal to or greater than θl in the activity label distribution is selected.

[0093] Finally, in step S75, the behavior estimation processing unit 17 outputs the set behavior label to the terminal MT from the input / output I / F unit 4. Note that in the inference phase, the parameter update processing unit 18 does not perform parameter update processing.

[0094] (Actions and Effects) As described above, in one embodiment, for each frame of video data captured by the camera CM, an environment segment indicating the area where an object is depicted and an environment segment label indicating the name of the object are extracted, and environment information representing actions that the actor can perform is generated for each environment segment based on the environment segment label. Furthermore, for each frame of the video data, the frame image area is divided into an action area whose degree of relevance to the person's action is greater than a threshold and a surrounding area other than the action area. For each environment segment, it is determined whether the environment segment is an attention area that overlaps with the action area or a surrounding area that does not overlap with it. Attention environment information is generated from the environment information in the attention area, and surrounding environment information is generated from the environment information in the surrounding area. The attention environment information and surrounding environment information are then fused with video feature information extracted from the video data, the fused information is converted into an action expression, and an action label distribution is estimated based on the action expression.

[0095] That is, in one embodiment, the area is divided into an action area where the degree of relevance to the person's action is greater than a threshold value and the surrounding area, and noteworthy environmental information in the action area, environmental information in the surrounding area, and video feature information extracted from the video data are combined, and the action label distribution is estimated based on this combined information.

[0096] Therefore, it is possible to estimate the behavior of the person being recognized by using environmental information of interest that is closely related to the behavior of the person being recognized and environmental information of the surrounding area, thereby effectively narrowing down candidates for the person's behavior and making it possible to accurately estimate the person's behavior.

[0097] For example, the hand movements and posture of a person when "cutting ingredients" are very similar to those when "driving a nail," and in this case it is difficult to accurately estimate a person's actions from hand movements and posture alone.

[0098] However, the action of "cutting ingredients" requires that there be a cutting board, knife, vegetables, fruits, and other ingredients around the person, and the action of "driving a nail" requires that there be a hammer, nails, wood, and other materials around the person.

[0099] In one embodiment of the present invention, a person's behavior is estimated by narrowing down candidates for the person's behavior based on environmental information such as surrounding objects that are closely related to the person's behavior. This makes it possible to distinguish between the behavior of "cutting ingredients" and the behavior of "hammering a nail," and accurately estimate the person's behavior based on the results.

[0100] Generally, environmental information surrounding a person often provides the person with options for their actions. For this reason, utilizing environmental information surrounding a person to estimate a person's actions, as in one embodiment of the present invention, is extremely effective. For example, consider a person in a kitchen environment where a frying pan containing ingredients is placed on a stovetop and a turner and condiments are placed next to it. In this case, the surrounding environment provides the person with limited options, such as "stir-frying" and "seasoning," while actions performed in the same kitchen, such as "chopping" and "washing dishes," can be excluded from the options. Therefore, when a person is performing a certain action, focusing on the environmental information surrounding the person can effectively narrow down the person's options for their actions. This can shorten the time required for estimation, reduce the processing load of the action recognition device AS, and further reduce the size of the learning model used in the series of estimation processes, resulting in extremely significant benefits.

[0101] [Other embodiments] (1) In the above embodiment, the surrounding environmental information that is closely related to the behavior of the person as the subject of the recognition target is used to narrow down candidates for the behavior of the person, and the behavior of the person is estimated from among these candidates. However, the present invention is not limited to this.

[0102] For example, information representing the movement and posture of the main body of the person who is the main actor may be detected from video data from a third-person perspective. The detected information representing the movement and posture of the main body of the person may then be input to the behavior estimation processing unit 17 together with the fusion information of the attention environment information and the surrounding environment information with the video feature information described in the embodiment, and the behavior estimation processing unit 17 may then estimate the behavior of the person. In this way, the behavior of the person who is the main actor may be effectively narrowed down based on environmental information that is closely related to the behavior of the person who is the main actor, and by further taking into account the information representing the movement and posture of the main body of the person, the behavior of the person who is the main actor may be estimated with higher accuracy.

[0103] (2) In the above embodiment, the functions of the behavior recognition device AS are provided in an information processing device such as a server computer or a personal computer that is provided independently of the camera CM and the terminal MT. However, the present invention is not limited to this, and all or part of the functions of the behavior recognition device AS may be provided in the camera CM and the terminal MT.

[0104] (3) In the above embodiment, a person is used as an example of an actor, and the case where the behavior of this person is estimated is described, but the actor to be recognized is not limited to a person, and may be, for example, an animal or a robot. Furthermore, the objects detected as environmental information may include, in addition to objects, other people different from the person to be recognized.

[0105] In addition, the type and configuration of the behavior recognition device, the processing procedures and processing contents of each processing unit, etc. can be modified in various ways without departing from the spirit of the present invention.

[0106] Although the embodiments of the present invention have been described in detail above, the above description is merely an example of the present invention in every respect. It goes without saying that various improvements and modifications can be made without departing from the scope of the present invention. In other words, when implementing the present invention, specific configurations according to the embodiments may be appropriately adopted.

[0107] In short, this invention is not limited to the above-described embodiments, and in the implementation stage, the components can be modified and embodied without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, some components may be omitted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined. [Explanation of symbols]

[0108] AS…Action recognition device CM...camera MT...Terminal 1...Control unit 2...Program memory section 3...Data storage unit 4...Input / output interface 5...Bus 11...Video data acquisition processing unit 12...Video feature information extraction processing section 13...Environment segment extraction processing section 14...Environmental information generation processing unit 15...Behavior area detection processing unit 16...Information Fusion Processing Unit 17...Behavior estimation processing unit 18...Parameter update processing section 31...Video data storage unit 32...Parameter storage section

Claims

1. a video data acquisition processing unit that acquires video data capturing an area including the subject and its surrounding area; a motion object detection processing unit that detects actor information representing a movement or posture of the actor from the acquired video data; an environmental information detection processing unit that detects environmental information of a surrounding area from the acquired video data, the environmental information being related to the behavior of the performer to a predetermined condition; an action estimation processing unit that estimates candidates for the action of the actor based on fusion information of the detected environmental information and video feature information extracted from the video data, and determines the action of the actor by further taking the actor information into account in the estimated candidates for the action; An activity recognition device comprising:

2. a video data acquisition processing unit that acquires video data capturing an area including at least a peripheral area of ​​the performer; an environmental information detection processing unit that detects environmental information of a surrounding area whose degree of association with the behavior of the performer satisfies a predetermined condition from the acquired video data; a behavior estimation processing unit that estimates candidates for the behavior of the subject based on the detected environmental information and video feature information extracted from the video data, and determines the behavior of the subject from the estimated candidates for the behavior; An activity recognition device comprising:

3. The environmental information detection processing unit a processing unit that detects an environment segment that represents an imaging area of ​​an object included in the acquired video data, and the environment information that represents an attribute of the object for each environment segment; a processing unit that detects, from the acquired video data, a peripheral area whose degree of association with the action of the subject satisfies the predetermined condition as an action area; a processing unit that classifies the environmental segments into attention environmental segments included in the behavioral area and other surrounding environmental segments, generates attention environmental information from the attention environmental segments and the corresponding environmental information, and generates surrounding environmental information from the surrounding environmental segments and the corresponding environmental information; Equipped with The behavior estimation processing unit estimates candidates for the behavior of the subject based on the generated attention environment information and the surrounding environment information and video feature information extracted from the video data, and determines the behavior of the subject from the estimated candidates for the behavior. The behavior recognition device according to claim 2 .

4. 4. The behavior recognition device according to claim 2, wherein the environmental information detection processing unit detects, as the environmental information of the surrounding area whose degree of involvement with the behavior of the performer satisfies the predetermined condition, information corresponding to either a surrounding first area having a base point at a part of the performer or a surrounding second area corresponding to the line of sight of the performer.

5. 4. The behavior recognition device according to claim 3, wherein the environmental information detection processing unit generates, for each environment segment, information representing an appearance frequency distribution or an appearance probability distribution of the objects appearing in the environment segment based on text data prepared in advance, and expresses the environmental information using the generated information representing the appearance frequency distribution or the information representing the appearance probability distribution.

6. at least one of the environmental information detection processing unit and the behavior estimation processing unit is configured by a learning model using a neural network; a parameter update processing unit that updates parameters of the learning model based on a difference between estimated behavior label distribution information representing candidates for the behavior of the subject estimated by the behavior estimation processing unit based on learning video data and correct behavior label distribution information that is set in advance corresponding to the learning video data; The behavior recognition device according to claim 2 , further comprising:

7. An action recognition method executed by an information processing device, a step of acquiring video data capturing an area including the subject of the action and its surrounding area; detecting actor information representing a movement or posture of the actor from the acquired video data; detecting environmental information of a surrounding area from the acquired video data, the degree of association with the action of the actor satisfying a predetermined condition; a process of estimating candidates for the action of the actor based on fusion information of the detected environmental information and video feature information extracted from the video data, and determining the action of the actor by further taking the actor information into account in the estimated candidates for the action; An activity recognition method comprising:

8. An action recognition method executed by an information processing device, a step of acquiring video data capturing an area including at least the area surrounding the performer; detecting, from the acquired video data, environmental information of a surrounding area whose degree of association with the action of the performer satisfies a predetermined condition; a step of estimating candidates for the action of the subject based on the detected environmental information and image feature information extracted from the image data, and determining the action of the subject from the estimated candidates for the action; An activity recognition method comprising:

9. 7. A program causing a processor included in the behavior recognition device to execute processing by each processing unit included in the behavior recognition device according to claim 1.

Citation Information

Patent Citations

  • Notification system

    JP2009260850A

  • Image processing device, image processing method and image processing program

    JP2018206321A