Behavior recognition learning device, behavior recognition estimation device, behavior recognition learning method, and behavior recognition learning program

The action recognition system integrates behavioral and textual features to accurately identify important basic actions related to procedures, addressing the challenges of existing technologies by enhancing the matching of procedural wording with actual actions.

JP7835304B2Active Publication Date: 2026-03-25NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing action recognition technologies struggle to accurately identify important basic actions related to procedures due to the redundancy and irrelevance of video data, making it difficult to match procedural wording with actual actions.

Method used

An action recognition system that includes an action feature extraction unit, a text feature generation unit, and a feature fusion unit to integrate behavioral and textual features, using a procedure judgment model to update and determine procedure information.

Benefits of technology

The system accurately identifies important basic actions related to procedures, enabling precise matching of abstract actions with procedural information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007835304000001
    Figure 0007835304000001
  • Figure 0007835304000002
    Figure 0007835304000002
  • Figure 0007835304000003
    Figure 0007835304000003
Patent Text Reader

Abstract

An action-recognition learning device according to the present invention includes an action-feature extraction unit, a text-feature generation unit, and a feature fusion unit. The action-feature extraction unit receives media information for learning as inputs and extracts action features from the media information for learning. The text-feature generation unit receives action information for learning as inputs and generates text features from the action information for learning. The feature fusion unit holds a procedure determination model, receives the action features, the text features, and procedure information for learning as inputs, and updates the procedure determination model on the basis of the action features, the text features, and the procedure information for learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005] ,

[0001] The present invention relates to an action recognition learning device, an action recognition estimation device, an action recognition learning method, and an action recognition learning program.

Background Art

[0002] In the background of labor shortages, digitalization of operations is required in various industries. When digitalizing operations, it is necessary to grasp how the operations to be digitalized are performed. By grasping the operations, for example, it becomes possible to analyze and improve the operations by visualizing the actions of store employees to improve efficiency, or to grasp medical practices and other actions by visualizing the actions of medical workers.

[0003] However, it has become difficult to accurately grasp operations. If an operation manual is prepared to grasp the operations, the operation manual becomes a useful information source. However, when comparing the operation manual with the actual operations, there are many cases where it is difficult to grasp the actions, such as the actions described in the operation manual being different from the actual actions, and the procedures in the operation manual being different from the actual operation procedures.

[0004] In addition, observing human actions using a camera or the like requires a huge amount of time. In order to automatically acquire human actions from a camera, it is conceivable to use action recognition technology to visualize the actual actions of a person and utilize it for analysis and improvement of operations. An example of action recognition technology is disclosed in Non-Patent Document 1, for example.

[0005] In behavior recognition technology, defining the behavior to be recognized (defining behavior labels) to match industry and business type would reduce versatility; therefore, it is usually defined as targeting general behaviors. Here, behavior refers to actions that represent human movements such as walking, running, and holding, and behavior labels are verbs that express actions such as walking and running. As a result, when trying to match these with procedures in a business manual, discrepancies occur because the procedural wording is more abstract. For example, if a manual describes placing a product on a shelf, it may implicitly include actions such as grasping the product or carrying the product before placing it. In other words, since the expression of a procedure described in a manual may consist of multiple basic actions, it is difficult to match the behavior recognition technology learned from basic actions with the expression in the manual. [Prior art documents] [Non-patent literature]

[0006] [Non-Patent Document 1] Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming He, “SlowFast Networks for Video Recognition”, ICCV2019. [Overview of the project] [Problems that the invention aims to solve]

[0007] When using information on basic actions output by an action recognition model to determine which procedure in a business manual or similar document corresponds to a given action, it is necessary to identify the basic actions associated with that procedure. However, video data is redundant, and therefore contains a large amount of irrelevant basic action information. Consequently, identifying the important basic actions related to a procedure is difficult.

[0008] This invention was made in view of the above circumstances, and its purpose is to provide an action recognition technology that can accurately identify important basic actions related to procedures. [Means for solving the problem]

[0009] One aspect of the present invention is an action recognition learning device. The action recognition learning device includes an action feature extraction unit, a text feature generation unit, and a feature fusion unit. The action feature extraction unit receives learning media information as input and extracts action features from the learning media information. The text feature generation unit receives learning action information as input and generates text features from the learning action information. The feature fusion unit holds a procedure judgment model and receives action features, text features, and learning procedure information as input and updates the procedure judgment model based on the action features, text features, and learning procedure information.

[0010] One aspect of the present invention is an action recognition estimation device. The action recognition estimation device comprises a procedure sentence vector DB and a procedure determination unit. The procedure sentence vector DB holds a plurality of procedure sentence vectors. The procedure determination unit receives a procedure determination model updated based on learning media information, learning action information, and learning procedure information, as input, media information, and action information. The procedure determination unit calculates the similarity between the estimated procedure sentence vector output by the procedure determination model, which takes the media information and action information as input, and each of the plurality of procedure sentence vectors held in the procedure sentence vector DB, and outputs the procedure sentence of the procedure sentence vector with the highest similarity as procedure information.

[0011] One aspect of the present invention is, The computer executes This is an action recognition learning method. The computer executes The behavior recognition learning method includes the steps of: receiving learning media information as input and extracting behavioral features from the learning media information; receiving learning behavioral information as input and generating textual features from the learning behavioral information; and receiving learning procedure information, behavioral features, and textual features as input and updating a procedure judgment model based on the learning procedure information, behavioral features, and textual features.

[0012] One aspect of the present invention is an action recognition learning program. The action recognition learning program is a computer having a processor and a memory device, and the above-mentioned action recognition learning device Function Make it run. [Effects of the Invention]

[0013] According to the present invention, an action recognition technology is provided that can accurately identify important basic actions associated with a procedure. [Brief explanation of the drawing]

[0014] [Figure 1] Figure 1 is a block diagram showing the functional configuration of the behavior recognition learning system according to the first embodiment. [Figure 2] Figure 2 is a block diagram showing the configuration of a computer that can constitute the action recognition learning device and the action recognition estimation device, respectively, according to the first embodiment. [Figure 3] Figure 3 is a flowchart showing the flow of the text feature generation process performed by the text feature generation unit of the behavior recognition learning device according to the first embodiment. [Figure 4] Figure 4 is an illustrative diagram showing how the text feature generation unit of the behavior recognition learning device according to the first embodiment generates prompt information from behavior information. [Figure 5] Figure 5 is a flowchart showing the flow of the procedure determination model update process executed by the feature fusion unit of the action recognition learning device according to the first embodiment. [Figure 6] Figure 6 is a flowchart showing the processing flow executed by the procedure determination unit of the action recognition estimation device according to the first embodiment. [Figure 7] Figure 7 is a block diagram showing the functional configuration of the action recognition learning system according to the second embodiment. [Figure 8] Figure 8 is a flowchart showing the flow of the procedure determination model update process executed by the feature fusion unit of the action recognition learning device according to the second embodiment. [Modes for carrying out the invention]

[0015] Embodiments of the present invention will be described below with reference to the drawings.

[0016] <First Embodiment> (Functional Configuration) First, referring to FIG. 1, the functional configuration of the action recognition system 10 according to the first embodiment will be described. FIG. 1 is a block diagram showing an example of the functional configuration of the action recognition system 10 according to the first embodiment.

[0017] The action recognition system 10 includes an action recognition learning device 20 and an action recognition estimation device 30. The action recognition learning device 20 and the action recognition estimation device 30 can transmit and receive information via, for example, a network. The network may be wireless or wired.

[0018] The action recognition learning device 20 holds a procedure determination model, receives learning media information, learning action information, and learning procedure information as inputs, and updates the procedure determination model based on these information. The action recognition learning device 20 is also a device that provides the procedure determination model to the action recognition estimation device 30 in response to a request from the action recognition estimation device 30.

[0019] The action recognition estimation device 30 receives media information and action information as inputs, and calculates and outputs procedure information using the procedure determination model.

[0020] (Action Recognition Learning Device 20) First, the action recognition learning device 20 will be described. The action recognition learning device 20 includes an action feature extraction unit 21, a text feature generation unit 22, and a feature fusion unit 23.

[0021] The action feature extraction unit 21 receives learning media information as an input, and extracts action features Vi from the learning media information. The action feature extraction unit 21 also outputs the extracted action features Vi to the feature fusion unit 23.

[0022] Here, the learning media information is, for example, moving image data. The behavior feature extraction unit 21 segments the received moving image data to obtain moving image segments. A moving image segment is data in a unit where a plurality of frames of a moving image are grouped together. For example, there is an example where the number of frames for 5 seconds is grouped as one segment. When considering long-term moving images of hundreds or thousands of frames, a set of frames sampled at regular intervals may also be used as a moving image segment.

[0023] Also, the learning media information may be not moving image data, but voice data, acoustic data, or the like. The behavior feature extraction unit 21 may segment and handle these voice data, acoustic data, etc. in the same way as moving image data. Hereinafter, the learning media information will be described as being moving image data.

[0024] The behavior feature V_i is a feature extracted using a behavior recognition model for each moving image segment i (0 < i ≤ N). An example of the extraction method of the behavior feature V_i using the behavior recognition model is disclosed in, for example, Non-Patent Document 1. Specifically, the behavior feature V_i is a multi-dimensional vector obtained by extracting the output of the final connection layer of a pre-trained behavior recognition model.

[0025] The text feature generation unit 22 receives the learning behavior information as an input and generates text features T_j from the learning behavior information. Specifically, the text feature generation unit 23 holds a plurality of templates prepared in advance, randomly selects a template for the learning behavior information, combines the selected template with the learning behavior information to generate prompt information, and generates the text features from the prompt information. The text feature generation unit 22 also outputs the generated text features T_j to the feature fusion unit 23.

[0026] Here, the learning behavior information is an action label for each video segment i including the basic action information. For example, when the video segment i has action labels of {Sitting in a chair}, {Holding a laptop}, and {Opening a laptop}, these are collectively regarded as the learning behavior information. The learning behavior information belonging to the video segment i is denoted as A_j (0 < j ≤ M). Here, M is the number of action labels. The above is an example where the number of action labels M = 3.

[0027] The plurality of templates prepared in advance are texts such as “A person is {___}.”, “He / she is {___}.”, “The person in the scene is {___}.”, …. The text feature generation unit 23 selects one template for each action label of the learning behavior information and combines them to generate prompt information. The prompt information is, for example, a text such as “A person is sitting in a chair.” obtained by combining an action label of the learning behavior information such as {Sitting in a chair} and a template such as “A person is {___}.”

[0028] The text feature T_j is a feature extracted by a text feature extractor using a pre-trained language model with the learning behavior information. An example of the text feature extractor is disclosed in the following Reference 1. 〔Reference 1〕Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding” NAACL, 2019.

[0029] The feature fusion unit 23 holds a procedure determination model. The feature fusion unit 23 receives behavioral features extracted by the behavioral feature extraction unit 21, textual features generated by the textual feature generation unit 22, and learning procedure information as input, and updates the procedure determination model based on this information. Here, the learning procedure information is text that describes the actions of the entire video, such as {A person is opening a laptop and making slides for a presentation.}. The feature fusion unit 23 also outputs the procedure determination model to the behavioral recognition estimation device 30 in response to a request from the behavioral recognition estimation device 30.

[0030] (Action recognition estimation device 30) Next, the behavior recognition estimation device 30 will be described. The behavior recognition estimation device 30 includes a procedure statement vector DB 31 and a procedure determination unit 32.

[0031] The procedure statement vector DB31 holds multiple procedure statement vectors. In other words, the procedure statement vector DB31 is a database that holds procedure statements and the textual features extracted from those procedure statements by a text feature extractor.

[0032] The procedure determination unit 32 receives a procedure determination model as input from the action recognition learning device 20. The procedure determination unit 32 also receives media information and action information as input. The procedure determination model outputs an estimated procedure sentence vector using the media information and action information as input. The procedure determination unit 32 calculates the similarity between the estimated procedure sentence vector output by the procedure determination model and each of the multiple procedure sentence vectors held by the procedure sentence vector DB 31, and outputs the procedure sentence of the procedure sentence vector with the highest similarity as procedure information.

[0033] Here, the media information, behavioral information, and procedure information are in the same data format as the learning media information, learning behavioral information, and learning procedure information mentioned above.

[0034] (Hardware configuration) Next, we will describe the hardware configurations of the behavior recognition learning device 20 and the behavior recognition estimation device 30. For example, the behavior recognition learning device 20 and the behavior recognition estimation device 30 are each composed of a computer in terms of hardware. The computer can be, for example, a personal computer or a server computer.

[0035] Figure 2 is a block diagram showing the configuration of a computer 40 that can constitute the behavior recognition learning device 20 and the behavior recognition estimation device 30 according to the embodiment. As shown in Figure 2, the computer 40 has a processor 41, a ROM (Read Only Memory) 42, a RAM (Random Access Memory) 43, an auxiliary storage device 44, an input / output interface 45, and a communication interface 46.

[0036] The processor 41, ROM 42, RAM 43, auxiliary storage device 44, input / output interface 45, and communication interface 46 are electrically connected to each other via a bus 47, and data is exchanged via the bus 47.

[0037] The processor 41 is composed of a general-purpose hardware processor, such as a CPU (Central Processing Unit) or a GPU (Graphical Processing Unit). The processor 41 controls the entire system, including the ROM 42, RAM 43, auxiliary storage device 44, input / output interface 45, and communication interface 46.

[0038] ROM42 is a non-volatile memory that constitutes part of the main memory. ROM42 non-temporarily stores the startup program required when the processor 41 starts up. The processor 41 starts up by executing the program in ROM42. ROM42 is, for example, composed of EPROM (Erasable Programmable Read Only Memory) and stores various startup settings in addition to the startup program.

[0039] RAM43 is a volatile memory that constitutes part of the main memory. RAM43 temporarily stores the program necessary for processing by the processor 41 and the data necessary for executing the program. The processor 41 executes the program in RAM43, performs calculations on the data in RAM43, and stores the calculation results in RAM43.

[0040] The auxiliary storage device 44 consists of non-volatile memory such as an HDD (Hard Disk Drive) or SSD (Solid State Drive). The auxiliary storage device 44 non-temporarily stores programs executed by the processor 41 and data necessary for program execution. The processor 41 reads the programs and data from the auxiliary storage device 44 into the RAM 43 and executes various functions by running the programs.

[0041] The input / output interface 45 is connected to an external input device 51 and an output device 52, etc., enabling the input of information from the input device 51 and the output of information to the output device 52. For example, the input / output interface 45 may be a wired interface or a wireless interface. A wired interface includes a port to which the device is connected. A wireless interface includes Bluetooth®, WiFi®, etc.

[0042] The input device 51 may include a keyboard, mouse, touch panel, receiver, disk drive, etc. The input device 51 is not limited to these and may include any other input device. The output device 52 may include a display, transmitter, disk drive, etc. The output device 52 is not limited to these and may include any other output device. The input device 51 and the output device 52 may be configured as an input / output device 53 that has the functions of both.

[0043] The input device 51 connected to the computer 40 that constitutes the behavior recognition learning device 20 has the function of inputting learning media information, learning behavior information, and learning procedure information to the behavior recognition learning device 20. The input device 51 connected to the computer 40 that constitutes the behavior recognition estimation device 30 has the function of inputting media information and behavior information to the behavior recognition estimation device 30.

[0044] The communication interface 46 is an interface that enables the sending and receiving of information with the network. This allows the action recognition learning device 20 and the action recognition estimation device 30 to send and receive information. In other words, the action recognition estimation device 30 can send information to the action recognition learning device 20 requesting the provision of a procedure judgment model. The action recognition learning device 20 can then provide the procedure judgment model to the action recognition estimation device 30.

[0045] A program stored non-temporarily in the auxiliary storage device 44 is provided to the computer 40, for example, via a recording medium 54 that is readable by the computer 40 on which the program was stored non-temporarily. Such a recording medium 54 is called a non-temporarily computer-readable recording medium. Non-temporarily computer-readable recording media include disks such as flexible disks, optical disks (CD-ROM, CD-R, DVD-ROM, DVD-R, etc.), magneto-optical disks (MO, etc.), and semiconductor memory.

[0046] The program stored non-temporarily in the auxiliary storage device 44 of the computer 40 that constitutes the behavior recognition learning device 20 includes a behavior recognition learning program. The behavior recognition learning program is a program that causes the computer 40 that constitutes the behavior recognition learning device 20 to execute the functions of the behavior feature extraction unit 21, the text feature generation unit 22, and the feature fusion unit 23.

[0047] The program stored non-temporarily in the auxiliary storage device 44 of the computer 40 that constitutes the behavior recognition estimation device 30 includes a behavior recognition estimation program. The behavior recognition estimation program is a program that causes the computer 40 that constitutes the behavior recognition estimation device 30 to execute the functions of the procedure statement vector DB 31 and the procedure determination unit 32.

[0048] Programs stored non-temporarily in the auxiliary storage device 44 are read into and stored non-temporarily in the auxiliary storage device 44 via the input device 51, which is a disk drive, and the input / output interface 45, if the recording medium 54 is a disk, or via the input / output interface 45, which is a port, if the recording medium 54 is semiconductor memory. Alternatively, the program may be stored on a server on a network, downloaded from the server, and stored non-temporarily in the auxiliary storage device 44.

[0049] When the computer 40 starts up, the processor 41 executes a program in the ROM 42 and loads the OS into the RAM 43 to start up. Under the control of the OS, the processor 41 monitors instruction inputs and the connection of external devices. Also, under the control of the OS, the processor 41 sets up a program area and a data area in the RAM 43. In response to an instruction input to start the device (behavior recognition learning device 20 or behavior recognition estimation device 30), the processor 41 loads a program (behavior recognition learning program or behavior recognition estimation program) from the auxiliary storage device 44 into the program area of ​​the RAM 43, and loads the data necessary for program execution from the auxiliary storage device 44 into the data area of ​​the RAM 43. The processor 41 performs calculations on the data in the data area according to the program and writes the calculation results to the data area. Through these operations, the processor 41, RAM 43, auxiliary storage device 44, input / output interface 45, and communication interface 46 work together to execute the functions of each component of the behavior recognition learning device 20 or behavior recognition estimation device 30.

[0050] (Text feature generation process) Next, with reference to Figure 3, the details of the text feature generation process performed by the text feature generation unit 22 of the behavior recognition learning device 20 will be explained. Figure 3 is a flowchart showing the flow of the text feature generation process performed by the text feature generation unit 22.

[0051] First, the text feature generation unit 22 receives learning behavior information A_j, which is present in each video segment i, as input (step S11).

[0052] The text feature generation unit 22 holds multiple pre-prepared templates. The text feature generation unit 22 randomly selects a template for each action information A_j and generates prompt information t_j from each action information A_j (step S12).

[0053] Here, the process of generating prompt information t_j in step S12 will be explained in more detail with reference to Figure 4. Figure 4 is a diagram illustrating an example of generating prompt information from action information belonging to video segment i.

[0054] In the example in Figure 4, video segment i is 5 seconds (5s) of video data, and the learning behavior information A_j belonging to video segment i has three behavior labels: {Sitting in a chair}, {Holding a laptop}, and {Opening a laptop}. The text feature generation unit 22 holds several pre-prepared templates such as "A person is {___}.", "He / she is {___}.", "The person in the scene is {___}.", etc.

[0055] The text feature generation unit 22 takes action information A_1 such as {Sitting in a chair}, selects and combines a template such as “A person is {___}.” to generate prompt information t_1 such as “A person is sitting in a chair.” Similarly, the text feature generation unit 22 takes action information A_2 such as {Holding a laptop}, selects and combines a template such as “He / she is {___}.” to generate prompt information t_2 such as “He / she is holding a laptop.” And takes action information A_3 such as {Opening a laptop}, selects and combines a template such as “The person in the scene is {___}.” to generate prompt information t_3 such as “The person in the scene is opening a laptop.”

[0056] Next, the text feature generation unit 22 generates text features T_j from prompt information t_j using a pre-trained language model as a text feature extractor (step S13). An example of a text feature extractor is disclosed in the aforementioned reference 1.

[0057] Finally, the text feature generation unit 22 outputs the text feature T_j to the feature fusion unit (step S14).

[0058] (Update process for the procedure decision model) Next, with reference to Figure 5, the details of the procedure determination model update process performed by the feature fusion unit 23 of the action recognition learning device 20 will be explained. Figure 5 is a flowchart showing the flow of the procedure determination model update process performed by the feature fusion unit 23.

[0059] First, the feature fusion unit 23 receives the behavioral features V_i and textual features T_j of each video segment i as input from the behavioral feature extraction unit 21 and the textual feature generation unit 22, respectively (step S21).

[0060] Next, the feature fusion unit 23 generates a semantic feature S_i of the video segment i using the behavioral feature V_i and the textual feature T_j (step S22).

[0061] Here, we will describe a specific method for generating semantic features S_i. For example, the feature fusion unit 23 generates semantic features S_i using an attention mechanism. Examples of attention mechanisms are disclosed in the aforementioned reference 1 and in the following reference 2. [Reference 2] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin, “Attention Is All You Need” NIPS, 2017.

[0062] The feature fusion unit 23 generates the behavioral feature V_i as the query Q_i, the textual feature T_j as the key K'_j, and the value V'_j. The query Q_i is obtained by the matrix product of the weight W_q and the behavioral feature V_i. The key K'_j is obtained by the matrix product of the weight W_k and the textual feature T_j. The value V'_j is obtained by the matrix product of the weight W_v and the textual feature T_j.

[0063] The feature fusion unit 23 calculates the weight for each value V'_j by taking the inner product of each key K'_j with the query Q_i. Then, the feature fusion unit 23A takes the weighted sum of these weights and each value V'_i as the semantic feature S_i. The weights W_q, W_k, and W_v are obtained through learning. The method for obtaining weights through learning is disclosed, for example, in the aforementioned references 1 and 2.

[0064] Here, we have described an example of generating semantic features S_i using an attention mechanism. However, instead of using an attention mechanism, a mechanism that can appropriately weight each sentence feature with behavioral features may also be used to generate semantic features S_i.

[0065] Next, the feature fusion unit 23 combines the behavioral feature V_i and semantic feature S_i of the video segment i to generate the segment feature F_i (step S23).

[0066] Next, the feature fusion unit 23 adds positional information obtained by positional encoding to the segment features F_i of each video segment i and generates a procedural feature F_V using a fully connected layer or the like (step S24). An example of positional encoding is disclosed in the aforementioned reference 2.

[0067] Next, the feature fusion unit 23 calculates the cosine similarity between the procedure statement feature F_V and the sentence features T_i extracted by the text feature extractor using each learning procedure information Act_i (step S25).

[0068] Next, the feature fusion unit 23 calculates a loss, for example, a cross-entropy loss, based on the determination by cosine similarity (step S26). If backpropagation is possible, a loss other than the cross-entropy loss may be used. Here, when calculating the cross-entropy loss, assuming there is a vector with the same number of dimensions as the number of procedure information, the one-hot vector is used as the ground truth data, in which the procedure information corresponding to the input video is assigned 1 and the other procedure information is assigned 0.

[0069] Next, the feature fusion unit 23 calculates the cross-entropy loss with respect to the cosine similarity vector with each procedure calculated in step S25. The calculated cross-entropy loss is backpropagated, and the parameters of the procedure decision model are updated (step S27).

[0070] Here, the procedure judgment model is a neural network model including the action recognition learning device 20. The parameters to be learned may be limited, such as fixing the parameters of the action feature extraction unit 21. Parameter updates are repeated with respect to the prepared training data until a certain criterion is reached, similar to the learning process of a general neural network. The criterion may be that the number of epochs reaches the upper limit, or that the loss falls below a certain value.

[0071] (Procedure information estimation process) Next, with reference to Figure 6, the details of the procedure information estimation process performed by the behavior recognition estimation device 30 will be explained. Figure 6 is a flowchart showing the flow of the procedure information estimation process performed by the behavior recognition estimation device 30.

[0072] First, the behavior recognition estimation device 30 requests the behavior recognition learning device 20 to provide a procedure determination model. In response, the behavior recognition learning device 20 outputs the procedure determination model to the procedure determination unit 32 of the behavior recognition estimation device 30. The procedure determination unit 32 receives the procedure determination model as input. The procedure determination unit 32 also receives media information and behavior information as input (step S31).

[0073] The procedure determination model takes media information and behavioral information as input and outputs an estimated procedure vector F'_V (step S32). In other words, the procedure determination unit 32 takes media information and behavioral information as input and uses the procedure determination model to obtain an estimated procedure vector F'_V.

[0074] The procedure determination unit 32 refers to the procedure sentence vector DB 31 and calculates the similarity with the estimated procedure vector F'_V (step S33). Here, the procedure sentence vector DB 31 is a database that holds procedure sentences and sentence features T_i' extracted from those procedure sentences by a text feature extractor. Here, i' is the subscript of the video of the procedure to be searched. The procedure determination unit 32 calculates the similarity between the estimated procedure vector F'_V and each of the procedure sentence vectors in the procedure sentence vector DB 31. Any search method is acceptable, such as clustering the procedure sentence vectors, calculating the similarity with the centroids of the clusters, and further searching for procedure sentence vectors with the most similar centroids.

[0075] Finally, the procedure determination unit 32 outputs the procedure statement of the procedure statement vector with the highest similarity as procedure information (step S34). The procedure information may include pairs of the similarity of the search results and the procedure statements.

[0076] (effect) As described above, according to this embodiment, a procedure determination model can be obtained that takes into account not only the behavioral features obtained from the learning media information, but also the textual features obtained from the learning behavioral information. This makes it possible to estimate procedures that include abstract actions rather than basic actions that represent physical actions. As a result, an action recognition technology is provided that can identify important basic actions related to procedures with high accuracy.

[0077] <Second Embodiment> (Functional Configuration) Next, with reference to Figure 7, the functional configuration of the behavior recognition system 10A according to the second embodiment will be described. Figure 7 is a block diagram showing an example of the functional configuration of the behavior recognition system 10A according to the second embodiment.

[0078] The behavior recognition system 10A according to the second embodiment differs from the behavior recognition learning device 20 according to the first embodiment in that the behavior recognition learning device 20A according to the second embodiment differs from the behavior recognition learning device 20 according to the first embodiment.

[0079] The behavior recognition learning device 20A has an added object feature extraction unit 24 compared to the behavior recognition learning device 20. Consequently, the processing of the feature fusion unit 23A of the behavior recognition learning device 20A differs from the processing of the feature fusion unit 23 of the behavior recognition learning device 20. Otherwise, the configuration of the behavior recognition learning device 20A is the same as that of the behavior recognition learning device 20. In Figure 7, the components indicated by the same reference numerals as those shown in Figure 1 are the same elements, and their detailed explanation is omitted. The following explanation will focus on the differences. In other words, the parts not mentioned in the following explanation are the same as in the first embodiment.

[0080] (Action Recognition Learning Device 20A) The behavior recognition learning device 20A includes a behavior feature extraction unit 21, a text feature generation unit 22, an object feature extraction unit 24, and a feature fusion unit 23A. The behavior feature extraction unit 21 and the text feature generation unit 22 are as described in the first embodiment.

[0081] The object feature extraction unit 24 receives learning media information as an input, and extracts an object feature O_k from the learning media information. The object feature extraction unit 24 also outputs the extracted object feature O_k to the feature fusion unit 23A.

[0082] In this embodiment, the learning media information is moving image data. Hereinafter, the case where the object feature extraction unit 24 uses one image at a time will be described.

[0083] The object feature extraction unit 24 segments the received moving image data to obtain moving image segments. The object feature extraction unit 24 extracts, for example, one image from a plurality of images included in the moving image segment. The extraction of one image may be performed, for example, by limiting the time axis of the moving image segment. For example, when the segment length of the moving image segment is 5 seconds, the object feature extraction unit 24 extracts an image corresponding to 2.5 seconds.

[0084] Next, the object feature extraction unit 24 performs object detection on the extracted single image. An example of object detection is disclosed in Reference 3 below. 〔Reference 3〕Shaoqing Ren, Kaiming He, Ross Girshick, Jian Sun, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks”, CVPR, 2015. Here, as an example, a method of extracting one image from a moving image segment and performing object detection on the extracted single image has been described. Instead of this method, object detection may be performed on all the images of the moving image segment to detect an object.

[0085] Then, the object feature extraction unit 24 extracts a feature vector obtained from a coupling layer close to the final layer of the object detection model for each detected object, and outputs it as object features O_k (0 < k ≦ K) corresponding to the number of objects to the feature fusion unit 23A. The feature fusion unit 23A uses the input object features O_k for generating semantic features S_i.

[0086] (Update process for the procedure decision model) Next, with reference to Figure 8, the details of the procedure determination model update process performed by the feature fusion unit 23A of the action recognition learning device 20A will be explained. Figure 8 is a flowchart showing the flow of the procedure determination model update process performed by the feature fusion unit 23A.

[0087] First, the feature fusion unit 23A receives the behavioral feature V_i, object feature O_k, and text feature T_j of each video segment i as input from the behavioral feature extraction unit 21, the object feature extraction unit 24, and the text feature generation unit 22, respectively (step S41).

[0088] Next, the feature fusion unit 23A generates a semantic feature S_i of the video segment i using the behavioral feature V_i, the text feature T_j, and the object feature O_k (step S42).

[0089] Here, we will describe a specific method for generating semantic features S_i. For example, the feature fusion unit 23 generates semantic features S_i using an attention mechanism. Examples of attention mechanisms are disclosed in the aforementioned references 1 and 2.

[0090] The feature fusion unit 23 generates the behavioral feature V_i as query Q_i, the textual feature T_j as key K'_j, and the value V'_j. The query Q_i is obtained by the matrix product of the weight W_q and a vector formed by combining the behavioral feature V_i and the object feature O_k.

[0091] Here, we describe how to combine behavioral features V_i and object features O_k. If there are multiple object features O_k, an average vector O_ave of the object features O_k is generated and combined with the behavioral feature V_i. If there is only one object feature O_k, it is combined directly with the behavioral feature V_i. The combination method only needs to be able to represent the two features, and dimensionality reduction using an autoencoder or similar method may be used.

[0092] The key K'_j is obtained by the matrix product of the weight W_k and the text feature T_j. The value V'_j is obtained by the matrix product of the weight W_v and the text feature T_j.

[0093] The feature fusion unit 23A calculates the weight for each value by taking the inner product of each key K'_j with the query Q_i. The feature fusion unit 23A defines the weighted sum of these weights and each value V_i as the semantic feature S_i. As mentioned above, each weight W_q, W_k, and W_v is obtained by learning. The method for obtaining weights by learning is disclosed, as mentioned above, for example in the aforementioned references 1 and 2.

[0094] Here, we have described an example of generating semantic features S_i using an attention mechanism. However, instead of using an attention mechanism, a mechanism that can appropriately weight each sentence feature with behavioral features may also be used to generate semantic features S_i.

[0095] The subsequent steps are the same as in the first embodiment. That is, the feature fusion unit 23A combines the behavioral feature V_i and semantic feature S_i of the video segment i to generate a segment feature F_i (step S43).

[0096] Next, the feature fusion unit 23 adds positional information obtained by positional encoding to the segment features F_i of each video segment i and generates a procedural feature F_V using a fully connected layer (step S44). An example of positional encoding is disclosed in the aforementioned reference 2.

[0097] Next, the feature fusion unit 23A calculates the cosine similarity between the procedure statement feature F_V and the sentence features T_i extracted by the text feature extractor using each learning procedure information Act_i (step S45).

[0098] Next, the feature fusion unit 23A calculates the cross-entropy loss based on the determination by cosine similarity (step S46).

[0099] Next, the feature fusion unit 23 backpropagates the cross-entropy loss and updates the parameters of the procedure decision model (step S47). As a result, object features are taken into consideration within the semantic features S_i.

[0100] (effect) According to this embodiment, a procedure determination model can be obtained that takes into account not only behavioral features obtained from training media information, but also textual features obtained from training behavior information and object information obtained from images contained in the training media information. This makes it possible to estimate procedures that include abstract actions rather than basic actions that represent physical actions. As a result, an action recognition technology is provided that can identify important basic actions related to procedures with high accuracy.

[0101] (Conclusion) As described above, the embodiments provide an action recognition learning device, an action recognition estimation device, an action recognition learning method, and an action recognition learning program that can identify important basic operations related to procedures with high accuracy.

[0102] In this embodiment, an example was described in which the behavior recognition learning device 20 is composed of a computer having a processor and a storage device, the storage device stores a behavior recognition learning program, and the processor executes the behavior recognition learning program to update the procedure determination model.

[0103] However, the behavior recognition learning program may cause the processor to execute a part of the functions of the behavior recognition learning device 20, that is, it may cause the processor to execute the functions of the behavior recognition learning device 20 in combination with a program already recorded in the computer. Alternatively, the behavior recognition learning program may cause the processor to execute the functions of the behavior recognition learning device 20 in combination with hardware such as a PLD (Programmable Logic Device), FPGA (Field Programmable Gate Array), or GPU (Graphic Processing Unit).

[0104] The same can be said for behavior recognition estimation devices and behavior recognition estimation programs.

[0105] Embodiments of the present invention have been described above with reference to the drawings. However, the above embodiments are merely examples of configurations that embody the present invention. In other words, it is clear that the present invention is not limited to the above embodiments. Therefore, additions, omissions, substitutions, and other modifications of components are permitted without departing from the technical spirit of the present invention.

[0106] In short, the present invention is not limited to the embodiments described above, and can be modified in various ways during implementation without departing from its essence. Furthermore, each embodiment may be combined as appropriate, and in that case, the combined effects can be obtained. Moreover, the above embodiments include various inventions, and various inventions can be extracted by selecting combinations from the multiple constituent elements disclosed. For example, if the problem can be solved and effects obtained even if some constituent elements are deleted from all the constituent elements shown in the embodiment, then the configuration with these deleted constituent elements can be extracted as an invention. [Explanation of symbols]

[0107] 10…Action recognition system 10A...Action Recognition System 20... Behavior Recognition Learning Device 20A...Action Recognition Learning Device 21... Behavioral feature extraction unit 22...Text Feature Generation Unit 23…Feature Fusion Section 23A…Feature Fusion Section 24...Object Feature Extraction Unit 30…Action recognition estimation device 31…Procedure statement vector DB 32…Procedure determination unit 40… Computers 41… Processor 42…ROM 43…RAM 44…Auxiliary storage device 45… Input / Output Interface 46…Communication Interface 47... Bus 51…Input device 52…Output device 53… Input / Output Devices 54…Recording media

Claims

1. A behavioral feature extraction unit receives learning media information as input and extracts behavioral features from the learning media information, A text feature generation unit receives learning behavior information as input and generates text features from the learning behavior information, It has a procedure determination model, a feature fusion unit that receives the behavioral features, textual features and learning procedure information as input, and updates the procedure determination model based on the behavioral features, textual features and learning procedure information, Action recognition learning device.

2. The behavioral feature extraction unit extracts the behavioral features for each segment obtained by segmenting the learning media information. The behavior recognition learning device according to claim 1.

3. The text feature generation unit holds a number of pre-prepared templates, randomly selects a template from the learning behavior information, combines the selected template with the learning behavior information to generate prompt information, and generates the text features from the prompt information. The behavior recognition learning device according to claim 2.

4. The feature fusion unit generates semantic features for each segment based on the behavioral features and textual features of each segment, combines the behavioral features and semantic features of each segment to generate segment features, generates procedure sentence features from the segment features of each segment, calculates similarity between the procedure sentence features and textual features, calculates a loss based on the similarity determination, backpropagates the loss, and updates the parameters of the procedure determination model. The behavior recognition learning device according to claim 2.

5. The aforementioned learning media information is video data, and the segment is a video segment. The system further includes an object feature extraction unit that receives the aforementioned learning media information as input, extracts images from each video segment, performs object detection on those images, and extracts object features of the detected objects. The feature fusion unit receives the object features in addition to the behavioral features and text features as input, generates semantic features for each video segment based on the behavioral features, text features and object features, combines the behavioral features and semantic features of each video segment to generate segment features, generates procedure sentence features from the segment features of each video segment, calculates the similarity between the procedure sentence features and the text features, calculates a loss based on the similarity determination, backpropagates the loss, and updates the parameters of the procedure determination model. The behavior recognition learning device according to claim 2.

6. A procedure statement vector DB that holds multiple procedure statement vectors, The system comprises a procedure determination model updated based on learning media information, learning behavior information, and learning procedure information; an estimated procedure sentence vector output by the procedure determination model, which receives media information and behavior information as input and takes the media information and behavior information as input; and a procedure determination unit that calculates the similarity between each of the plurality of procedure sentence vectors held in the procedure sentence vector DB and the estimated procedure sentence vector, and outputs the procedure sentence of the procedure sentence vector with the highest similarity as procedure information. Behavior recognition estimation device.

7. The process involves receiving learning media information as input and extracting behavioral characteristics from the learning media information, The process involves receiving learning behavior information as input and generating textual features from the learning behavior information, The process includes receiving learning procedure information, the behavioral features, and the textual features as input, and updating a procedure determination model based on the learning procedure information, the behavioral features, and the textual features. A method of learning behavior recognition performed by computers.

8. A computer having a processor and memory, To perform the function of the behavior recognition learning device described in any one of claims 1 to 5, Action recognition program.

Citation Information

Patent Citations

  • System and method for attention-based configurable convolutional neural network (abc-CNN) for visual question answering

    JP2017091525A

  • Image description position determination method and device, electronic device, and storage medium

    JP2021509979A

  • Behavior recognition system and behavior recognition method

    WO2018159542A1