Image recognition device, image recognition method, image recognition program, and learning method

The image recognition apparatus enhances the accuracy of recognizing fine human movements and impressions by extracting features from divided image regions and utilizing a Transformer model, addressing the limitations of CNNs in spatial resolution loss.

WO2025141830A1PCT designated stage expired Publication Date: 2025-07-03NT T INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/047133
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Conventional image recognition technologies using CNNs struggle to accurately recognize fine human movements due to the reduction of spatial resolution and loss of fine spatial information.

Method used

An image recognition apparatus that extracts features from divided regions of the input image, including the face, head, and mouth, using a CNN, and identifies gestures or impressions using a Transformer model, optimizing model parameters through learning.

Benefits of technology

Accurately recognizes fine human movements and impressions by dividing the input image into regions, improving recognition accuracy compared to conventional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023047133_03072025_PF_FP_ABST
    Figure JP2023047133_03072025_PF_FP_ABST
Patent Text Reader

Abstract

An acquisition unit (11) has a feature amount extraction unit (113) and a label identification unit (115). The feature amount extraction unit (113) extracts a feature amount from each of a plurality of images obtained by dividing an input image obtained by imaging a person. The label identification unit (115) identifies, by using a model, the gesture of the person or the impression of the person in the input image on the basis of the feature amount.
Need to check novelty before this filing date? Find Prior Art

Description

Image recognition device, image recognition method, image recognition program, and learning method

[0001] The present invention relates to an image recognition device, an image recognition method, an image recognition program, and a learning method.

[0002] Conventionally, known techniques for recognizing human gestures based on images include a technique for estimating emotions based on facial expressions or body gestures from videos (see, for example, Non-Patent Documents 1 and 2), and a technique for recognizing explicit poses as a hand gesture recognition task (see, for example, Non-Patent Document 3).

[0003] Shan Li and Weihong Deng, “Deep Facial Expression Recognition: A survey,” IEEE Transactions on Affective Computing, vol. 13, no. 3, pp. 1195-1215, 2020. on Emotional Body Gesture Recognition,” IEEE Transactions on Affective Computing, vol. 12, no. 2, pp. 505-523, 2018.SHI Yuanyuan, LI Yunan, FU Xiaolong, MIAO Kaibin, and MIAO Qiguang, “Review of dynamic gesture recognition,” Virtual Reality & Intelligent Hardware, vol. 3, no. 3, pp. 183- 206, 2021.

[0004] However, conventional techniques have a problem in that they may not be able to accurately recognize information about a person based on an image.

[0005] Here, the information about a person recognized based on an image includes human gestures. To identify a person's gestures, it is considered useful to recognize not only the overall body movements contained in the image, but also the detailed movements occurring in specific parts of the body. Meanwhile, the CNN (Convolutional Neural Network) used for image recognition has a structure that reduces the spatial resolution of the image and removes detailed spatial information. Therefore, conventional CNN-based technologies can have difficulty recognizing detailed human movements.

[0006] In order to solve the above-mentioned problems and achieve the objective, the image recognition device is characterized by having a feature extraction unit that extracts features from each of multiple images obtained by dividing an input image of a person, and an identification unit that uses a model to identify the gestures or impression of the person in the input image based on the features.

[0007] According to the present invention, human gestures can be recognized with high accuracy based on images.

[0008] FIG. 1 is a diagram illustrating an example of the configuration of an image recognition device according to a first embodiment. FIG. 2 is a diagram illustrating an example of the configuration of a learning unit. FIG. 3 is a diagram illustrating an example of an image of a mouth region. FIG. 4 is a diagram illustrating an example of an image of a face region. FIG. 5 is a diagram illustrating an example of an image of a head region. FIG. 6 is a diagram illustrating an example of an image of an upper body region. FIG. 7 is a diagram illustrating the structure of an acquisition unit. FIG. 8 is a diagram illustrating an example of the configuration of a prediction unit. FIG. 9 is a flowchart illustrating the processing flow of the learning unit. FIG. 10 is a flowchart illustrating the processing flow of acquisition processing. FIG. 11 is a flowchart illustrating the processing flow of the prediction unit. FIG. 12 is a diagram illustrating a dataset used in an experiment. FIG. 13 is a diagram illustrating the results of the experiment. FIG. 14 is a diagram illustrating the results of the experiment. FIG. 15 is a diagram illustrating the results of the experiment. FIG. 16 is a diagram illustrating an example of a computer that executes an image recognition program.

[0009] Hereinafter, embodiments of an image recognition device, an image recognition method, an image recognition program, and a learning method according to the present application will be described in detail with reference to the accompanying drawings. Note that the present invention is not limited to the embodiments described below.

[0010] [Configuration of First Embodiment] Fig. 1 is a diagram showing an example of the configuration of an image recognition device according to a first embodiment. The image recognition device 1 shown in Fig. 1 identifies a gesture of a person appearing in an input image. For example, the image recognition device 1 outputs a label corresponding to the identified gesture.

[0011] Here, there are two approaches to human communication: an approach based on linguistic information such as the content of conversation, and an approach based on non-linguistic information such as visual information. Non-linguistic information is, for example, the movement of a person.

[0012] For example, it is known that the impression one receives during communication changes depending on non-verbal information such as gestures and facial movements (Reference 1: Albert Mehrabian, Silent Messages, vol. 8, Wadsworth, Belmont, CA, 1971).

[0013] For example, by checking the user's own gestures recognized by the image recognition device 1, the user can review the gestures used in communication so as to improve the impression given to the other person. Note that the "gestures" in the embodiment are an example of movements that are non-verbal information. Furthermore, the "gestures" may be rephrased as "gestures," "manner," "behavior," etc.

[0014] 1, the image recognition device 1 includes a learning unit 10 and a prediction unit 30. The image recognition device 1 also stores model parameters 20. Note that the learning unit 10 and the prediction unit 30 may be realized by different devices.

[0015] The learning unit 10 accepts input of a training gesture dataset. The training gesture dataset is a collection of video clips associated with teacher labels. A video clip is video data organized into predetermined units, and may be simply referred to as a video. A video clip may be video data shot for a predetermined period of time, or video data extracted from an entire video. Details of the video clips and labels will be described later. The learning unit 10 uses the training gesture dataset to learn a model used for image recognition, which is constructed based on model parameters 20. The learning unit 10 updates the model parameters 20. The model may also be referred to as a trained model, a machine learning model, or the like.

[0016] The prediction unit 30 receives an input of a gesture data set for prediction. The gesture data set for prediction is a collection of one or more video clips. For example, the prediction unit 30 processes the video clips one by one. The prediction unit 30 predicts a gesture from the gesture data set for prediction using a model constructed based on the updated model parameters 20. The prediction unit 30 outputs a gesture label (hereinafter simply referred to as a label) representing the gesture.

[0017] The configuration of the learning unit 10 will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of the configuration of the learning unit. As shown in Fig. 2, the learning unit 10 includes an acquisition unit 11 and an update unit 12. It is also assumed here that the acquisition unit 11 includes model parameters 20.

[0018] The acquisition unit 11 acquires labels based on the input data set. The update unit 12 updates the model parameters 20.

[0019] The acquisition unit 11 includes an image sequence extraction unit 111, an area division unit 112, a feature extraction unit 113a, a feature extraction unit 113b, a feature extraction unit 113c, a feature extraction unit 113d, a feature connection unit 114, and a label identification unit 115. In the following description, the feature extraction units 113a, 113b, 113c, and 113d may be collectively referred to as the feature extraction unit 113.

[0020] The image sequence extraction unit 111 extracts an image sequence from the input data set. For example, the training gesture data set is K video clips and is expressed as in equation (1) (where K is an integer equal to or greater than 1).

[0021]

[0022] vd k is the k-th video clip (where k is an integer from 1 to K). k is the k-th label. Note that a video is an example of a sequence of images that are continuous in time.

[0023] The labels correspond to gestures, including face-hand gestures (Face-Hand), head-hand gestures (Head-Hand), face gestures (Face), and a neutral state.

[0024] For example, gestures related to the face and hands include "covering mouth with hand," "touching ear with hand," "touching nose with hand," "rubbing eyes with hand," "touching eyebrow with hand," "touching cheek with hand," "touching chin with hand," and "resting chin on hand." Gestures related to the head and hands include "scratching head with hand," "stroking hair with hand," "grabbing hair with hand," "touching back of head with hand," and "ruffling hair with hand." Gestures related to the face include "licking lips," "biting lip," "sticking out tongue," "yawning," "sniffing nose," and "puffing cheeks."

[0025] In this embodiment, there are 20 labels, including the 19 gestures listed here and a state where no gesture is being performed. However, the content and number of gestures are not limited to those described here. Note that gestures associated with labels include gestures that people make unconsciously.

[0026] For example, video clips of the training gesture dataset are obtained by filming a subject being instructed to perform a gesture. In this case, the label of each video clip is known. For example, a video clip labeled "cover mouth with hand" is obtained by filming a subject being instructed to "cover mouth with hand."

[0027] The image sequence extraction unit 111 extracts an image sequence V from the C-th video clip included in the data set, as shown in equation (2). C (where C is an integer between 1 and K). The image sequence extraction unit 111 can extract an image sequence using a program prepared as a library.

[0028]

[0029] Image sequence V C The image sequence V contains N images (still images) as elements. C Each element of v may be a predetermined frame included in a video clip. 1 , …, v N , means that the elements are arranged in chronological order (where N is an integer equal to or greater than 1). That is, in a video clip, 2 , v 1 The image is a frame that comes later in time than the image in the previous frame.

[0030] The region dividing unit 112 divides an input image of a person into a plurality of regions each including a part of the person's face. The region dividing unit 112 can divide the image using a program prepared as a library.

[0031] The region dividing unit 112 divides the image sequence V C V shown in equation (3) C i Here, i=1, 2, 3, 4. That is, the region dividing unit 112 divides the image sequence V C V C 1 , V C 2 , V C 3 , V C 4 The image is divided into four image sequences.

[0032]

[0033] V C 1 is the upper body region image series data. C 2 is the face region image series data.C 3 is the head region image series data. C 4 is the mouth region image series data. In this way, the region dividing unit 112 divides the input image into a plurality of regions including the face region, head region, and mouth region of a person.

[0034] Fig. 3 is a diagram showing an example of an image of a mouth region. The area surrounded by a dashed line in Fig. 3 is an example of a mouth region. For example, the image of the mouth region shows a detailed action of "touching the cheek," which is one of the gestures related to the hands and face.

[0035] 4 is a diagram showing an example of an image of a face region. The area surrounded by a dashed line in Fig. 4 is an example of a face region. For example, the image of the face region shows a detailed action such as "licking lips," which is one of facial gestures.

[0036] Fig. 5 is a diagram showing an example of an image of the head region. The area surrounded by a dashed line in Fig. 5 is an example of the head region. For example, the image of the head region shows a detailed movement such as "touching the back of the head," which is one of the gestures related to the head and hands.

[0037] Fig. 6 is a diagram showing an example of an image of the upper body region. The region surrounded by a dashed line in Fig. 6 is an example of the upper body region. For example, from the image of the upper body region, it can be seen that the person is not making any gesture (Neutral).

[0038] The regions divided by the region dividing unit 112 are not limited to those described here. The lower half of the face may be used instead of the mouth region. Also, the upper half of the face may be used instead of the head region.

[0039] Returning to FIG. 2 , the feature extraction unit 113 extracts features from each image divided by the region division unit 112. The feature extraction unit 113 inputs image sequence data to a CNN corresponding to each region to acquire features. The CNN is constructed based on model parameters 20. The model parameters 20 are CNN parameters, such as θ 1 , θ 2 , θ 3 , θ 4 Includes.

[0040] The feature extraction unit 113a extracts the parameter θ 1 The upper body image sequence data V C 1 and the feature F 1 The feature extraction unit 113b obtains the parameter θ 2 The face region image sequence data V C 2 and the feature F 2 The feature extraction unit 113c obtains the parameter θ 3 The head region image sequence data V C 3 and the feature F 3 The feature extraction unit 113d obtains the parameter θ 4 The mouth region image sequence data V C 4 and the feature F 4 get.

[0041] The feature extraction unit 113 also edits the images to unify their formats before inputting them to the CNN. For example, the feature extraction unit 113 resizes the images to 128×128 pixels and converts the color space to three RGB channels.

[0042] The feature amount connection unit 114 concatenates the feature amounts extracted by the feature amount extraction unit 113. For example, the feature amount connection unit 114 concatenates each feature amount that is a vector.

[0043] The label identification unit 115 uses a model that captures time-series changes in the input data to identify the linked feature F CA In other words, the label identification unit 115 identifies the gesture of the person in the input image using a time-series processing algorithm. A model that captures time-series changes in input data is, for example, the Transformer described in References 2 or 3.

[0044] Reference 2: Zehua Sun, Qiuhong Ke, Hossein Rahmani, Mohammed Bennamoun, Gang Wang, and Jun Liu, “Human Action Recognition From Various Data Modalities: A review,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 3, pp. 3200-3225, 2023

[0045] Reference 3: Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, Zhaohui Yang, Yiman Zhang, and Dacheng Tao, “A survey on vision transformer,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 1, pp. 87-110, 2022

[0046] The Transformer is constructed based on model parameters 20. The model parameters 20 are parameters of the Transformer, such as θ 5 Includes.

[0047] The Transformer outputs a score for each of the 20 labels. The label identification unit 115 identifies the label with the highest score as the label representing the gesture of the person in the input image.

[0048] The update unit 12 updates the model parameters so as to optimize the gesture identified by the label identification unit 115. For example, the update unit 12 updates the model parameters 20 so as to reduce the difference between the label identified for each image sequence in equation (3) by the label identification unit 115 and the teacher label shown in equation (4). For example, the update unit 12 updates the parameters of the CNN and the Transformer using the backpropagation algorithm.

[0049]

[0050] Here, the Transformer outputs a vector in which each element corresponds to a gesture, the value of each element (corresponding to a score) is non-negative, and the sum of the values ​​of the elements is 1. In addition, in the training gesture dataset, the teacher label associated with a video clip is assumed to have been converted into a vector in which the value of the element corresponding to the correct gesture is 1 and the values ​​of the other elements are 0, as shown in equation (4). For example, even if the teacher label is given as a string such as "cover mouth with hand," it is converted into a vector when input to the image recognition device 1.

[0051] The update unit 12 calculates the error (e.g., cross entropy) between the vector output by the Transformer and the vector of the teacher label. The update unit 12 changes the parameters of the CNN and the Transformer by an amount corresponding to a preset learning rate so as to reduce the calculated error.

[0052] A model of the acquisition unit 11 is shown in FIG. 7. FIG. 7 is a diagram illustrating the structure of the acquisition unit. As shown in FIG. 7, the acquisition unit 11 divides an image into an upper body region (Base), a face region (Face), a head region (Head), and a mouth region (Mouth), and inputs the results to a CNN corresponding to each region. The acquisition unit 11 concatenates the features output from the CNN and inputs the results to a Transformer. The acquisition unit 11 further performs Attentive pooling and normalization using Softmax layers on the output from the Transformer, and finally identifies a label.

[0053] The region dividing unit 112 can divide the image by trimming it. For example, the region dividing unit 112 can obtain an image of a face region by deleting from the image all regions other than the rectangular region containing the face.

[0054] Furthermore, the images obtained by division may have overlapping regions. For example, in the example of Fig. 7, the image of the face region partially overlaps with the image of the head region and the image of the mouth region. Note that the images obtained by division may include the input image itself.

[0055] Here, the upper body region image in Fig. 7 is referred to as a first image. On the other hand, the face region image, the head region image and the mouth region image are referred to as second images. In the example of Fig. 7, the first image includes a plurality of second images. The acquisition unit 11 inputs the first image together with one or more second images into the model.

[0056] That is, the feature extraction unit 113 extracts features from each of a first image obtained by dividing an input image of a person and one or more second images included in the first image, which are also obtained by dividing the input image. This allows the image recognition device 1 to recognize the movement of the entire body from the first image, while also being able to recognize finer movements of the head, face, mouth, etc., thereby improving accuracy.

[0057] The first image may include a plurality of second images, or may overlap at least partially with each of the plurality of second images. The first image may also be the input image itself. The first image may also be an image of the largest area among the images obtained by the area division unit 112. For example, the area division unit 112 obtains the image of the largest area by trimming an area that does not include a human body from the input image.

[0058] The configuration of the prediction unit 30 will be described with reference to Fig. 8. Fig. 8 is a diagram showing an example of the configuration of the prediction unit. As shown in Fig. 8, the prediction unit 30 includes an acquisition unit 11.

[0059] The configuration of the acquisition unit 11 is as described with reference to Fig. 2 etc. However, unlike the case of Fig. 2, the acquisition unit 11 here identifies a label from the gesture data set for prediction.

[0060] For example, the prediction gesture data set is K' video clips and is expressed as in equation (5).

[0061]

[0062] In this case, the image sequence V′ extracted from the Cth video clip included in the prediction gesture dataset is C is as shown in equation (6).

[0063]

[0064] [Processing of First Embodiment] The processing flow of the learning unit 10 will be described with reference to Fig. 9. Fig. 9 is a flowchart showing the processing flow of the learning unit.

[0065] 9, the learning unit 10 first receives an input of a training gesture data set (step S11). Next, the learning unit 10 executes an acquisition process to identify labels (step S12). In the acquisition process, labels are identified using a model constructed based on model parameters.

[0066] Next, the learning unit 10 updates the model parameters based on the identified labels (step S13). If the termination condition is not satisfied (step S14, No), the learning unit 10 returns to step S12 and repeats the process. On the other hand, if the termination condition is satisfied (step S14, Yes), the learning unit 10 terminates the process. The termination condition is that the parameter update has been repeated a certain number of times, the parameter update amount has converged, etc.

[0067] The flow of the acquisition process (step S12 in FIG. 9) will be described using FIG. 10. FIG. 10 is a flowchart showing the flow of the acquisition process. The acquisition process is performed by the acquisition unit 11 in accordance with a request from the learning unit 10 or the prediction unit 30. A training gesture data set or a prediction gesture data set is input to the acquisition process.

[0068] 10, first, the acquisition unit 11 acquires an image sequence from the input data set (step S21), and then divides each image included in the image sequence into images of the upper body region, face region, head region, and mouth region (step S22).

[0069] Next, the acquiring unit 11 extracts features of each divided image using a model (e.g., CNN) constructed from the model parameters (step S23), and then connects the extracted features (step S24).

[0070] The acquisition unit 11 then performs time-series processing of the linked features using a model (e.g., Transformer) constructed from the model parameters (step S25), and identifies labels based on the results of the time-series processing (step S26).

[0071] The flow of processing by the prediction unit 30 will be described with reference to Fig. 11. Fig. 11 is a flowchart showing the flow of processing by the prediction unit.

[0072] 11 , first, the prediction unit 30 receives an input of a gesture data set for prediction (step S31). Next, the prediction unit 30 executes an acquisition process to identify labels (step S32). Details of the acquisition process are as described in FIG. 10. Then, the prediction unit 30 outputs the identified labels (step S33).

[0073] Here, the target predicted by the image recognition device 1 is not limited to the movement of a person appearing in an image. Depending on the configuration of the model (e.g., the type and number of labels to be output) and the configuration of the training gesture dataset, the image recognition device 1 can predict any target.

[0074] For example, the image recognition device 1 can predict the impression that a person in an image gives to a communication partner. In this case, a label indicating an impression is associated with each video clip in the training gesture dataset. The label may be, for example, a binary value indicating whether the impression is good or bad, a score indicating whether the impression is good or bad, or a value indicating the type of impression. The type of impression may be expressed in natural language, such as "energetic," "gloomy," "intelligent," "like," or "dislike." Furthermore, information related to impressions, such as "good impression," "bad impression," and "no effect on impression," may be associated in advance with the aforementioned 20 labels. Furthermore, depending on how the labels are set, the image recognition device 1 can predict non-verbal information, including attitudes, emotions, intentions, and the like, in addition to impressions.

[0075] As described above, the feature extraction unit 113 extracts features from each of a plurality of images obtained by dividing an input image of a person. The label identification unit 115 uses a model to identify the gestures or impressions of the person in the input image based on the features.

[0076] The image recognition device 1 of this embodiment divides an input image into facial regions, thereby enabling recognition of minute movements occurring in parts of the face, etc. As a result, according to this embodiment, it is possible to accurately recognize human gestures or impressions based on an image.

[0077] The feature extraction unit 113 also extracts features from each of the images of the person's face region, head region, and mouth region obtained by dividing the input image. This makes it easier to recognize detailed movements in the head, face, mouth, etc. regions compared to when gestures are recognized using only an image of the whole body or upper body.

[0078] The feature extraction unit 113 extracts features using a CNN corresponding to each of the multiple images. The label identification unit 115 uses a Transformer to identify a person's gestures or an impression of the person based on features obtained by combining the features corresponding to each of the multiple images extracted by the feature extraction unit 113. The region division unit 112 divides an input image, which is a photograph of a person and is included in a series of temporally consecutive images, into multiple regions including a partial region of the person's face. In this way, the image recognition device 1 can accurately recognize gestures or an impression of a person by capturing temporal changes using a series of temporally consecutive images (e.g., a video).

[0079] The update unit 12 updates the parameters of the model so as to optimize the gesture identified by the label identification unit 115. This allows the image recognition device 1 to learn the model.

[0080] The image recognition device 1 according to the first embodiment provides specific improvements over conventional techniques for recognizing gestures from images, such as those described in non-patent document 3, and represents an advancement in the technical field of recognizing human gestures from images.

[0081] [Experiment] Here, an experiment conducted by the inventor to confirm the effects of the embodiment will be described. The experiment mainly compared the gesture recognition accuracy between the first embodiment (segmenting the image and using CNN and Transformer) and a baseline (using CNN and Transformer without segmenting the image). The baseline is described in Non-Patent Document 1. The accuracy is the accuracy rate of the identified labels.

[0082] The experimental conditions were as follows: Training gesture dataset size (number of video clips): 23,496 Prediction gesture dataset size (number of video clips): 6,228 Input image size: 128 x 128 (after resizing) Learning rate: 0.0001 Pre-training model: 128 x 128 ImageNet (Reference 5) trained with Efficientnet (Reference 4) Image segmentation method: Upper body, face, head, and mouth regions were trimmed (segmented) using dlib (Reference 6) Batch size: 4 Optimizer: Radam (Reference 7) Training termination condition: EarlyStopping (stop if accuracy does not improve after 5 epochs)

[0083] Reference 4: Mingxing Tan and Quoc Le, “Efficientnet: Rethinking Model Scaling for Convolutional Neural Networks,” Proceedings of the 36th International Conference on Machine Learning, pp. 6105-6114, 2019

[0084] Reference 5: Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei, “ImageNet: A large-scale hierarchical image database,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 248-255, 2009.

[0085] Reference 6: Davis E King, “Dlib-ml: A Machine Learning Toolkit,” The Journal of Machine Learning Research, vol. 10, pp. 1755- 1758, 2009.

[0086] Reference 7: Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han, “On the Variance of the Adaptive Learning Rate and Beyond,” Proceedings of the 8th International Conference on Learning Representations, 2020.

[0087] Fig. 12 is a diagram illustrating the dataset used in the experiment. The categories include gestures related to the face and hands (Face-Hand), gestures related to the head and hands (Head-Hand), gestures related to the face (Face), and a state in which no gesture is being made (Neutral).

[0088] The number of categories is the number of gestures included in each category. There are eight gestures related to the face and hands: "covering mouth with hand," "touching ear with hand," "touching nose with hand," "rubbing eyes with hand," "touching eyebrows with hand," "touching cheek with hand," "touching chin with hand," and "resting chin on hand." There are five gestures related to the head and hands: "scratching head with hand," "stroking hair with hand," "grabbing hair with hand," "touching back of head with hand," and "ruffling hair with hand." There are six gestures related to the face: "licking lips," "biting lip," "sticking out tongue," "yawning," "sniffing nose," and "puffing cheeks."

[0089] The number of video clips is the number of video clips for each category. For example, the number of videos in the training gesture dataset for face-hand gestures is 8,199. Also, for example, the number of videos in the prediction gesture dataset for face-hand gestures is 2,182.

[0090] 13, 14, and 15 are diagrams showing the results of the experiment. As shown in Fig. 13, the accuracy of the embodiment is improved by 2.2% compared to the baseline. Note that Previous is a method in which the baseline Transformer is replaced with LSTM (Long Short Term Memory), and its accuracy is lower than the baseline.

[0091] In this embodiment, all of the four divided images are basically used (Base, Face, Head, and Mouth are all checked). However, each figure shows the results of a pattern in which all of the four divided images are used, as well as a pattern in which only part of the four divided images are used (Base, Face, Head, and Mouth are checked).

[0092] As shown in FIG. 14, the present embodiment exhibited higher accuracy than the baseline, particularly for Face-Hand, Face, and Neutral, among the four categories.

[0093] 15, this embodiment achieved particularly high recognition accuracy for gestures such as "licking lips," "sticking out tongue," and "sniffing." It is believed that the input image series of head region images and mouth region images contributes to improved accuracy when recognizing gestures such as these that involve moving the face vertically and horizontally.

[0094] [System Configuration, etc.] The components of each device shown in the figure are conceptual functional units and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of the devices can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc. Furthermore, all or any part of the processing functions performed by each device can be realized by a CPU (Central Processing Unit) and a program analyzed and executed by the CPU, or can be realized as hardware using wired logic. The program may be executed not only by the CPU but also by other processors such as a GPU.

[0095] Furthermore, among the processes described in this embodiment, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method.In addition, the information including the processing procedures, control procedures, specific names, various data and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified.

[0096] [Program] In one embodiment, the image recognition device 1 can be implemented by installing an image recognition program that executes the above-described processes as package software or online software on a desired computer. For example, by executing the image recognition program on an information processing device, the information processing device can function as the image recognition device 1. The information processing device referred to here includes desktop and notebook personal computers. In addition, the information processing device also includes mobile communication terminals such as smartphones, mobile phones, and PHS (Personal Handyphone Systems), as well as slate terminals such as PDAs (Personal Digital Assistants).

[0097] The image recognition device 1 may also be implemented as a server device that provides a service related to the above-described processing to a client terminal device used by a user. For example, the server device may be implemented as a server device that receives history information as an input and outputs candidates for users who are candidates for intervention. In this case, the server device may be implemented as a web server or as a cloud that provides a service related to the above-described processing by outsourcing.

[0098] 16 is a diagram showing an example of a computer that executes an image recognition program. The computer 1000 includes, for example, a memory 1010 and a CPU 1020. The computer 1000 also includes a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0099] The memory 1010 includes a read-only memory (ROM) 1011 and a random access memory (RAM) 1012. The ROM 1011 stores a boot program such as a basic input / output system (BIOS). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.

[0100] The hard disk drive 1090 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. That is, a program that defines each process of the image recognition device 1 is implemented as a program module 1093 in which computer-executable code is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing processes similar to those of the functional configuration of the image recognition device 1 is stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).

[0101] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in the memory 1010 or the hard disk drive 1090. Then, the CPU 1020 reads the program module 1093 or the program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as necessary, and executes the processing of the above-described embodiment.

[0102] The program module 1093 and program data 1094 may not necessarily be stored in the hard disk drive 1090, but may also be stored in a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or a wide area network (WAN)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.

[0103] The following additional notes are provided regarding the above-described embodiments.

[0104] (Supplementary Item 1) An image recognition device including: a memory; and at least one processor connected to the memory, wherein the processor extracts features from each of a plurality of images obtained by dividing an input image of a person; and identifies a gesture or an impression of the person in the input image based on the features using a model.

[0105] (Supplementary Item 2) The image recognition device according to Supplementary Item 1, wherein the processor extracts features from each of an image of a face region of the person, an image of a head region of the person, and an image of a mouth region of the person obtained by dividing the input image.

[0106] (Supplementary Item 3) The image recognition device according to Supplementary Item 1 or 2, wherein the processor extracts the feature amounts using a CNN corresponding to each of the plurality of images, and identifies the person's gestures or the impression of the person using a Transformer based on a feature amount obtained by combining the extracted feature amounts corresponding to each of the plurality of images.

[0107] (Supplementary Item 4) The image recognition device according to any one of Supplementary Items 1 to 3, wherein the processor extracts features from each of a first image obtained by dividing the input image and one or more second images obtained by dividing the input image and included in the first image.

[0108] (Supplementary Item 5) A non-transitory storage medium readable by a computer, storing a program for causing a computer to execute image recognition processing by the image recognition device according to Supplementary Item 1.

[0109] (Supplementary Item 6) A learning method executed by a computer, comprising: a feature extraction step of extracting features from each of a plurality of images obtained by dividing an input image of a person; an identification step of identifying information about the person in the input image based on the features using a model; and an update step of updating parameters of the model so that the identification result obtained by the identification step is optimized.

[0110] REFERENCE SIGNS LIST 1 Image recognition device 10 Learning unit 11 Acquisition unit 12 Update unit 20 Model parameters 30 Prediction unit 111 Image sequence extraction unit 112 Region division unit 113a, 113b, 113c, 113d Feature extraction unit 114 Feature connection unit 115 Label identification unit

Claims

1. A feature quantity extraction unit that extracts feature quantities from each of a plurality of images obtained by dividing an input image of a person, and a specification unit that uses a model to specify the gesture or impression of the person in the input image based on the feature quantities, wherein the image recognition apparatus is characterized by having the above components.

2. The feature quantity extraction unit according to claim 1, wherein the feature quantity extraction unit extracts feature quantities from each of an image of the face region of the person, an image of the head region of the person, and an image of the mouth region of the person, which are obtained by dividing the input image.

3. The feature quantity extraction unit according to claim 1 or 2, wherein the feature quantity extraction unit extracts the feature quantities using a CNN corresponding to each of the plurality of images, and the specification unit uses a Transformer to specify the gesture or impression of the person based on the combined feature quantities of the feature quantities corresponding to each of the plurality of images extracted by the feature quantity extraction unit.

4. The feature quantity extraction unit according to claim 1, wherein the feature quantity extraction unit extracts feature quantities from each of a first image obtained by dividing the input image and an image obtained by dividing the input image and including one or more second images included in the first image.

5. A feature quantity extraction unit that extracts feature quantities from each of a plurality of images obtained by dividing an input image of a person, a specification unit that uses a model to specify information about the person in the input image based on the feature quantities, and an update unit that updates the parameters of the model so that the specification result by the specification unit is optimized, wherein the image recognition apparatus is characterized by having the above components.

6. An image recognition method executed by an image recognition apparatus, the method including a feature quantity extraction step of extracting feature quantities from each of a plurality of images obtained by dividing an input image of a person, and a specification step of using a model to specify the gesture or impression of the person in the input image based on the feature quantities.

7. An image recognition program, characterized in that it causes a computer to execute: a feature amount extraction step of extracting a feature amount from each of a plurality of images obtained by dividing an input image of a person; and an identification step of identifying a gesture or an impression of the person in the input image based on the feature amount using a model.

8. A learning method executed by a computer, comprising: a feature amount extraction step of extracting a feature amount from each of a plurality of images obtained by dividing an input image of a person; an identification step of identifying information about the person in the input image based on the feature amount using a model; and an update step of updating parameters of the model so that the identification result by the identification step is optimized.

Citation Information

Patent Citations

  • Plasticized object feature detection device

    JP1996263623A