Collection system and collection method

The collection system efficiently assigns labels to moving images by extracting and altering facial movements, addressing privacy issues in crowdsourced video labeling.

JP7768243B2Active Publication Date: 2025-11-12NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023562026
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-18
Publication Date
2025-11-12
Estimated Expiration
2041-11-18

AI Technical Summary

Technical Problem

Conventional techniques face challenges in efficiently assigning labels to moving images due to issues with personal information protection and psychological barriers when making videos public for crowdsourcing.

Method used

A collection system comprising an extraction device to extract head and eyeball movements from videos, a collection device to create representative facial pose videos, and an evaluation device to assign labels efficiently while protecting personal information by altering the face in the video to maintain privacy.

Benefits of technology

Enables efficient labeling of moving images by allowing evaluators to intuitively recognize psychological states without revealing the identity of the subject, thus overcoming privacy concerns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007768243000001
    Figure 0007768243000001
  • Figure 0007768243000002
    Figure 0007768243000002
  • Figure 0007768243000003
    Figure 0007768243000003
Patent Text Reader

Abstract

A collection device (30) creates a video image representing a facial posture on the basis of time-series data that is extracted from a captured video image of a subject, and that indicates the movement of a user's head and eyeballs. The collection device (30) collects a label assigned to the created video image. It is not desirable to disclose the captured video image to an unspecified number of users. In contrast, the created video image is created on the basis of the time-series data, and is a video image that poses no problem even when disclosed to an unspecified number of users.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention Collection stem and collection method Regarding. [Background technology]

[0002] There are known machine learning models that estimate the psychological state of a user (hereinafter referred to as the subject) from video images of the user's face. Furthermore, training such machine learning models requires training data that combines video images with labels indicating the psychological state.

[0003] Here, the learning data is obtained by manually labeling videos. Crowdsourcing is known as a technique for making the manual labeling process more efficient (see, for example, Non-Patent Document 1).

[0004] For example, in crowdsourcing, an unspecified number of users (hereinafter referred to as raters) actually watch videos and assign labels to them on a platform accessible via the Internet. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Hisashi Kashima, Akira Kajino, "Crowdsourcing and Machine Learning (Special Issue: Knowledge Transfer)," Journal of the Japanese Society for Artificial Intelligence, Vol. 27, No. 4 (July 2012), pp. 381-388 (https: / / www.jstage.jst.go.jp / article / jjsai / 27 / 4 / 27_381 / _article / -char / ja / ) Summary of the Invention [Problem to be solved by the invention]

[0006] However, the conventional techniques have a problem in that labels cannot be efficiently assigned to moving images in some cases.

[0007] For example, when labeling using conventional crowdsourcing, video images of the subject's face must be made public to an unspecified number of raters.

[0008] On the other hand, it may become difficult to make the videos public from the perspective of protecting the personal information of the photographed person. Even if the issue of personal information protection is resolved legally, psychological barriers on the part of the photographed person may affect their decision to make the videos public. [Means for solving the problem]

[0009] In order to solve the above problems and achieve the objectives, The collection system is a collection system having an extraction device, a collection device, a providing device, and an evaluation device, wherein the extraction device has an acquisition unit that acquires a first video of a user, and an extraction unit that extracts data indicating the movement of the user's head and eyeballs from the first video, the collection device has a creation unit that creates a second video representing facial poses based on the data extracted by the extraction device, and a collection unit that collects labels assigned to the second video in the evaluation device via the providing device, the providing device has a provision unit that provides the second video created by the collection device to the evaluation device, and the evaluation device has a display control unit that causes an output device to output a screen including the second video and a message regarding changes in facial pose due to changes in concentration level, and an assignment unit that assigns a label specified by the user to the second video. [Effects of the Invention]

[0010] According to the present invention, labels can be efficiently assigned to moving images. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a diagram illustrating an overview of a collection system according to a first embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of the configuration of the extraction device according to the first embodiment. [Figure 3] FIG. 3 is a diagram showing an example of data indicating head and eye movements. [Figure 4] FIG. 4 is a diagram illustrating a method for extracting data indicating head and eye movements. [Figure 5] FIG. 5 is a diagram illustrating the coverage rate of the eyeball by the eyelid. [Figure 6] FIG. 6 is a diagram illustrating a method for extracting data indicating head and eye movements. [Figure 7] FIG. 7 is a diagram illustrating a method for extracting data indicating head and eye movements. [Figure 8] FIG. 8 is a diagram illustrating an example of the configuration of the collection device according to the first embodiment. [Figure 9] FIG. 9 is a diagram illustrating an example of the configuration of a providing device according to the first embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of the configuration of the rating device according to the first embodiment. [Figure 11] FIG. 11 is a diagram showing an example of a screen displayed on the rating device. [Figure 12] FIG. 12 is a flowchart showing the flow of processing in the collection system according to the first embodiment. [Figure 13] FIG. 13 is a diagram illustrating an example of the configuration of a collection device according to another embodiment. [Figure 14] FIG. 14 is a diagram illustrating an example of a computer that executes a collection program. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments of a collection device, a collection method, a collection program, and a collection system according to the present application will be described in detail with reference to the accompanying drawings. However, the present invention is not limited to the embodiments described below.

[0013] [Configuration of the first embodiment] 1 is a diagram illustrating an overview of a collection system according to the first embodiment. The collection system 1 is a system for manually assigning labels to videos used as training data for a machine learning model.

[0014] 1, the collection system 1 includes an extraction device 10, a photographing device 20, a collection device 30, a providing device 40, and an evaluation device 50. The extraction device 10, the photographing device 20, the collection device 30, the providing device 40, and the evaluation device 50 are capable of communicating with each other via a network (e.g., the Internet).

[0015] A user who uses the extraction device 10 is called a person to be photographed. The person to be photographed has a moving image captured by the photographing device 20.

[0016] The extraction device 10 is, for example, a personal computer, a smartphone, an ECU (Electronic Control Unit) mounted on an automobile, a drive recorder, etc. The image capturing device 20 is, for example, a camera.

[0017] The image capturing device 20 may be integrated with the extraction device 10 or may be separate from the extraction device 10. The image capturing device 20 is provided in a position that allows it to capture an image of the head of the person being photographed from the front. For example, the image capturing device 20 is installed in a position that faces the face of the person being photographed.

[0018] For example, if the extraction device 10 is a portable terminal device such as a laptop PC or a smartphone, the image capturing device 20 may be a camera provided in the terminal device.

[0019] Furthermore, when the extraction device 10 is an in-vehicle device such as an ECU or a drive recorder, the image capturing device 20 may be a camera provided at a position inside the vehicle that can capture an image of the driver.

[0020] A user who uses the collection device 30 is called an analyst. The analyst collects labels assigned to videos via the collection device 30. The analyst can use the collected combinations of labels and videos as training data to train a machine learning model.

[0021] The collection device 30 is, for example, a personal computer, a smartphone, or the like.

[0022] The providing device 40 is a device that provides a platform and data to an unspecified number of users (raters) and allows the users to perform predetermined tasks.

[0023] For example, the providing device 40 is a server connected to the rating device 50 via the Internet.

[0024] A user who uses the rating device 50 is called an evaluator. The evaluator performs the work in crowdsourcing. In this embodiment, the evaluator watches a video and then assigns a label to the video.

[0025] For example, the label may be related to the psychological state of a person appearing in a video, and may represent the degree of concentration (level of concentration) of the person on a task (desk work, driving a car), emotions, etc.

[0026] The evaluation device 50 is, for example, a personal computer, a smartphone, or the like.

[0027] The processing flow of the collection system 1 will be described with reference to Fig. 1. First, as shown in Fig. 1, the photographing device 20 photographs a person to be photographed (step S1).

[0028] Next, the extraction device 10 extracts time-series data from the images (moving images) captured by the image capturing device 20 (step S2). The time-series data is data that indicates the movements of the head and eyeballs of the person being photographed.

[0029] Then, the extraction device 10 transmits the extracted time series data to the collection device 30 (step S3). Note that the extraction device 10 may extract and transmit the time series data in response to an instruction from the collection device 30.

[0030] The collection device 30 then creates a video from the time-series data (step S4). The video created by the collection device 30 represents the facial features of a person. The facial features include head and eye movements, facial expressions, skin color and texture, etc.

[0031] The collecting device 30 transmits the created video to the providing device 40 (step S5). The providing device 40 transmits the video received from the collecting device 30 to the rating device 50 (step S6).

[0032] The rating device 50 assigns a label to the video in accordance with the designation of the rater who viewed the video received from the providing device 40 (step S7). Then, the rating device 50 transmits the assigned label to the providing device 40.

[0033] For example, the rating device 50 can realize steps S7 and S8 by associating the label specified by the rater with information (ID, file name, etc.) that identifies the video to which the label has been assigned and notifying the providing device 40 of the association.

[0034] Here, the face appearing in the video created by the collection device 30 and displayed by the evaluation device 50 is different from the face of the person being photographed. In other words, the evaluator can watch the video without recognizing that the person appearing in the video is the person being photographed.

[0035] The providing device 40 transmits the labels received from the rating device 50 to the collecting device 30 (step S9). This allows the collecting device 30 to collect the labels assigned to the videos.

[0036] Here, the configurations and processes of the extraction device 10, the collection device 30, the provision device 40, and the evaluation device 50 will be described.

[0037] 2 is a diagram illustrating an example of the configuration of an extraction device according to the first embodiment. As illustrated in FIG. 2, the extraction device 10 includes a communication unit 11, an input unit 12, an output unit 13, a storage unit 14, and a control unit 15.

[0038] The communication unit 11 performs data communication with other devices via a network. For example, the communication unit 11 is a network interface card (NIC).

[0039] The input unit 12 is an interface for receiving input of data. For example, the input unit 12 is connected to input devices such as a mouse and a keyboard. The input unit 12 is also connected to the image capturing device 20 and receives input of images (moving images) captured by the image capturing device 20.

[0040] The output unit 13 is an interface for outputting data and is connected to output devices such as a display and a speaker.

[0041] The storage unit 14 is a storage device such as a hard disk drive (HDD), a solid state drive (SSD), an optical disk, etc. The storage unit 14 may also be a data-rewritable semiconductor memory such as a random access memory (RAM), a flash memory, or a non-volatile static random access memory (NVSRAM).

[0042] The storage unit 14 stores an OS (Operating System) and various programs executed by the extraction device 10. The storage unit 14 stores image data 141 and extraction model information 142.

[0043] The image data 141 is data of a moving image captured by the image capturing device 20. The extraction model information 142 is information such as parameters for constructing a model for extracting time-series data from a moving image.

[0044] The control unit 15 controls the entire extraction device 10. The control unit 15 is, for example, an electronic circuit such as a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or a GPU (Graphics Processing Unit), or an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0045] The control unit 15 also has an internal memory for storing programs that define various processing procedures and control data, and executes each process using the internal memory.

[0046] The control unit 15 also functions as various processing units by running various programs. For example, the control unit 15 includes an acquisition unit 151 and an extraction unit 152.

[0047] The acquisition unit 151 acquires moving images of a person being photographed. In the following description, the moving images acquired by the acquisition unit 151 will be referred to as captured moving images.

[0048] The acquisition unit 151 may acquire captured moving images from the image capturing device 20, or may acquire captured moving images that have already been stored in the storage unit 34 as image data 341.

[0049] The extraction unit 152 extracts data indicating the movements of the head and eyes of the person being photographed from the photographed video. The photographed video is an example of a first video.

[0050] The extraction unit 152 extracts information relating to the position and orientation of the head and eyeballs at each time, for example, as shown in Fig. 3. Fig. 3 is a diagram showing an example of data indicating the movements of the head and eyeballs.

[0051] 3 is the coordinate of a predetermined part of the head (for example, the center of gravity) in three-dimensional space, and the direction of the head is the angle relative to a predetermined axis or plane.

[0052] Fig. 4 is a diagram illustrating a method for extracting data showing head and eye movements. Fig. 4 shows the faces of people captured in a captured video. Note that the dashed lines, character strings, and arrows representing axes in Fig. 4 are auxiliary and are not actually displayed in the captured video.

[0053] The three-dimensional space is represented by the x-axis, y-axis, and z-axis, which are perpendicular to each other. In FIG. 4, the front side is the positive direction of the x-axis, the right side is the positive direction of the y-axis, and the top side is the positive direction of the z-axis.

[0054] For example, the extraction unit 152 extracts O in FIG. fThe coordinates of the center of gravity of the head when the origin O is taken as the origin are extracted as the position of the head. f The angles θ and φ of rotation from the initial position of the head are extracted as the head direction.

[0055] For example, the extraction unit 152 estimates the position and direction of the head by template matching using the inverted triangular outline of the face, the arrangement of the eyes and nose, etc., or a machine learning model such as a neural network. The extraction unit 152 can construct a trained machine learning model from the extraction model information 142.

[0056] Also, for example, the extraction unit 152 may e1 The coordinates on the plane of the pupil or cornea (black part of the eye) when the origin O is taken as the origin are extracted as the pupil position. e1 is located at the center of the conjunctiva (white of the eye). It is known that the position of the pupil reflects the gaze position.

[0057] Furthermore, the extraction unit 152 extracts the coverage rate of the eyeball by the eyelid (eyelid coverage rate). Fig. 5 is a diagram illustrating the coverage rate of the eyeball by the eyelid. As shown in Fig. 5, if the vertical length (z-axis direction) of the entire eyeball including the eyelid is L and the length from the upper end to the lower end of the eyelid is L1, the extraction unit 152 can extract L1 / L as the coverage rate of the eyeball by the eyelid.

[0058] For example, the extraction unit 152 estimates the pupil position and the coverage rate of the eyeball by the eyelid using template matching or a machine learning model such as a neural network, using the presence of a black circle in a horizontally elongated white oval at the top of the face as a clue. The extraction unit 152 can construct a trained machine learning model from the extraction model information 142.

[0059] Although an example in which the extraction unit 152 extracts the movement of the right eyeball is described here, the extraction unit 152 can similarly extract the movement of the left eyeball.

[0060] Furthermore, the extraction unit 152 can improve the accuracy of extraction by using a dedicated device for extracting not only the captured video but also the movements of the head and eyes.

[0061] For example, by using an eye tracker, the extraction unit 152 can extract the head position, head direction, pupil position, and eyelid coverage ratio more accurately.

[0062] For example, Figure 3 shows that the head position at "2021 / 11 / 5 12:00:03" is (x, y, z) = (0, -3, 1), the head direction is (θ, φ) = (15°, 10°), the eyelid coverage ratio is 0.1, and the pupil position is (y, z) = (0.3, 0).

[0063] Furthermore, when the captured video image transitions from the state in Fig. 4 to the state in Fig. 6 or Fig. 7, the information extracted by the extraction unit 152 changes. Fig. 6 and Fig. 7 are diagrams for explaining a method for extracting data indicating the movement of the head and eyeballs.

[0064] For example, in Figure 6, the head position has moved in the negative direction of the y-axis, the pupil position has moved in the positive direction of the y-axis, and the eyelid coverage ratio has decreased compared to Figure 4. Also, for example, in Figure 7, the eyelid coverage ratio has increased compared to Figure 4.

[0065] 8 is a diagram illustrating an example of the configuration of a collection device according to the first embodiment. As illustrated in FIG. 8, the collection device 30 includes a communication unit 31, an input unit 32, an output unit 33, a storage unit 34, and a control unit 35.

[0066] The communication unit 31 performs data communication with other devices via a network. For example, the communication unit 31 is an NIC.

[0067] The input unit 32 is an interface for receiving data input, and is connected to input devices such as a mouse and a keyboard, for example.

[0068] The output unit 33 is an interface for outputting data, and is connected to output devices such as a display and a speaker.

[0069] The storage unit 34 is a storage device such as an HDD, an SSD, an optical disk, etc. The storage unit 34 may also be a data-rewritable semiconductor memory such as a RAM, a flash memory, or an NVSRAM.

[0070] The storage unit 34 stores an OS and various programs executed by the collection device 30. The storage unit 34 stores image data 341, created model information 342, and learning data 343.

[0071] The image data 341 is data of moving images created by the collection device 30. The created model information 342 is information such as parameters for constructing a model for creating moving images from time-series data. The learning data 343 is information related to collected labels.

[0072] The control unit 35 controls the entire collection device 30. The control unit 35 is, for example, an electronic circuit such as a CPU, an MPU, or a GPU, or an integrated circuit such as an ASIC or an FPGA.

[0073] The control unit 35 also has an internal memory for storing programs that define various processing procedures and control data, and executes each process using the internal memory.

[0074] The control unit 35 also functions as various processing units by running various programs. For example, the control unit 35 includes a creating unit 351 and a collecting unit 352.

[0075] The creation unit 351 creates a moving image showing the facial appearance based on data extracted from a captured moving image of the user and indicating the movement of the head and eyes of the person being photographed. In the following description, the moving image created by the creation unit 351 will be referred to as a created moving image. The created moving image is an example of a second moving image.

[0076] The creating unit 351 can create a created moving image using 3D modeling based on the data (time-series data) that indicates the movements of the head and eyes of the person being photographed, extracted by the extraction device 10.

[0077] First, the creation unit 351 draws a 3D model of a face using computer graphics. The creation unit 351 arranges the drawn 3D model in chronological order in a three-dimensional space based on the time-series data.

[0078] Furthermore, the creation unit 351 implements the eyelids, pupils, and corneas (black parts) of the eyeballs in the 3D face model to function as moving parts.

[0079] The creating unit 351 moves the eyelids in the created moving image in accordance with the eyelid coverage ratio included in the time-series data, and adds motions that reproduce the degree of eye opening and closing.

[0080] Furthermore, the creating unit 351 changes the positions of the pupil (pupil) and cornea (black part of the eye) on the cornea (white part of the eye) according to the pupil position included in the time-series data, and adds a motion that reproduces the change in gaze position.

[0081] As a result, the creating unit 351 creates a created moving image that reflects the movements of the head and eyeballs of the person being photographed.

[0082] The creating unit 351 may create a created video from time-series data using a deep learning model (for example, a generative adversarial network). In this case, the deep learning model is constructed from the created model information 342.

[0083] Furthermore, the face in the created video is different from the face of the person being photographed. The face in the created video may be, for example, that of an actor, a character, or a fictional person.

[0084] Furthermore, the creating unit 351 changes the appearance of a predetermined part of the face in the created video to a different appearance from that appearing in the captured video, according to the movements of the head and eyes indicated in the time-series data.

[0085] For example, the creating unit 351 can exaggerate the movements of the head and eyes in the created moving image to a range that does not appear in the time-series data.

[0086] Such intentional transformation of movements emphasizes the psychological states that appear in the created video, making it easier for evaluators to assign labels.

[0087] For example, if the eye movement indicated in the data satisfies a predetermined condition, the creating unit 351 increases the proportion of the eyelid covering the eye in the facial video, which makes it easier to assign a label related to the concentration level.

[0088] It is known that when the level of concentration decreases, the size of the pupil generally becomes smaller and fluctuates. However, it is thought to be difficult for evaluators to intuitively recognize pupil movement from the created video.

[0089] On the other hand, it is known that when concentration levels are low, drowsiness occurs and the eyelids tend to close. This can be intuitively recognized by the evaluators when viewing the created video, and can be used as a basis for judging concentration levels.

[0090] The creating unit 351 may divide the created moving image created from the time-series data into segments of a predetermined length (for example, one minute) and transmit the segments to the providing device 40.

[0091] The collection unit 352 collects the labels assigned to the created videos. The collection unit 352 associates the labels received from the providing device 40 with time-series data or information specifying the captured videos, and stores them in the storage unit 34 as learning data 343.

[0092] 9 is a diagram showing an example of the configuration of a providing device according to the first embodiment. As shown in FIG. 9, the providing device 40 includes a communication unit 41, a storage unit 44, and a control unit 45.

[0093] The communication unit 41 performs data communication with other devices via a network. For example, the communication unit 41 is a NIC.

[0094] The storage unit 44 is a storage device such as an HDD, an SSD, an optical disk, etc. The storage unit 44 may also be a data-rewritable semiconductor memory such as a RAM, a flash memory, or an NVSRAM.

[0095] The storage unit 44 stores the OS and various programs executed by the providing device 40. The storage unit 44 stores image data 441.

[0096] The image data 441 is data of the created moving image received from the collection device 30 .

[0097] The control unit 45 controls the entire providing device 40. The control unit 45 is, for example, an electronic circuit such as a CPU, an MPU, or a GPU, or an integrated circuit such as an ASIC or an FPGA.

[0098] The control unit 45 also has an internal memory for storing programs that define various processing procedures and control data, and executes each process using the internal memory.

[0099] The control unit 45 also functions as various processing units by running various programs. For example, the control unit 45 includes a providing unit 451.

[0100] The providing unit 451 provides the created video to the rating device 50. The number of rating devices 50 and raters is not limited to those shown in FIG.

[0101] Furthermore, the providing unit 451 may provide one created video or a plurality of created videos to the rating device 50. The providing unit 451 may provide the number of created videos designated by the rater via the rating device 50.

[0102] 10 is a diagram showing an example of the configuration of the rating device according to the first embodiment. As shown in FIG. 10, the rating device 50 includes a communication unit 51, an input unit 52, an output unit 53, a storage unit 54, and a control unit 55.

[0103] The communication unit 51 performs data communication with other devices via a network. For example, the communication unit 51 is an NIC.

[0104] The input unit 52 is an interface for receiving data input, and is connected to input devices such as a mouse and a keyboard, for example.

[0105] The output unit 53 is an interface for outputting data and is connected to output devices such as a display and a speaker.

[0106] The storage unit 54 is a storage device such as an HDD, an SSD, an optical disk, etc. The storage unit 54 may also be a data-rewritable semiconductor memory such as a RAM, a flash memory, or an NVSRAM.

[0107] The storage unit 54 stores an OS and various programs executed by the rating device 50. The storage unit 54 stores image data 541, created model information 342, and learning data 343.

[0108] The image data 541 is data of the created moving image received from the providing device 40 .

[0109] The control unit 55 controls the entire evaluation device 50. The control unit 55 is, for example, an electronic circuit such as a CPU, an MPU, or a GPU, or an integrated circuit such as an ASIC or an FPGA.

[0110] The control unit 55 also has an internal memory for storing programs that define various processing procedures and control data, and executes each process using the internal memory.

[0111] The control unit 55 also functions as various processing units by running various programs. For example, the control unit 55 includes a display control unit 551 and an attachment unit 552.

[0112] The display control unit 551 displays the created moving image on the output unit 53. In this case, the output unit 53 is a display capable of displaying a screen. The display control unit 551 is an example of an output control unit.

[0113] Furthermore, the display control unit 551 causes the output device to output information indicating a policy for assigning a label together with the created video, thereby enabling the rating device 50 to guide or guide the rater in assigning a label.

[0114] For example, as shown in FIG. 11, the display control unit 551 outputs, together with the created moving image, information indicating that the coverage rate of the eyeball by the eyelid increases as the concentration level decreases.

[0115] In the example of Figure 11, the display control unit 551 displays the message "It is known that when the level of concentration decreases, the eyelids droop and the time the eyes are closed becomes longer" in the "Advice" item on the screen displayed on the output unit 53.

[0116] This allows the evaluator to be guided to focus on the eyelids of the face in the created video.

[0117] The assigning unit 552 assigns a label designated by the user to the created video. For example, the assigning unit 552 assigns to the created video a label corresponding to a selected radio button among the radio buttons displayed below the message "Please select the concentration level of people appearing in the video" on the screen of Fig. 11.

[0118] For example, the assigning unit 552 associates information specifying the created moving image being output with one of "high," "medium," and "low" selected as the concentration level and notifies the providing device 40 of the association.

[0119] [Processing of the first embodiment] FIG. 12 is a flowchart showing the flow of processing in the collection system according to the first embodiment.

[0120] First, the image capturing device 20 captures a moving image of the subject (step S101). Next, the extraction device 10 extracts time-series data representing the movements of the subject's head and eyes from the captured moving image (step S102).

[0121] Here, the collecting device 30 creates a video from the time-series data (step S103), and the providing device 40 provides the created video to the rating device (step S104).

[0122] The rating device 50 receives labels from the rater to the videos (step S105). Then, the collection device 30 associates the videos with the labels and stores or outputs them (step S106).

[0123] [Effects of the first embodiment] As explained above, the creating unit 351 creates a created video showing the facial pose based on data extracted from a captured video of the subject, which data indicates the movements of the subject's head and eyes. The collecting unit 352 collects labels assigned to the created videos.

[0124] The captured video may contain the face of the person being photographed. For this reason, it may not be desirable for the captured video to be made public to an unspecified number of users. On the other hand, the captured video may contain signals that reflect the psychological state of the person being photographed but that are not problematic if made public to an unspecified number of users. For example, such signals are signals related to the movement of the head and eyes.

[0125] In this embodiment, signals that can be made public without any problems are used to create images that reflect the psychological state of the person being photographed, and these images are used to collect labels. Therefore, this embodiment makes it possible to efficiently assign labels to moving images.

[0126] The creation unit 351 changes the appearance of a predetermined part of the face in the created video to a different appearance from that in the captured video, in accordance with the head and eye movements indicated in the data, thereby enabling the evaluator to intuitively recognize the psychological state from the created video.

[0127] If the eyeball movement shown in the data satisfies a predetermined condition, the creation unit 351 increases the proportion of the eyeball covered by the eyelids in the facial video image, thereby emphasizing the eyelid movement, which is intuitively recognizable to the evaluator.

[0128] The display control unit 551 causes the output device to output information indicating the policy for assigning labels together with the created video, thereby allowing the evaluator to assign labels according to the criteria desired by the analyst.

[0129] The display control unit 551 outputs, together with the created video, information indicating that the proportion of the eyeball covered by the eyelid increases as the concentration level decreases. This allows the evaluator to pay attention to the eyelid movement, which is likely to reveal the concentration level, and assign a label to it.

[0130] The functions of the extraction device 10, the collection device 30, and the providing device 40 of the collection system 1 may be realized by a single device. Fig. 13 is a diagram showing an example of the configuration of a collection device according to another embodiment.

[0131] The collection device 30a shown in FIG. 13 is a device that has the same functions as the extraction device 10, the collection device 30, and the providing device 40 of the collection system 1.

[0132] The communication unit 31a performs data communication with other devices via a network. The communication unit 31a performs data communication with the evaluation device 50.

[0133] The input unit 32a is an interface for receiving input of data, and the output unit 33a is an interface for outputting data.

[0134] The storage unit 34a stores image data 341a, extracted model information 342a, created model information 343a, and learning data 344a.

[0135] Image data 541 is data on captured videos and created videos. Extraction model information 342a is information for constructing a model that extracts time-series data from captured videos. Creation model information 343a is information for constructing a model that creates created videos from time-series data. Learning data 344a is information related to collected labels.

[0136] The control unit 35a includes an extraction unit 351a, a creation unit 352a, a provision unit 353a, and a collection unit 354a.

[0137] The extraction unit 351a extracts time-series data from the captured video images, similar to the extraction unit 152 of the extraction device 10.

[0138] The creating unit 352a creates a created moving image from time-series data, similar to the creating unit 351 of the collecting device 30.

[0139] The providing unit 353 a provides the created moving image to the rating device 50 in the same manner as the providing unit 451 of the providing device 40 .

[0140] The collection unit 354 a collects the labels assigned by the evaluation device 50 in the same manner as the collection unit 352 of the collection device 30 .

[0141] [System configuration, etc.] Furthermore, the components of each device shown in the figure are functional concepts and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown, and all or part of the devices can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc. Furthermore, all or any part of the processing functions performed by each device can be realized by a CPU (Central Processing Unit) and a program analyzed and executed by the CPU, or can be realized as hardware using wired logic. Note that the program may be executed not only by the CPU but also by other processors such as a GPU.

[0142] Furthermore, among the processes described in this embodiment, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method.In addition, the information including the processing procedures, control procedures, specific names, various data and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified.

[0143] [program] In one embodiment, the collection device 30 can be implemented by installing a collection program that executes the above-described collection process as package software or online software on a desired computer. For example, by having an information processing device execute the above-described collection program, the information processing device can function as the collection device 30. The information processing device referred to here includes desktop and notebook personal computers. In addition, the information processing device also includes mobile communication terminals such as smartphones, mobile phones, and PHS (Personal Handyphone Systems), as well as slate terminals such as PDAs (Personal Digital Assistants).

[0144] The collection device 30 may also be implemented as a server device that provides services related to the collection process to a client terminal device used by a user. For example, the server device may be implemented as a collection service that receives time-series data and outputs labels. In this case, the server device may be implemented as a web server or as a cloud that provides services related to the collection process through outsourcing.

[0145] 14 is a diagram showing an example of a computer that executes a collection program. The computer 1000 includes, for example, a memory 1010 and a CPU 1020. The computer 1000 also includes a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0146] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM (Random Access Memory) 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.

[0147] The hard disk drive 1090 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. That is, the program that defines each process of the collection device 30 is implemented as a program module 1093 in which computer-executable code is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing processes similar to those of the functional configuration of the collection device 30 is stored in the hard disk drive 1090. The hard disk drive 1090 may be replaced by an SSD (Solid State Drive).

[0148] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in the memory 1010 or the hard disk drive 1090. Then, the CPU 1020 reads the program module 1093 or the program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as necessary, and executes the processing of the above-described embodiment.

[0149] The program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090, but may also be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or a wide area network (WAN)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070. [Explanation of symbols]

[0150] 1. Collection System 10 Extraction device 11, 31, 31a, 41, 51 Communication Department 12, 32, 32a, 52 Input section 13, 33, 33a, 53 output section 14, 34, 34a, 44, 54 storage section 15, 35, 35a, 45, 55 Control section 20 Imaging equipment 30, 30a Collector 40 Providing device 50 Grading device 141, 341, 341a, 441, 541 Image data 142, 342a Extraction model information 151 Acquisition Department 152 Extraction part 342, 343a Creation Model Information 343, 344a Training data 351 Creation Department 352 Collection Department 451 Providing Department 551 Display control unit 552 Granting Department

Claims

1. A collection system having an extraction device, a collection device, a providing device, and an evaluation device, The extraction device comprises: an acquisition unit that acquires a first moving image captured of a user; an extractor that extracts data indicating movements of the user's head and eyeballs from the first moving image; and The collection device a creating unit that creates a second moving image representing a facial aspect based on the data extracted by the extracting device; a collection unit that collects labels assigned to the second video in the rating device via the providing device; and The providing device is a providing unit that provides the second moving image created by the collecting device to the evaluation device; The evaluation device is a display control unit that causes an output device to output a screen including the second moving image and a message regarding changes in facial appearance due to changes in concentration level; an assigning unit that assigns a label designated by a user to the second video; A collection system comprising:

2. The collection system described in claim 1, characterized in that the creation unit changes the appearance of a specified part of the face in the second video to an appearance different from the appearance in the first video, depending on the movement of the head and eyes indicated in the data.

3. The collection system according to claim 2, wherein the creation unit increases the proportion of the eyelids covering the eyeballs in the facial video image when the eyeball movement shown in the data satisfies a predetermined condition.

4. The evaluation device The collection system described in any one of claims 1 to 3, further comprising an output control unit that causes the output device to output information indicating a policy for assigning labels together with the second moving image.

5. The collection system of claim 4, characterized in that the output control unit outputs, together with the second moving image, information indicating that the proportion of the eyelid covering the eyeball increases as the concentration level decreases.

6. A collection method performed by a collection system having an extraction device, a collection device, a provision device, and an evaluation device, comprising: the extraction device acquires a first moving image of a user; the extraction device extracts data indicating movements of the user's head and eyeballs from the first moving image; the collecting device creates a second moving image representing a facial aspect based on the data extracted by the extracting device; the collection device collects, via the provision device, the labels assigned to the second video by the rating device; the providing device provides the second moving image created by the collecting device to the rating device; the evaluation device causes an output device to output a screen including the second moving image and a message regarding a change in facial appearance due to a change in concentration level; The rating device assigns a label designated by a user to the second video. A collection method characterized by:

Citation Information

Patent Citations

  • Agent device

    JP1999272640A

  • Image processor and program

    JP2005100139A

  • Image recognition device, image recognition method, and program

    JP2015133065A

  • Expression control program, recording medium, expression control device and expression control method

    JP2020091909A

  • Image processing device and image processing method

    JP2021077996A