Video session evaluation terminal, video session evaluation system, and video session evaluation program

The video session evaluation system objectively assesses online communication by analyzing moving images from video sessions, addressing the need for efficient communication in digital transformation and remote work environments.

JP7694961B2Active Publication Date: 2025-06-18IMBESIDEYOU INC
View PDF 15 Cites 0 Cited by

Patent Information

Application Number
JP2022518709
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-02-02
Publication Date
2025-06-18
Estimated Expiration
2041-02-02

AI Technical Summary

Technical Problem

Existing technologies are inadequate for objectively evaluating online communication, which is essential for efficient communication in digital transformation and remote work settings.

Method used

A video session evaluation system that includes a camera unit, gaze acquisition unit, display unit, position acquisition unit, and output unit to analyze and evaluate moving images from video sessions, providing objective feedback on communication efficiency.

Benefits of technology

The system enables objective evaluation of video sessions, improving communication efficiency by identifying key factors influencing emotional changes and providing actionable insights for improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007694961000001
    Figure 0007694961000001
  • Figure 0007694961000002
    Figure 0007694961000002
  • Figure 0007694961000003
    Figure 0007694961000003
Patent Text Reader

Abstract

[Problem] To evaluate a video session by evaluating video acquired in the video session. [Solution] A video session evaluation system according to the present disclosure comprises: an acquisition means for acquiring at least a video; a face recognition means for recognizing at least a face image of a subject included in the video for each prescribed frame; a speech recognition means for recognizing at least the speech of the subject included in the video; an evaluation means for calculating evaluation values for a prescribed aspect on the basis of both the recognized face images and the recognized speech; an output means for outputting the evaluation values as change information in a time series; and a specification means for referencing other change information associated with other videos and specifying other videos that include the same pattern as the pattern extracted from the change information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a video session evaluation terminal, a video session evaluation system, and a video session evaluation program.

Background Art

[0002] Conventionally, techniques for analyzing the emotions received by others in response to the speech of a speaker are known (see, for example, Patent Document 1). Also known are techniques for analyzing the changes in the facial expressions of a subject over time in a time series and estimating the emotions held during that period (see, for example, Patent Document 2). Furthermore, techniques for identifying the elements that most influenced the change in emotions are also known (see, for example, Patent Documents 3 to 5). Still further, techniques for comparing the subject's usual facial expression with the current facial expression and issuing an alert when the facial expression is gloomy are also known (see, for example, Patent Document 6). Also known are techniques for comparing the subject's facial expression in a normal state (when expressionless) with the current facial expression and determining the degree of the subject's emotion (see, for example, Patent Documents 7 to 9). Furthermore, techniques for analyzing the emotions of an organization and the atmosphere within a group felt by an individual are also known (see, for example, Patent Documents 10 and 11).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Patent Document 3

Patent Document 4

Patent Document 5

Patent Document 6

Patent Document 7

[0004] All of the above-mentioned technologies are merely secondary functions in a situation where communication in the real space is mainly involved. That is, it was not born in a situation where communication such as business and classes is mainly conducted online due to the recent DX (Digital Transformation) of business and the spread of infectious diseases around the world.

[0005] An object of the present invention is to objectively evaluate the exchanged communication in order to perform more efficient communication in a situation where online communication is mainly involved. [Means for Solving the Problems]

[0006] According to the present invention, a camera unit that acquires a moving image obtained by photographing a subject; a gaze acquisition unit that acquires the movement of the gaze of the subject based on the acquired moving image; a display unit that continuously displays a plurality of images to the subject; a position acquisition unit that acquires the positional relationship between the camera unit and the display unit; an output unit that associates and outputs the movement of the gaze for each of the plurality of displayed images; are obtained.

[0007] Also, according to the present invention, At least acquisition means for acquiring a moving image, Face recognition means for recognizing at least a face image of a target person included in the moving image for each predetermined frame, Voice recognition means for recognizing at least the voice of the target person included in the moving image, Evaluation means for calculating an evaluation value from a predetermined viewpoint based on both the recognized face image and the voice, Output means for outputting the evaluation value as change information along a time series, Specific means for specifying another moving image including the same pattern as the pattern extracted from the change information by referring to other change information regarding another moving image, comprising A moving image analysis system. is obtained.

[0008] Also, according to the present invention, Acquisition means for acquiring a moving image related to a video session carried out between at least two user terminals, Face recognition means for recognizing a face image of a user included in the moving image for each predetermined frame, Voice recognition means for recognizing at least the voice of the user included in the moving image, Evaluation means for calculating an evaluation value from a plurality of viewpoints based on both the recognized face image and the voice, Storage means for storing the evaluation value as change information along a time series, Detection means for detecting that only the evaluation value in one of the plurality of viewpoints has changed beyond a predetermined range, Specific frame acquisition means for acquiring a specific frame including the detected range, comprising A moving image analysis system. is obtained.

[0009] Also, according to the present invention, Acquisition means for acquiring a moving image related to a video session carried out between at least two user terminals, Face recognition means for recognizing a face image of a user included in the moving image for each predetermined frame, voice recognition means for recognizing at least the voice of the user included in the moving image; facial expression evaluation means for calculating facial expression evaluation values from a plurality of viewpoints based on the recognized facial image; voice evaluation means for calculating voice evaluation values from a plurality of viewpoints based on the recognized voice; facial expression-voice correlation evaluation means for evaluating the correlation between the facial expression evaluation value and the voice evaluation value of at least one of the users; detection means for detecting that the facial expression evaluation value and the voice evaluation value have changed beyond a predetermined range from the correlation; peculiar frame acquisition means for acquiring a peculiar frame including the detected range, A moving image analysis system. is obtained.

[0010] Also, according to the present invention, acquisition means for acquiring a moving image related to a video session carried out between at least two user terminals; face recognition means for recognizing the face image of the user included in the moving image for each predetermined frame; emotion evaluation means for analyzing the standard facial expression of the user from the recognized facial image and evaluating the degree of deviation from the standard facial expression of the user; concentration evaluation means for evaluating at least the amount of movement of the pupil or face of the user from the recognized facial image; safety evaluation means for evaluating the emotion related to the anxiety of the user from the recognized facial image; score generation means for generating a score based on two or more of the emotion evaluation means, the concentration evaluation means, and the safety evaluation means, A moving image analysis system. is obtained.

[0011] Also, according to the present invention, acquisition means for acquiring a moving image of a video session conducted with another terminal; Face recognition means for recognizing at least the face image of the subject included in the moving image for each predetermined frame, Subject identification means for identifying the subject frames in which the subject is recognized from each of the moving images related to a plurality of the video sessions, Further comprising a digest generation video means for generating a digest video by connecting a plurality of the subject frames, Video analysis system. is obtained.

[0012] Also, according to the present invention, Acquisition means for acquiring a moving image of a video session conducted with another terminal, Face recognition means for recognizing at least the face image of the subject included in the moving image for each predetermined frame, Evaluation means for evaluating at least the amount of movement of the pupil and the face of the user respectively from the recognized face image, Score calculation means for calculating a score regarding concentration based on the evaluation, comprising Video analysis system. is obtained.

[0013] Also, according to the present invention, Acquisition means for acquiring a moving image of a video session conducted with another terminal, Face recognition means for recognizing at least the face image of the user included in the moving image for each predetermined frame, Voice recognition means for recognizing at least the voice of the user included in the moving image Evaluation means for calculating evaluation values from a plurality of viewpoints based on both the recognized face image and the voice, State evaluation means for evaluating the state of the user based on a first evaluation analysis value obtained by analyzing the evaluation value in a first period and a second evaluation analysis value obtained by analyzing the evaluation value in a second period longer than the first period for the same user, comprising Video analysis system. is obtained.

[0014] Also, according to the present invention, acquisition means for acquiring a moving image of a video session conducted with another terminal; face recognition means for recognizing at least a face image of a target person included in the moving image for each predetermined frame; voice recognition means for recognizing at least the voice of the user included in the moving image; evaluation means for calculating evaluation values from a plurality of viewpoints based on both the recognized face image and the voice; response acquisition means for acquiring response information of the user to question information created based on the plurality of viewpoints; a state evaluation system for evaluating the state of the user by comparing the evaluation value and the response information. is obtained. is obtained.

[0015] Also, according to the present invention, acquisition means for acquiring a moving image of a video session conducted with another terminal; face recognition means for recognizing at least a face image of the user included in the moving image for each predetermined frame; voice recognition means for recognizing at least the voice of the target person included in the moving image; evaluation means for calculating evaluation values from a plurality of viewpoints based on both the recognized face image and the voice; annotation reception means for receiving an annotation from the user with respect to the evaluation value; a video session evaluation system comprising display means for simultaneously displaying the evaluation value and the received annotation. is obtained. is obtained.

[0016] Also, according to the present invention, video acquisition means for acquiring a moving image of a business video session conducted between a business-side terminal of a business-side person in charge and a business-site terminal of a business-site person in charge; contract information acquisition means for acquiring contract information of the business video session; Face recognition means for recognizing at least one of the face images of the business-side person in charge or the business partner person in charge included in the moving image for each predetermined frame; Voice recognition means for recognizing at least one of the voices of the business-side person in charge or the business partner person in charge included in the moving image; Evaluation means for calculating evaluation values from a plurality of viewpoints based on both the recognized face image and the voice; Model generation means for generating a contract estimation model for estimating the contract rate of the moving image of another business video session as a plurality of ranks using the evaluation value and the contract information as teacher data; Determination means for associating one of the plurality of ranks with a new business video session using the model; A video session evaluation system. is obtained.

[0017] Also, according to the present invention, A lecturer terminal having a lecturer-side camera for capturing at least the face of the lecturer user; A learner terminal communicably connected to the lecturer terminal via a network, the learner terminal having a learner-side face camera for capturing at least the face of the learner and a hand camera for capturing the hands of the learner; Acquisition means for acquiring a moving image of a video session performed between; Hand movement recognition means for recognizing the movement of the hands of the learner in the moving image acquired from at least the hand camera for each predetermined frame; Estimation means for estimating the degree of understanding of the learner based on the recognized movement of the hands; Comprising A video session evaluation system. is obtained.

Effect of the Invention

[0018] According to the present disclosure, by analyzing and evaluating the moving image of a video session, it is possible to objectively evaluate the content in particular.

[0019] In particular, according to the present invention, in a situation mainly involving online communication, in order to perform more efficient communication, the exchanged communication can be objectively evaluated.

Brief Description of the Drawings

[0020]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Embodiments for Carrying Out the Invention

[0021] The contents of the embodiments of the present disclosure will be listed and described. The present disclosure has the following configuration. [Item 1] A camera unit that acquires a moving image obtained by photographing a subject, A gaze acquisition unit that acquires the movement of the gaze of the subject based on the acquired moving image, A display unit that continuously displays a plurality of images to the subject, A position acquisition unit that acquires the positional relationship between the camera unit and the display unit, An output unit that associates and outputs the movement of the gaze for each of the plurality of displayed images, A gaze evaluation system comprising: [Item 2] The gaze evaluation system according to Item 1, wherein The output unit outputs by superimposing a heatmap indicating the fixation time generated based on the movement of the line of sight on the image. Eye line evaluation system [Item 3] The eye line evaluation system according to Item 1, The output unit further outputs in association with the movement of the line of sight of other subjects who displayed the same image. Eye line evaluation system [Item 4] The eye line evaluation system according to Item 3, further comprising a specific determination unit that determines whether the movement of the line of sight associated with the subject is specific compared to the movement of the line of sight associated with the other subjects. Eye line evaluation system. [Item 5] The eye line evaluation system according to any one of Items 1 to 4, The output unit outputs in association with a normalized heatmap obtained by normalizing the movement of the line of sight of a plurality of the subjects for each image. Eye line evaluation system. [Item 6] Acquisition means for acquiring at least a moving image, Face recognition means for recognizing at least a face image of a subject included in the moving image for each predetermined frame, Voice recognition means for recognizing at least the voice of the subject included in the moving image, Evaluation means for calculating an evaluation value from a predetermined viewpoint based on both the recognized face image and the voice, Output means for outputting the evaluation value as change information along a time series, specific means for specifying another moving image including the same pattern as the pattern extracted from the change information by referring to other change information regarding another moving image, Moving image analysis system. [Item 7] The moving image analysis system according to Item 6, The output means outputs the evaluation value as graph information along a time series, The specific means includes specific means for receiving a selection operation of a part of the graph information from an analyzing user and specifying corresponding frames of other moving images including a graph pattern identical to the graph pattern of the selected part. Moving image analysis system. [Item 8] The moving image analysis system according to Item 6 or Item 7, wherein the specific means specifies other moving images including a pattern identical to the pattern extracted from the change information in the same time zone. Moving image analysis system. [Item 9] An acquisition means for acquiring a moving image related to a video session implemented between at least two user terminals, a face recognition means for recognizing a face image of a user included in the moving image for each predetermined frame, a voice recognition means for recognizing at least the voice of the user included in the moving image, an evaluation means for calculating evaluation values from a plurality of viewpoints based on both the recognized face image and the voice, a storage means for storing the evaluation values as change information along a time series, a detection means for detecting that only the evaluation value in one of the plurality of viewpoints has changed beyond a predetermined range, and a specific frame acquisition means for acquiring a specific frame including the detected range. Moving image analysis system. [Item 10] The moving image analysis system according to Item 9, wherein the plurality of viewpoints includes a first viewpoint and a second viewpoint associated with mutually opposite attributes, and the detection means detects that the evaluation value according to the first viewpoint and the evaluation value according to the second viewpoint have deviated beyond the predetermined range. Moving image analysis system. [Item 11] The moving image analysis system according to Item 9, wherein The detection means detects that within a predetermined time after the elapse of the first time point, the evaluation value in one of the viewpoints changes beyond a predetermined range and immediately returns to a value substantially the same as the evaluation value at the first time point. Moving image analysis system. [Item 12] The moving image analysis system according to any one of Items 9 to 11, further comprising digest generation video means for generating a digest video by concatenating a plurality of the specific frames acquired from the moving image. Video analysis system. [Item 13] The moving image analysis system according to any one of Items 9 to 12, further comprising specific frame corresponding text output means for converting the voice corresponding to the specific frame into text and outputting it. Moving image analysis system. [Item 14] The moving image analysis system according to any one of Items 9 to 13, the video session can share screen information displayed on the screen of one user terminal, further comprising shared screen output means for outputting at least the screen information corresponding to the shared specific frame. Moving image analysis system. [Item 15] acquisition means for acquiring a moving image related to a video session implemented between at least two user terminals, face recognition means for recognizing a face image of a user included in the moving image for each predetermined frame, voice recognition means for recognizing at least the voice of the user included in the moving image, facial expression evaluation means for calculating facial expression evaluation values from a plurality of viewpoints based on the recognized face image, voice evaluation means for calculating voice evaluation values from a plurality of viewpoints based on the recognized voice, facial expression and voice correlation evaluation means for evaluating the correlation between the facial expression evaluation value and the voice evaluation value of at least one of the users, Detection means for detecting that the facial expression evaluation value and the voice evaluation value have changed beyond a predetermined range from the correlation relationship; and specific frame acquisition means for acquiring a specific frame including the detected range. Moving image analysis system. [Item 16] The moving image analysis system according to Item 15, further comprising attribute evaluation means for associating attributes corresponding to each of the facial expression evaluation value and the voice evaluation value; The detection means detects that the attribute of the facial expression evaluation value and the attribute of the voice evaluation value are opposite to each other. Moving image analysis system. [Item 17] The moving image analysis system according to any one of Items 15 or 16, further comprising digest generation video means for concatenating a plurality of the specific frames obtained from the moving image to generate a digest video. Video analysis system. [Item 18] The moving image analysis system according to any one of Items 15 to 17, further comprising specific frame corresponding text output means for converting the voice corresponding to the specific frame into text and outputting it. Moving image analysis system. [Item 19] The moving image analysis system according to any one of Items 15 to 18, the video session can share screen information displayed on the screen of one user terminal, further comprising shared screen output means for outputting at least the screen information corresponding to the shared specific frame. Moving image analysis system. [Item 20] acquisition means for acquiring a moving image related to a video session implemented between at least two user terminals; face recognition means for recognizing a face image of a user included in the moving image for each predetermined frame; Emotion evaluation means for analyzing the standard expression of the user from the recognized face image and evaluating the degree of deviation from the standard expression of the user, Concentration evaluation means for evaluating at least the amount of movement of the user's pupils or face movement from the recognized face image, Safety evaluation means for evaluating the emotion related to the user's anxiety from the recognized face image, Score generation means for generating a score based on two or more evaluations among the emotion evaluation means, the concentration evaluation means, and the safety evaluation means, Moving image analysis system. [Item 21] Acquisition means for acquiring a moving image of a video session conducted with another terminal, Face recognition means for recognizing at least the face image of the subject included in the moving image for each predetermined frame, Subject identification means for identifying the subject frames in which the subject is recognized from each of the moving images related to a plurality of the video sessions, Further comprising a digest generation moving image means for concatenating a plurality of the subject frames to generate a digest video, Moving image analysis system. [Item 22] Acquisition means for acquiring a moving image of a video session conducted with another terminal, Face recognition means for recognizing at least the face image of the subject included in the moving image for each predetermined frame, Evaluation means for evaluating at least the amount of movement of the user's pupils and the amount of face movement respectively from the recognized face image, Score calculation means for calculating a score related to concentration based on the evaluation, Moving image analysis system. [Item 23] Acquisition means for acquiring a moving image of a video session conducted with another terminal, Face recognition means for recognizing at least the face image of the user included in the moving image for each predetermined frame, Voice recognition means for recognizing at least the voice of the user included in the moving image Evaluation means for calculating evaluation values from a plurality of viewpoints based on both the recognized face image and the voice State evaluation means for evaluating the state of the user based on a first evaluation analysis value obtained by analyzing the evaluation value in a first period and a second evaluation analysis value obtained by analyzing the evaluation value in a second period longer than the first period, for the same user Moving image analysis system [Item 24] The moving image analysis system according to Item 23, comprising Trend detection means for detecting a predetermined trend regarding the second evaluation analysis value Correction means for correcting the first evaluation analysis value according to the detected trend Video analysis system [Item 25] Acquisition means for acquiring a moving image of a video session conducted with another terminal Face recognition means for recognizing at least the face image of a subject included in the moving image for each predetermined frame Voice recognition means for recognizing at least the voice of the user included in the moving image Evaluation means for calculating evaluation values from a plurality of viewpoints based on both the recognized face image and the voice Answer acquisition means for acquiring answer information of the user with respect to question information created based on the plurality of viewpoints Evaluating the state of the user by comparing the evaluation value and the answer information State evaluation system [Item 26] The state evaluation system according to Item 25, further comprising Organization scoring means for scoring the organization state based on the states of all the users belonging to the same organization State evaluation system [Item 27] Acquisition means for acquiring a moving image of a video session conducted with another terminal Face recognition means for recognizing at least the face image of the user included in the moving image for each predetermined frame, Voice recognition means for recognizing at least the voice of the target person included in the moving image, Evaluation means for calculating evaluation values from a plurality of viewpoints based on both the recognized face image and the voice, Annotation reception means for receiving an annotation from the user with respect to the evaluation value, A video session evaluation system comprising display means for simultaneously displaying the evaluation value and the received annotation. Video session evaluation system. [Item 28] The video session evaluation system according to Item 27, further comprising graph output means for outputting the evaluation value as graph information along a time series, The display means superimposes and displays the annotation on the graph information. Video session evaluation system. Video session evaluation system. [Item 29] Video acquisition means for acquiring a moving image of a business video session conducted between a business-side terminal of a business-side person in charge and a business-site terminal of a business-site person in charge, Contract information acquisition means for acquiring contract information of the business video session, Face recognition means for recognizing at least the face image of either the business-side person in charge or the business-site person in charge included in the moving image for each predetermined frame, Voice recognition means for recognizing at least the voice of either the business-side person in charge or the business-site person in charge included in the moving image, Evaluation means for calculating evaluation values from a plurality of viewpoints based on both the recognized face image and the voice, Model generation means for generating a contract estimation model that estimates the contract rate of the moving image of another business video session as a plurality of ranks using the evaluation value and the contract information as teacher data, Determination means for associating one of the plurality of ranks with a new business video session using the model. Video session evaluation system. [Item 30] The video session evaluation system according to Item 29, further comprising an expected value calculation means for calculating an expected business value in a predetermined period based on the estimated transaction amount related to the business video session and the rank within the organization. Video session evaluation system. [Item 31] A lecturer terminal having a lecturer-side camera for capturing at least the face of the lecturer user; a learner terminal communicably connected to the lecturer terminal via a network, the learner terminal having a learner-side face camera for capturing at least the face of the learner and a hand camera for capturing the area in front of the learner's hands; an acquisition means for acquiring a moving image of a video session conducted between the two; a hand movement recognition means for recognizing the movement of the learner's hands in the moving image acquired from at least the hand camera for each predetermined frame; an estimation means for estimating the degree of understanding of the learner based on the recognized movement of the hands; comprising Video session evaluation system. [Item 32] The video session evaluation system according to Item 31, a voice recognition means for recognizing at least the voice of the learner included in the moving image; an evaluation means for calculating evaluation values from a plurality of viewpoints based on both the recognized face image and the voice; the estimation means estimates the degree of understanding of the learner based on the movement of the hands and the evaluation value. Video session evaluation system. [Item 33] The video session evaluation system according to Item 31 or Item 32, wherein the estimation means estimates the degree of understanding of the learner according to the amount of the recognized movement of the hands. Video session evaluation system. [Item 34] A video session evaluation system according to any one of Items 31 to 33, wherein the estimation means further has an alert means for issuing an alert when the face of the attendee is not shown by the attendee side face camera and the movement of the hand of the attendee is not recognized from the hand camera. Video session evaluation system.

[0022] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In the present specification and drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant description is omitted.

[0023] <Basic functions> The video session evaluation system of the present embodiment is a system for analyzing and evaluating a specific emotion (a feeling that occurs in response to one's own or other people's words and actions. Pleasure, discomfort, or the degree thereof, etc.) different from others for an analysis target person among a plurality of people in an environment where a video session (hereinafter referred to as an online session including one-way and two-way) is performed by a plurality of people.

[0024] The online session is, for example, an online meeting, an online class, an online chat, etc., and terminals installed at a plurality of locations are connected to a server via a communication network such as the Internet so that moving images can be exchanged between the plurality of terminals through the server.

[0025] The moving images handled in the online session include the face images and voices of the users using the terminals. The moving images also include images such as materials shared and viewed by a plurality of users. It is possible to switch between the face image and the material image on the screen of each terminal and display only one of them, or to divide the display area and display the face image and the material image simultaneously. It is also possible to display the image of one person among a plurality of people in full screen, or to divide and display the images of some or all of the users in a small screen.

[0026] Among a plurality of users who participate in an online session using terminals, it is possible to specify any one or more of them as the analysis target. For example, the leader, facilitator, or administrator of the online session (hereinafter collectively referred to as the organizer) specifies any user as the analysis target. The organizer of the online session is, for example, a lecturer in an online class, a chairperson or facilitator in an online meeting, a coach in a session for coaching purposes, etc. The organizer of the online session is usually one of the plurality of users who participate in the online session, but may be another person who does not participate in the online session. Note that it is also possible not to specify the analysis target and analyze all participants as the analysis target.

[0027] Also, it is possible for the leader, facilitator, or administrator of the online session (hereinafter collectively referred to as the organizer) to specify any user as the analysis target. The organizer of the online session is, for example, a lecturer in an online class, a chairperson or facilitator in an online meeting, a coach in a session for coaching purposes, etc. The organizer of the online session is usually one of the plurality of users who participate in the online session, but may be another person who does not participate in the online session.

[0028] When a video session is established among a plurality of terminals, the video session evaluation system according to the present embodiment displays at least a moving image acquired from the video session. The displayed moving image is acquired by the terminal, and at least a face image included in the moving image is identified for each predetermined frame unit. Then, an evaluation value regarding the identified face image is calculated. The evaluation value is shared as necessary.

[0029] In particular, in this embodiment, the acquired moving image is stored in the terminal, analyzed and evaluated on the terminal, and the result is provided to the user of the terminal. Therefore, for example, even in the case of a video session including personal information or a video session including confidential information, it is possible to analyze and evaluate without providing the video itself to an external evaluation institution or the like. Further, if necessary, only the evaluation result (evaluation value) can be provided to an external terminal, so that the result can be visualized or cross-analyzed.

[0030] As shown in FIG. 1, the video session evaluation system according to this embodiment includes user terminals 10 and 20 each having at least an input unit such as a camera unit and a microphone unit, a display unit such as a display, and an output unit such as a speaker, a video session service terminal 30 that provides a two-way video session to the user terminals 10 and 20, and an evaluation terminal 40 that performs part of the evaluation related to the video session.

[0031] <Hardware configuration example> FIG. 2 is a diagram showing a hardware configuration example of a computer that realizes each of the terminals 10 to 40 according to this embodiment. The computer includes at least a control unit 110, a memory 120, a storage 130, a communication unit 140, and an input / output unit 150. These are electrically connected to each other through a bus 160.

[0032] The control unit 110 is an arithmetic device that controls the operation of the entire terminal, controls the transmission and reception of data between each element, and performs information processing necessary for the execution and authentication processing of applications. For example, the control unit 110 is a processor such as a CPU, and executes programs and the like stored in the storage 130 and expanded in the memory 120 to perform each information process.

[0033] The memory 120 includes a main memory composed of a volatile storage device such as a DRAM, and an auxiliary memory composed of a non-volatile storage device such as a flash memory or an HDD. The memory 120 is used as a work area of the control unit 110, and stores a BIOS executed at the startup of each terminal and various setting information.

[0034] Storage 130 stores various programs such as application programs. A database storing data used for each process may be constructed in Storage 130. In particular, in the present embodiment, the moving images in the online session are not recorded in the storage 130 of the video session service terminal 30, but are stored in the storage 130 of the user terminal 10. Further, the evaluation terminal 40 stores applications and other programs necessary for evaluating the moving images acquired on the user terminal 10, and appropriately provides them for the user terminal 10 to use. Note that, for example, only the results analyzed and evaluated by the user terminal 10 may be shared in the storage 13 managed by the evaluation terminal 40.

[0035] The communication unit 140 connects the terminal to the network. The communication unit 140 communicates with external devices directly or via a network access point by, for example, a method such as a wired LAN, a wireless LAN, Wi-Fi (registered trademark), infrared communication, Bluetooth (registered trademark), short-distance or non-contact communication.

[0036] The input / output unit 150 is, for example, an information input device such as a keyboard, a mouse, a touch panel, and an output device such as a display.

[0037] The bus 160 is commonly connected to the above elements and transmits, for example, an address signal, a data signal, and various control signals.

[0038] In particular, the evaluation terminal according to the present embodiment acquires a moving image from the video session service terminal, identifies at least a face image included in the moving image for each predetermined frame unit, and calculates an evaluation value related to the face image (details will be described later). <Method for acquiring video> As shown in FIG. 3, the video session service provided by the video session service terminal (hereinafter sometimes simply referred to as "this service") enables two-way communication with the user terminals 10 and 20 by means of images and sounds. This service can display a moving image acquired by the camera unit of the other user terminal on the display of the user terminal, and output the sound acquired by the microphone unit of the other user terminal from the speaker.

[0039] In addition, this service is configured such that either or both user terminals can record (record) moving images and sounds (collectively referred to as "moving images, etc.") in the storage unit on at least one of the user terminals. The recorded moving image information Vs (hereinafter referred to as "recorded information") is cached in the user terminal that started the recording and is recorded only locally on one of the user terminals. If necessary, the user can view the recorded information by himself / herself within the scope of using this service, share it with others, etc.

[0040] The user terminal 10 acquires the recorded information and performs analysis and evaluation as described below.

[0041] The user terminal 10 evaluates the video acquired as described above by the following analysis.

[0042] <Functional Configuration Example 1> FIG. 4 is a block diagram showing a configuration example according to this embodiment. As shown in FIG. 4, the video session evaluation system of this embodiment is realized as a functional configuration of the user terminal 10. That is, the user terminal 10 includes, as its functions, a moving image acquisition unit 11, a biological reaction analysis unit 12, a specific determination unit 13, a related event identification unit 14, a clustering unit 15, and an analysis result notification unit 16.

[0043] Each of the above functional blocks 11 to 16 can be configured by, for example, any of the hardware, DSP (Digital Signal Processor), and software provided in the user terminal 10. For example, when configured by software, each of the above functional blocks 11 to 16 is actually configured with a computer's CPU, RAM, ROM, etc., and is realized by the operation of a program stored in a recording medium such as RAM, ROM, hard disk, or semiconductor memory.

[0044] The moving image acquisition unit 11 acquires, from each terminal, a moving image obtained by photographing a plurality of people (a plurality of users) with a camera provided in each terminal during an online session. Whether the moving image acquired from each terminal is set to be displayed on the screen of each terminal or not does not matter. That is, the moving image acquisition unit 11 acquires moving images from each terminal, including the moving images being displayed and not being displayed on each terminal.

[0045] The biological reaction analysis unit 12 analyzes changes in the biological reactions of each of the plurality of people based on the moving image (regardless of whether it is being displayed on the screen or not) acquired by the moving image acquisition unit 11. In the present embodiment, the biological reaction analysis unit 12 separates the moving image acquired by the moving image acquisition unit 11 into a set of images (a collection of frame images) and audio, and analyzes changes in the biological reactions from each of them.

[0046] For example, the biological reaction analysis unit 12 analyzes changes in the biological reaction related to at least one of the expression, eye line, pulse, and facial movement by analyzing the user's face image using the frame images separated from the moving image acquired by the moving image acquisition unit 11. In addition, the biological reaction analysis unit 12 analyzes changes in the biological reaction related to at least one of the user's speech content and voice quality by analyzing the audio separated from the moving image acquired by the moving image acquisition unit 11.

[0047] When a person's emotions change, they manifest as changes in biological reactions such as facial expressions, eye lines, pulse, facial movements, speech content, and voice quality. In this embodiment, the changes in the user's emotions are analyzed by analyzing the changes in the user's biological reactions. The emotion analyzed in this embodiment is, for example, the degree of pleasure / displeasure. In this embodiment, the biological reaction analysis unit 12 calculates a biological reaction index value that reflects the content of the change in the biological reaction by quantifying the change in the biological reaction according to a predetermined standard.

[0048] The analysis of the change in facial expression is performed, for example, as follows. That is, for each frame image, the face region is identified from the frame image, and the facial expressions of the identified face are classified into a plurality according to a pre-trained image analysis model. Then, based on the classification results, it is analyzed whether a positive facial expression change has occurred, a negative facial expression change has occurred, and the magnitude of the facial expression change that has occurred between consecutive frame images, and a facial expression change index value corresponding to the analysis result is output.

[0049] The analysis of the change in eye line is performed, for example, as follows. That is, for each frame image, the eye region is identified from the frame image, and by analyzing the directions of both eyes, it is analyzed where the user is looking. For example, it is analyzed whether the user is looking at the face of the speaker being displayed, the shared material being displayed, or outside the screen. Also, it may be analyzed whether the movement of the eye line is large or small, and whether the frequency of movement is high or low. The change in eye line is also related to the user's concentration. The biological reaction analysis unit 12 outputs an eye line change index value according to the analysis result of the change in eye line.

[0050] The analysis of the change in pulse rate is performed as follows, for example. That is, for each frame image, the face region is specified from within the frame image. Then, using a learned image analysis model that captures the numerical values of the face color information (G of RGB), the change in the G color on the face surface is analyzed. By arranging the results along the time axis, a waveform representing the change in color information is formed, and the pulse rate is specified from this waveform. When a person is nervous, the pulse rate increases, and when the person is calm, the pulse rate decreases. The biological reaction analysis unit 12 outputs a pulse rate change index value according to the analysis result of the change in pulse rate.

[0051] The analysis of the change in face movement is performed as follows, for example. That is, for each frame image, the face region is specified from within the frame image, and by analyzing the orientation of the face, it is analyzed where the user is looking. For example, it is analyzed whether the user is looking at the face of the speaker being displayed, at the shared material being displayed, outside the screen, etc. Also, it may be analyzed whether the face movement is large or small, whether the frequency of movement is high or low, etc. The face movement and the movement of the line of sight may be analyzed together. For example, it may be analyzed whether the user is looking straight at the face of the speaker being displayed, looking up or down, or looking obliquely. The biological reaction analysis unit 12 outputs a face orientation change index value according to the analysis result of the change in face orientation.

[0052] The analysis of the speech content is performed as follows, for example. That is, the biological reaction analysis unit 12 performs known speech recognition processing on the speech for a specified time (for example, a time of about 30 to 150 seconds) to convert the speech into a character string, and by performing morphological analysis on the character string, words that are unnecessary for representing conversation such as particles and articles are removed. Then, the remaining words are vectorized, and it is analyzed whether a positive emotional change is occurring, whether a negative emotional change is occurring, and the magnitude of the emotional change occurring, and a speech content index value according to the analysis result is output.

[0053] The analysis of voice quality is performed as follows, for example. That is, the biological reaction analysis unit 12 identifies the acoustic characteristics of the voice by performing known voice analysis processing on the voice for a specified time (for example, a time of about 30 to 150 seconds). Then, based on the acoustic characteristics, it analyzes whether a positive voice quality change has occurred, whether a negative voice quality change has occurred, and the magnitude of the voice quality change, and outputs a voice quality change index value corresponding to the analysis result.

[0054] The biological reaction analysis unit 12 calculates a biological reaction index value using at least one of the facial expression change index value, eye line change index value, pulse change index value, face orientation change index value, speech content index value, and voice quality change index value calculated as described above. For example, the biological reaction index value is calculated by performing weighted calculation on the facial expression change index value, eye line change index value, pulse change index value, face orientation change index value, speech content index value, and voice quality change index value.

[0055] The specific determination unit 13 determines whether the change in the biological reaction analyzed for the analysis target person is specific compared to the change in the biological reaction analyzed for other persons other than the analysis target person. In the present embodiment, the specific determination unit 13 determines whether the change in the biological reaction analyzed for the analysis target person is specific compared to others based on the biological reaction index values calculated for each of a plurality of users by the biological reaction analysis unit 12.

[0056] For example, the specific determination unit 13 calculates the variance of the biological reaction index values calculated for each of a plurality of persons by the biological reaction analysis unit 12, and determines whether the change in the biological reaction analyzed for the analysis target person is specific compared to others by comparing the biological reaction index value calculated for the analysis target person with the variance.

[0057] The following three patterns can be considered as cases where the change in the biological reaction analyzed for the subject to be analyzed is specific compared to others. The first case is when there is no particularly large change in the biological reaction for others, but a relatively large change in the biological reaction occurs for the subject to be analyzed. The second case is when there is no particularly large change in the biological reaction for the subject to be analyzed, but a relatively large change in the biological reaction occurs for others. The third case is when a relatively large change in the biological reaction occurs for both the subject to be analyzed and others, but the content of the change is different between the subject to be analyzed and others.

[0058] When a change in the biological reaction determined to be specific by the specificity determination unit 13 occurs, the related event identification unit 14 identifies an event occurring with respect to at least one of the subject to be analyzed, others, and the environment. For example, when a specific change in the biological reaction occurs for the subject to be analyzed, the related event identification unit 14 identifies the actions and words of the subject to be analyzed himself / herself from the moving image. In addition, the related event identification unit 14 identifies the actions and words of others when a specific change in the biological reaction occurs for the subject to be analyzed from the moving image. Further, the related event identification unit 14 identifies the environment when a specific change in the biological reaction occurs for the subject to be analyzed from the moving image. The environment is, for example, shared materials being displayed on the screen, things reflected in the background of the subject to be analyzed, and the like.

[0059] The clustering unit 15 analyzes the degree of correlation between a change in the biological reaction (for example, one or a combination of one or more of eye line, pulse, facial movement, speech content, voice quality) determined to be specific by the specificity determination unit 13 and an event (an event identified by the related event identification unit 14) that occurred when the specific change in the biological reaction occurred. When it is determined that the correlation is at a certain level or higher, the subject to be analyzed or the event is clustered based on the analysis result of the correlation.

[0060] For example, when a specific change in a biological reaction corresponds to a negative emotional change and the event occurring when such a specific change in the biological reaction occurs is also a negative event, a correlation equal to or higher than a certain level is detected. The clustering unit 15 clusters the subject or event to be analyzed into one of a plurality of pre-segmented classifications according to the content of the event, the degree of negativity, the magnitude of the correlation, and the like.

[0061] Similarly, when a specific change in a biological reaction corresponds to a positive emotional change and the event occurring when such a specific change in the biological reaction occurs is also a positive event, a correlation equal to or higher than a certain level is detected. The clustering unit 15 clusters the subject or event to be analyzed into one of a plurality of pre-segmented classifications according to the content of the event, the degree of positivity, the magnitude of the correlation, and the like.

[0062] The analysis result notification unit 16 notifies at least one of the change in the biological reaction determined to be specific by the specificity determination unit 13, the event specified by the related event identification unit 14, and the classification clustered by the clustering unit 15 to the designated person of the subject to be analyzed (the subject to be analyzed or the organizer of the online session).

[0063] For example, when a specific change in a biological reaction different from others occurs in the subject to be analyzed (any of the three patterns described above; the same applies hereinafter), the analysis result notification unit 16 notifies the subject to be analyzed of his or her own speech and actions as the event occurring at that time. Thereby, the subject to be analyzed can grasp that he or she has different emotions from others when performing a certain speech or action. At this time, the specific change in the biological reaction specified for the subject to be analyzed may also be notified to the subject to be analyzed. Further, the change in the biological reaction of the other person to be compared may be further notified to the subject to be analyzed.

[0064] For example, when there is a difference between the emotion received by others in response to the words and deeds of the person to be analyzed that were performed without particular awareness with the usual emotions, or the words and deeds that were performed with particular awareness with a certain emotion, and the emotion that the person to be analyzed himself / herself had at the time of the words and deeds, the words and deeds of the person to be analyzed himself / herself at that time are notified to the person to be analyzed. As a result, it is also possible to discover words and deeds that are well-received by others against one's own will, or words and deeds that are not well-received by others, etc.

[0065] In addition, when a change in a specific biological reaction different from others occurs in the person to be analyzed, the analysis result notification unit 16 notifies the event that is occurring together with the change in the specific biological reaction to the organizer of the online session. As a result, the organizer of the online session can know what events are affecting what changes in emotions as specific phenomena unique to the specified person to be analyzed. And it becomes possible to take appropriate measures for the person to be analyzed according to the grasped content.

[0066] In addition, when a change in a specific biological reaction different from others occurs in the person to be analyzed, the analysis result notification unit 16 notifies the event that is occurring or the clustering result of the person to be analyzed to the organizer of the online session. As a result, the organizer of the online session can grasp the tendency of the behavior unique to the specified person to be analyzed, or predict the behavior and state that may occur in the future, etc., depending on which classification the specified person to be analyzed is clustered into. And it becomes possible to take appropriate measures for the person to be analyzed accordingly.

[0067] Note that in the above embodiment, an example has been described in which a biological reaction index value is calculated by quantifying a change in a biological reaction according to a predetermined standard, and it is determined whether the change in the biological reaction analyzed for the person to be analyzed is specific compared to others based on the biological reaction index values calculated for each of a plurality of people, but it is not limited to this example. For example, it may be as follows.

[0068] That is, the biological reaction analysis unit 12 analyzes the movement of the line of sight for each of a plurality of persons and generates a heat map indicating the direction of the line of sight. The specific determination unit 13 determines whether the change in the biological reaction analyzed for the person to be analyzed is specific compared to the change in the biological reaction analyzed for others by comparing the heat map generated for the person to be analyzed by the biological reaction analysis unit 12 with the heat map generated for others.

[0069] In this way, in the present embodiment, the moving image of the video session is stored in the local storage of the user terminal 10, and the above-described analysis is performed on the user terminal 10. Although it may depend on the machine specifications of the user terminal 10, it is possible to perform the analysis without providing the information of the moving image to the outside.

[0070] <Functional Configuration Example 2> As shown in FIG. 5, the video session evaluation system of the present embodiment may include a moving image acquisition unit 11, a biological reaction analysis unit 12, and a reaction information presentation unit 13a as functional configurations.

[0071] The reaction information presentation unit 13a presents information indicating the change in the biological reaction analyzed by the biological reaction analysis unit 12a including participants not displayed on the screen. For example, the reaction information presentation unit 13a presents information indicating the change in the biological reaction to the leader, facilitator, or administrator of the online session (hereinafter collectively referred to as the host). The host of the online session is, for example, a lecturer of an online class, a chairperson or facilitator of an online meeting, a coach of a session for coaching, etc. The host of the online session is usually one of the plurality of users participating in the online session, but may be another person not participating in the online session.

[0072] By doing so, the host of the online session can also grasp the state of the participants not displayed on the screen in an environment where the online session is held by a plurality of persons.

[0073] <Functional Configuration Example 3> FIG. 6 is a block diagram showing a configuration example according to the present embodiment. As shown in FIG. 6, for the video session evaluation system of the present embodiment, for functions similar to those in the above-described Embodiment 1 in terms of functional configuration, the same reference numerals may be used and the description may be omitted.

[0074] The system according to the present embodiment includes a camera unit that acquires video of a video session and a microphone unit that acquires audio, an analysis unit that analyzes and evaluates moving images, an object generation unit that generates a display object (described later) based on information obtained by evaluating the acquired moving images, and a display unit that displays both the moving images of the video session and the display object during the execution of the video session.

[0075] Similar to the above description, the analysis unit includes a moving image acquisition unit 11, a biological reaction analysis unit 12, a specific determination unit 13, a related event identification unit 14, a clustering unit 15, and an analysis result notification unit 16. The functions of each element are as described above.

[0076] As shown in FIG. 7, based on the result of analyzing the moving images acquired from the video session by the analysis unit, the object generation unit superimposes and displays, as necessary, an object 50 indicating the recognized face portion and information 100 indicating the above-described analyzed and evaluated content on the moving images. When there are multiple people's faces moving in the moving images, the object 50 may identify and display the faces of all the multiple people.

[0077] Also, for example, when the camera function of the video session is stopped on the other party's terminal (i.e., not physically covering the camera, etc., but software - stopped within the video - session application), if the other party's face was recognized by the other party's camera, the object 50 or object 100 may be displayed at the portion where the other party's face is located. This enables both parties to confirm that the other party is in front of the terminal even when the camera function is off. In this case, for example, in the video - session application, while the information acquired from the camera is made non - visible, only the object 50 or object 100 corresponding to the face recognized by the analysis unit may be displayed. Also, the video information acquired from the video session and the information that can be recognized by the analysis unit may be separated into different display layers, and the layer related to the former information may be made non - visible.

[0078] When there are areas for displaying a plurality of moving images, the object 50 or object 100 may be displayed only in all areas or some of the areas. For example, as shown in FIG. 8, it may be displayed only in the moving image on the guest side.

[0079] As described above, the preferred embodiments of the present disclosure have been described in detail with reference to the accompanying drawings. However, the technical scope of the present disclosure is not limited to such examples. It is obvious that those having ordinary knowledge in the technical field of the present disclosure can come up with various modification examples or correction examples within the scope of the technical idea described in the claims, and these are of course understood to belong to the technical scope of the present disclosure.

[0080] The devices described in this specification may be realized as a single device, or may be realized by a plurality of devices (e.g., cloud servers) partially or entirely connected by a network. For example, the control unit 110 and the storage 130 of each terminal 10 may be realized by different servers connected to each other by a network.

[0081] That is, this system includes user terminals 10 and 20, a video session service terminal 30 that provides a two-way video session to the user terminals 10 and 20, and an evaluation terminal 40 that evaluates the video session. In this case, the following variation combinations of configurations can be considered. (1) Process everything only with user terminals As shown in FIG. 9, by performing the processing by the analysis unit at the terminal where the video session is being conducted (although a certain processing capacity is required), it is possible to obtain analysis and evaluation results in real time simultaneously with the time of the video session. (2) Process with user terminals and evaluation terminals As shown in FIG. 10, it may be possible to provide the evaluation terminal connected by a network or the like with an analysis unit. In this case, the moving image acquired by the user terminal is shared with the evaluation terminal simultaneously with or after the video session, and after being analyzed and evaluated by the analysis unit in the evaluation terminal, the information of the object 50 and the object 100 is shared with the user terminal together with the moving image data or separately (that is, at least the information including the analysis data) and displayed on the display unit.

[0082] <First Embodiment> With reference to FIGS. 11 and 12, the first embodiment of the present invention will be described. The system according to this embodiment generally analyzes and evaluates which part of the displayed material has been gazed at for how long from the information on where the gaze of the person to be evaluated is focused on the screen and the information on the material being displayed at that time.

[0083] That is, the system according to this embodiment has a camera means for acquiring a moving image obtained by photographing the person to be evaluated, a gaze acquisition means for acquiring the movement of the gaze of the subject based on the acquired moving image, and a display means for continuously displaying a plurality of images to the subject.

[0084] In particular, this system has position acquisition means for acquiring the positional relationship between the camera means and the display means. As a result, it becomes possible to calibrate the eye movement and the fixation point of the subject. The calibration process is, for example, as shown in FIGS. 11 and 12, the state of the subject's eyes is acquired by the camera unit of the display, and then the eye movement when looking at a predetermined location (calibration points: center, four corners of the screen, etc.) on the screen is acquired. Regarding the acquisition of eye movement, for example, it may be possible to intentionally have the subject look at the calibration point by playing an announcement on the screen. It may also be possible to prominently display a sign only in the center in an eye-catching manner and estimate that the eye movement at that moment is in a state of gazing at the center.

[0085] As shown in FIG. 11, in this embodiment, the fixation point is associated with the fixation time on the material (shared) displayed on the screen and output like a heatmap. As a result, it is possible to grasp which part of the material was paused for how long, and the part of interest of the subject can be understood.

[0086] In addition, the system according to this embodiment may generate a heatmap considering the movements of other subjects (other business partners, other course participants, etc.) for the same material. In this case, the fixation points of other subjects may also be displayed on the material. At this time, it may be possible to output specific characteristics unique to the subject, such as parts that the subject did not look at even though other subjects were looking at them, or parts that the subject looked at even though other subjects did not look at them.

[0087] Note that as a method for evaluating the material itself, for each material, a normalized heatmap obtained by normalizing the movement of the subject's line of sight may be associated and output. For example, the necessity of the material can be grasped from the perspective of which material was looked at well. On the other hand, it can also be understood that materials with short fixation times are not very necessary.

[0088] <Second Embodiment> The second embodiment of the present invention will be described with reference to FIGS. 13 to 14. The system according to this embodiment visualizes the evaluation values analyzed and evaluated based on the above-described expressions and voices as a graph, and extracts other subjects having the same pattern as the pattern readable from the graph.

[0089] That is, the system according to this embodiment outputs the evaluation values calculated by recognizing the face images and voices of the subjects included in the acquired moving image as change information along the time series (for example, like a graph shown in FIG. 13).

[0090] As shown in FIG. 14, for example, in the graphs of the evaluated subjects A to C, the value "safety" indicating a sense of security will be described as an example. The illustrated graph plots the time axis on the horizontal axis and the evaluation value indicating the degree of security on the vertical axis. It can be seen that the value of the graph (A) indicating the subject A drops significantly from time t1 to t2. For example, when it is detected by comprehensively judging the expression and voice that A felt anxious or scared, it appears as such a graph.

[0091] Similarly, in the graph (B) representing the subject B, it can be seen that the value drops significantly from time t1 to t2. For example, when it is detected by comprehensively judging the expression and voice that A felt anxious or scared, it appears as such a graph.

[0092] Although the degree of security of both A and B drops significantly from time t1 to t2, there is no significant change in the degree of security of the subject C (C) from time t1 to t2.

[0093] When such a graph is obtained, the system extracts the information of (B) including the same pattern with reference to the change from time t1 to t2 of (A).

[0094] In this way, by extracting similar patterns, it becomes possible to extract subjects who have similar emotions in the same time period. Also, it becomes possible to extract subjects who have opposite emotions in the same time period.

[0095] When extracting similar patterns, a selection operation of a part of the graph information as the source from the analyst may be received (for example, by selecting the time t1 to t2 in (A) of FIG. 17), and the corresponding frames of other moving images including the graph pattern identical to the graph pattern of the selected part may be specified.

[0096] <The Third Embodiment> Referring to FIGS. 15 and 16, the third embodiment of the present invention will be described. The system according to this form detects that the above-described facial expressions and voices have instantaneously and greatly changed beyond a predetermined threshold. In particular, by detecting that only one of the evaluations from a plurality of viewpoints has greatly changed, it becomes possible to analyze the deep psychology of the subject. By cutting out and evaluating the video in which such a change has occurred, the analysis can also be efficiently performed.

[0097] As shown in FIG. 14, for the graphs of "safety" indicating the degree of security and "happy" indicating the degree of happiness of the subject, there is no significant change in the graph of safety, but only the graph of happy drops significantly at times t1 and t2 and then returns to the original level immediately afterwards.

[0098] This system cuts out and combines a predetermined length portion L1 before and after including the change at t1 and a predetermined length portion L2 before and after including the change at t2 in this way to generate a digest video. Thereby, it becomes possible to extract a video including the moment when the deep psychology appears.

[0099] As shown in FIG. 16, the system according to this embodiment may also be configured to detect that two different graphs change significantly instantaneously. For example, as shown in FIG. 16, in the evaluation values associated with mutually opposite characteristics such as happy and sad, at time t1, happy instantaneously drops and sad instantaneously rises. Thus, when one evaluation value instantaneously changes and the other opposite evaluation value instantaneously rises, the rising emotion is often the true emotion.

[0100] Note that even if the characteristics of the evaluation values are not opposite, by identifying the evaluation value that has risen significantly when some evaluation value has dropped significantly, in-depth psychological analysis can be readily performed.

[0101] Note that a plurality of frames (specific frames) obtained from a moving image may be concatenated to generate a digest moving image. The voice corresponding to the specific frame may be converted into text and output. When it is possible to share the screen information displayed on the screen of the user terminal, the screen information corresponding to the specific frame may be output when there is the above-described instantaneous change.

[0102] <Fourth Embodiment> Referring to FIG. 17, the system according to the fourth embodiment of the present invention will be described. The system according to this embodiment can detect false emotions or communication such as flattery, for example, when the emotion "happy" drops extremely even though a positive word such as "thank you" is being uttered.

[0103] That is, the system according to this embodiment associates attributes with the facial expression evaluation value and the voice evaluation value in advance, and detects that the change exceeds a predetermined range from the correlation relationship between the attributes. For example, positive labels are associated with words such as "thank you" and "I understand well", and the correlation relationship with the evaluation of facial expressions (happy, sad, safety) is defined in advance.

[0104] As shown in FIG. 17, in the graph of (A), words such as "Thank you." and "I understand very well." are extracted, and the value of the happy graph is high. On the other hand, in (B), although saying "Thank you for the explanation." and "I could understand it very well.", happy has dropped significantly after time t1. Here, although "I could understand it very well." is defined as a positive word, the corresponding happy emotion has dropped significantly.

[0105] The system shown in FIG. 18 can share the screen information displayed on the screen. In this case, it may be possible to associate the content of the screen, the text information, and the graph information indicating the emotion as described above.

[0106] <The Fifth Embodiment> The system according to the fifth embodiment of the present invention will be described. The system according to this time difference form analyzes the standard expression of the user from the recognized face image and analyzes and evaluates it from three viewpoints: emotion, concentration, and safety.

[0107] That is, evaluation is performed from three viewpoints: emotion that evaluates the degree of deviation from the standard expression of the user, concentration that evaluates at least the amount of movement of the user's pupils or face movement from the recognized face image, and safety that evaluates the emotion related to the user's anxiety from the recognized face image.

[0108] The evaluation may use a learning device that has learned each viewpoint, or may be performed by other methods. A score is generated based on two or more evaluations among the evaluated viewpoints.

[0109] <The Sixth Embodiment> Referring to FIG. 19, a system according to the sixth embodiment of the present invention will be described. The system according to this embodiment is, for example, for identifying a video in which a specific target person appears from a plurality of business videos, lecture videos, etc. This makes it possible to focus on and evaluate a specific person from various sessions conducted online.

[0110] As shown in the figure, when there is video data related to an online lecture by a certain lecturer, usually, if there is no labeling, title, etc. associated with it, it is impossible to determine whether a specific person appears (such as whether it is the lecture of that lecturer) until the content is played. According to this embodiment, it is possible to analyze within each video file and detect whether a person whose face image has been registered in advance appears in the video.

[0111] Note that in the moving image, it may be possible to cut out only the part where the target person appears to generate a digest video. For example, if among the lecture videos 001 to 004 of the moving image shown, the moving images in which the target person appears are lecture 001, lecture 002, and lecture 004, then this system extracts the parts of t1, t2, and t3 from each moving image. The extracted parts can be played as a digest moving image.

[0112] <Seventh Embodiment> Referring to FIG. 20, a system according to the seventh embodiment of the present invention will be described. The system according to this embodiment calculates the so-called concentration score (degree of concentration) of a target person participating in an online session. In the case of online, especially in the case of a webinar format, the cameras of the listeners are often turned off. According to this system, in such a case, it is possible to quantitatively determine how concentrated each listener is on the lecture.

[0113] This system recognizes the face images captured by the camera during a session (regardless of whether or not to share one's own camera image with others), and evaluates the amount of movement of the target person's pupils and face respectively. For example, as shown in FIG. 20, the amounts regarding how much the face has moved and how much the eyes have moved from the initial position are evaluated as absolute values. For example, in the period L1 of the illustrated graph, it can be seen that both the movement of the face and the movement of the eyes are small, and it is presumed that the person is staring at the screen or the like. On the other hand, in the period L2, although the face has not moved, the eyes have moved in various directions, and it is presumed that the person is reading the contents within the material. The degree of concentration is grasped as these two patterns, and in the former case, it can be presumed that the person is staring at the face of the lecturer and listening intently to the speech, and in the latter case, it can be presumed that the person is concentrating on reading the displayed materials or the like.

[0114] Also, a score regarding the degree of concentration may be calculated based on the degree of movement (Value in the illustrated graph). Various forms of calculation methods, statistical methods can be selected. For example, when both face and eyes are 0, the degree of concentration may be set to 100, and when both are at the maximum value, the degree of concentration may be set to 0.

[0115] <Eighth Embodiment> With reference to FIG. 21, a system according to the eighth embodiment of the present invention will be described. The system according to this embodiment attempts to obtain a true evaluation by correcting seasonal and time factors by performing the evaluations performed by the system in the above-described first to seventh embodiments over different spans.

[0116] That is, this system evaluates the state of the target person based on a short-term evaluation value (analysis value) obtained by analyzing the evaluation values in a short period and a long-term evaluation value (analysis value) obtained by analyzing the evaluation values in a long period for the same user. For example, it may be to analyze the evaluation values (long-term evaluation values) of a certain student over one year and the evaluation values (short-term evaluation values) on a monthly basis.

[0117] For example, as the content analyzed from the long-term evaluation value, the long-term characteristics of the subject can be analyzed, such as communication becoming dull in winter, being lively in summer, and the number of communications being extremely small on winter nights. On the other hand, as the content analyzed from the short-term evaluation value, the short-term characteristics of the subject, such as being in a depressed mood at the end of the month and having a bright expression on Friday, can be analyzed.

[0118] The above-mentioned long term is, for example, a cycle of 3 months, half a year, 1 year, etc., and the short term is, for example, a cycle of 1 day, 1 week, 1 month, etc., but it is not limited to this, and it may be evaluated in two spans with different periods. In this case, the evaluation may adopt the average value or median value of the evaluation value, or appropriate values may be calculated by various statistical methods.

[0119] In the present invention, when analyzing the trend of the above-mentioned evaluation value, for example, when an evaluation that the mood drops at the end of the month is made, it may be corrected by multiplying the happiness score generated at the end of the month by a predetermined coefficient. That is, when an evaluation different from the trend occurs, the true feelings can be analyzed by performing the evaluation with a higher weighting than the evaluation.

[0120] For example, as shown in FIG. 21, for a subject with a trend of happiness (happy) indicated by a solid line, assume that the value in the latter half of February (Feb.) is P1. Although the happy score should originally be low at the end of the month, it can be seen that it deviates from the trend. In this case, the evaluation value may be treated as something that should be specially evaluated by correcting (adding the deviation) the evaluation value by the deviation from the trend. The correction method may be to add or subtract the positive or negative deviation from the trend as an addition circle or subtraction, or other methods may also be used.

[0121] <The Ninth Embodiment> A system according to the ninth embodiment of the present invention will be described. The system according to this embodiment acquires subjective responses (such as questionnaires and interviews) from the target person in advance regarding a certain theme, and compares them with the acquired evaluation values. Thereby, for example, even if a subjective response such as "no dissatisfaction with the lecture" is obtained in a questionnaire or the like, when the actual facial expressions and voices are analyzed and the happiness evaluation is low, it can be grasped that some speculation is at work. The system according to this embodiment is particularly suitable in the field of awareness surveys from a company to employees.

[0122] For example, as a questionnaire for the target person, questions regarding happiness (job satisfaction, ventilation, etc.), questions regarding anxiety (things that are difficult, things that are feared, etc.), and questions regarding future safety (securing a career path, promotion, salary increase, etc.) may be prepared and answered, and the evaluation values regarding happiness, anxiety, and safety in the facial expressions and voices of the target person may be compared.

[0123] By performing such a comparison on all employees of the organization to which the target person belongs, the true feelings of the employees can be analyzed, and the evaluation of the entire organization becomes possible.

[0124] <Tenth Embodiment> The tenth embodiment of the present invention will be described with reference to FIG. 22. The system according to this embodiment accepts, after the fact, labeling (annotation) regarding the situation at that time from the target person for the evaluation value obtained from the facial expressions and voices of the target person. Thereby, the evaluation value can be subjectively evaluated after the fact, and the algorithm can also be updated by feeding back the evaluation result.

[0125] As shown in the figure, for the graph of happiness, it may be labeled for each time period such as Labels Lab.1 to Lab.4, and the evaluation value and the content of the received label may be superimposed and displayed.

[0126] As the content of labeling, for example, various information such as "correct / incorrect about the evaluation value", "the situation at that time", and "the numerical value based on self-standards" can be added.

[0127] <The 11th Embodiment> The system according to the 11th embodiment of the present invention will be described. The system according to this embodiment relates to a business video session conducted between a business-side terminal of a business-side person in charge and a customer-side terminal of a customer-side person in charge. Conventionally, a salesperson has sometimes made a prediction about the closing rate based on the feel and experience of a confident sales interview or the like, and calculated an expected value by multiplying the transaction amount of the sales case by the predicted closing rate to formulate a sales plan. According to this system, by analyzing the business video session, the closing rate can be presented by machine learning statistical processing based on the data of past business video sessions and their closing results.

[0128] This system has a closing information acquisition means for acquiring the closing information of past business video sessions, a face recognition means for recognizing at least one of the face images of the business-side person in charge or the customer-side person in charge included in the moving image of the business video session for each predetermined frame, a voice recognition means for recognizing at least one of the voices of the business-side person in charge or the customer-side person in charge included in the moving image, and an evaluation means for calculating evaluation values from a plurality of viewpoints based on both the recognized face image and voice. In particular, this system is provided with a model generation means for generating a closing estimation model for estimating the closing rate of the moving image of another business video session using the recorded evaluation value and the closing information as teacher data, and determines the closing rate of a new business video session using this model.

[0129] The conversion rate may be a numerical value such as, for example, 50% or 70%, or may be ranks (zones) such as A, B, and C. The conversion rate may be calculated based on the similarity to the past converted moving images. For example, if the similarity between the moving image of a new sales video session and the converted moving image of the same (or similar) past sales destination is 70%, it may be determined that the conversion rate of the new sales is also 70%.

[0130] Based on the estimated transaction amount related to the sales video session and the estimated conversion rate (rank) within the organization, this system calculates the numerical value of the expected sales prospect for a predetermined period such as the current month or per quarter.

[0131] <The 12th Embodiment> Referring to FIG. 23, a system according to the 12th embodiment of the present invention will be described. The system according to this embodiment is mainly suitable for the use of online learning guidance. This system includes an instructor terminal and a student terminal that are communicably connected to each other via a network.

[0132] The instructor terminal has an instructor-side camera for capturing at least the face of the instructor user. As shown in the figure, the student terminal has a face camera for capturing at least the face of the student and a hand camera for capturing the student's hand (while writing in notes or prints, or the state on the desk).

[0133] This system particularly includes hand motion recognition means for recognizing the motion of the student's hand in the moving image obtained from the hand camera for each predetermined frame, and estimation means for estimating the understanding degree of the student based on the recognized hand motion.

[0134] The estimation means estimates the understanding degree of the student according to the amount of the recognized hand motion. For example, evaluations can be made from the perspective of whether the amount of writing on the blackboard matches the amount of writing on the instructor's side (whether the notes are taken properly), or whether the student is making efforts such as color-coding important points by analyzing the color of the pen being used.

[0135] In addition, even when no face image is captured by the face camera (i.e., the user is looking down), it is also possible to detect situations where hand movements are not detected (not taking notes on the blackboard, dozing off), etc., and issue an alert to the instructor terminal or the attendee terminal.

[0136] As described above, it is also possible to estimate the degree of understanding based on the evaluation value based on the expression and voice of the attendee and the hand movements. In this case, if it is detected that the evaluated attendee is irritated or feels anxious, and at the same time, no hand movement of the attendee can be detected by the hand camera, it can be presumed that the lecture is not being conducted effectively.

[0137] <Supplement to the Hardware Configuration> A series of processes performed by the apparatus described in this specification may be realized using any of software, hardware, and a combination of software and hardware. It is possible to create a computer program for realizing each function of the information sharing support apparatus 10 according to this embodiment and install it on a PC or the like. It is also possible to provide a computer-readable recording medium storing such a computer program. The recording medium is, for example, a magnetic disk, an optical disk, a magneto-optical disk, a flash memory, or the like. Further, the above computer program may be distributed via a network, for example, without using a recording medium.

[0138] Also, the processes described using flowcharts in this specification do not necessarily have to be executed in the order shown in the figures. Some processing steps may be executed in parallel. Additionally, additional processing steps may be adopted, and some processing steps may be omitted.

[0139] It is also possible to implement the embodiments described above by appropriately combining them. Further, the effects described in this specification are merely illustrative or exemplary and not limiting. That is, the technology according to the present disclosure may exhibit other effects apparent to those skilled in the art from the description of this specification, together with or instead of the above effects.

Explanation of Reference Numerals

[0140] 10, 20 User terminals 30 Video session service terminal 40 Evaluation terminal

Claims

1. Acquisition means for acquiring a moving image related to a video session implemented between at least two user terminals; Facial recognition means for recognizing a user's facial image included in the moving image for each predetermined frame; Voice recognition means for recognizing at least the voice of the user included in the moving image; Facial expression evaluation means for calculating facial expression evaluation values from a plurality of viewpoints based on the recognized facial image; Voice evaluation means for calculating voice evaluation values from a plurality of viewpoints based on the recognized voice; Facial expression and voice correlation evaluation means for evaluating the correlation between the facial expression evaluation value and the voice evaluation value of at least one of the users; Detection means for detecting that the facial expression evaluation value and the voice evaluation value have changed beyond a predetermined range from the correlation; Specific frame acquisition means for acquiring a specific frame including the detected range, comprising: Further comprising attribute evaluation means for associating attributes corresponding to each of the facial expression evaluation value and the voice evaluation value; The detection means detects that the attribute of the facial expression evaluation value and the attribute of the voice evaluation value are contrary to each other; The detection means detects that a change in the facial expression evaluation value or the voice evaluation value related to one of the plurality of viewpoints is greater than or equal to a predetermined amount compared to a change in the facial expression evaluation value or the voice evaluation value related to another viewpoint. Moving image analysis system.

2. The moving image analysis system according to claim 1, Further comprising digest video generation means for concatenating the plurality of specific frames acquired from the moving image to generate a digest video. Moving image analysis system.

3. The moving image analysis system according to claim 1 or claim 2, Further comprising specific frame corresponding text output means for converting the voice corresponding to the specific frame into text and outputting the text. Moving image analysis system.

4. The moving image analysis system according to any one of claims 1 to 3, The video session is capable of sharing screen information displayed on the screen of one user terminal, Further comprising shared screen output means for outputting at least the screen information corresponding to the shared specific frame. Moving image analysis system.

Citation Information

Patent Citations

  • Feeling analyzing system

    JP2000076421A

  • Content retrieval / recommendation method, content retrieval / recommendation device, and content retrieval / recommendation program

    JP2008204193A

  • Expression change analysis system

    JP2011154665A

  • Emotion estimation device and emotion estimation method

    JP2011186521A

  • Face expression amplification device, expression recognition device, face expression amplification method, expression recognition method and program

    JP2012008949A